10 interesting stories served every morning and every evening.

Introducing Claude Opus 5

www.anthropic.com

Claude Opus 5 is avail­able to­day. It’s a thought­ful and proac­tive model that comes close to the fron­tier in­tel­li­gence of Claude Fable 5 at half the price.

On cod­ing and knowl­edge work eval­u­a­tions like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it re­mains be­hind Mythos 5 on cy­ber­se­cu­rity tasks.

Opus 5 is de­signed to be used every day: it works more ef­fi­ciently than other mod­els. It’s the new de­fault model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-ef­fec­tive­ness

Claude Opus 5 pro­vides greatly im­proved per­for­mance for the same cost as its pre­de­ces­sor, Opus 4.8. The charts in this sec­tion show how per­for­mance changes ac­cord­ing to the mod­el’s ef­fort set­ting, which cus­tomers can use to op­ti­mize for in­tel­li­gence or con­serve to­kens for faster and cheaper re­sults.

Opus 5 ex­cels on valu­able soft­ware en­gi­neer­ing tasks. For ex­am­ple, on Frontier-Bench v0.1, Opus 5 sur­passes all other mod­els, and more than dou­bles Opus 4.8’s per­for­mance at a lower cost per task. On CursorBench 3.2, at max ef­fort, the model per­forms within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater per­for­mance at a given cost than all other mod­els on high, xhigh, and max ef­fort.

We see sim­i­lar re­sults on knowl­edge work and prob­lem-solv­ing tasks. For ex­am­ple:

On ARC-AGI 3, an eval­u­a­tion where the model has to solve novel prob­lems, Opus 5’s score is three times as high as the next-best model.

On Zapier AutomationBench, which mea­sures whether mod­els can com­plete busi­ness tasks from start to fin­ish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its low­est ef­fort set­ting, Opus 5 passes more tasks than any other model.

On OSWorld 2.0, a com­puter use bench­mark, Opus 5 out­per­forms every other model at any given cost, sur­pass­ing Fable 5’s best re­sult at just over a third of the cost.

It’s also our best and most cost-ef­fi­cient model on sev­eral re­lated eval­u­a­tions:

Opus 5 is a mean­ing­ful im­prove­ment over Opus 4.8 for sci­en­tific re­search. It shows bet­ter per­for­mance than Opus 4.8 on every one of our life sci­ences eval­u­a­tions, which cover top­ics in­clud­ing struc­tural bi­ol­ogy, or­ganic chem­istry, and bioin­for­mat­ics. Its im­prove­ments are most no­table on or­ganic chem­istry tasks, like in­fer­ring mol­e­c­u­lar struc­tures from spec­troscopy data (it scores 10.2 per­cent­age points higher than Opus 4.8 on our in­ter­nal bench­mark), and on pro­tein-re­lated tasks like pre­dict­ing how vari­a­tions in a pro­tein’s se­quence af­fect how it func­tions (here, it scores 7.7 per­cent­age points higher).

Finally, Opus 5 is ca­pa­ble of pro­duc­ing much stronger vi­sual out­puts:

Working with Claude Opus 5

Claude Opus 5 is much stronger at ver­i­fy­ing its work and it­er­at­ing care­fully un­til it suc­ceeds. In eval­u­a­tions and early-ac­cess test­ing, we and our users found many ex­am­ples of Opus 5’s agency and thor­ough­ness:

On one Frontier-Bench task, Opus 5 was given a draw­ing of a ma­chine part and asked to write code to re­build it as a 3D FreeCAD model. However, in this task, the model was in­ten­tion­ally given no way to di­rectly view the draw­ing. Opus 5 re­sponded by writ­ing its own com­puter vi­sion pipeline to pull the geom­e­try from the raw pix­els, then re­con­structed the full ma­chine part. It suc­ceeded in do­ing so re­peat­edly; no com­pet­ing model with the same setup could solve it af­ter five at­tempts.

Given a real bug in a pop­u­lar open-source pack­age man­ager, Opus 5 found the root cause and fixed an edge case that the com­mu­ni­ty’s patch had missed. A com­pet­ing model fixed only the sur­face symp­tom (not the un­der­ly­ing cause), then re­ported the bug re­solved.

An en­gi­neer at a trad­ing firm used Opus 5 to build a mar­ket data feed for a new ex­change in a sin­gle ses­sion. Previous mod­els could not com­plete this task at all, even given ex­ten­sive plans from the en­gi­neer. Finding no live feed to val­i­date against, Opus 5 even built its own test har­ness to check that its code parsed the ex­change’s data cor­rectly.

Below are fur­ther re­ports from our early-ac­cess cus­tomers on their ex­pe­ri­ence of work­ing with Opus 5:

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 lets mon­i­tor­ing agents man­age parts of their own mem­ory in pro­duc­tion, mak­ing them more au­tonomous and re­li­able over longer hori­zons. The agent treats its con­text as a liv­ing doc­u­ment: af­ter flag­ging a po­ten­tial anom­aly in one of our ser­vices, it re-checked its own as­sump­tion against pro­duc­tion, found the sig­nal was be­nign, wrote the cor­rec­tion into its mem­ory, and re­tired its mon­i­tor­ing queries on its own.

Claude Opus 5 lets mon­i­tor­ing agents man­age parts of their own mem­ory in pro­duc­tion, mak­ing them more au­tonomous and re­li­able over longer hori­zons. The agent treats its con­text as a liv­ing doc­u­ment: af­ter flag­ging a po­ten­tial anom­aly in one of our ser­vices, it re-checked its own as­sump­tion against pro­duc­tion, found the sig­nal was be­nign, wrote the cor­rec­tion into its mem­ory, and re­tired its mon­i­tor­ing queries on its own.

Claude Opus 5 is a strong agen­tic cod­ing model built for long-run­ning, multi-step work. It deeply un­der­stands your code­base, holds the thread across com­plex tasks, and pins down re­quire­ments for fea­ture de­vel­op­ment and bug-fix­ing more ef­fec­tively than Opus 4.8. Developers can now build with Opus 5 in Kiro, ac­cess­ing its ad­vanced ca­pa­bil­i­ties to tackle am­bi­tious pro­jects.

Claude Opus 5 is a strong agen­tic cod­ing model built for long-run­ning, multi-step work. It deeply un­der­stands your code­base, holds the thread across com­plex tasks, and pins down re­quire­ments for fea­ture de­vel­op­ment and bug-fix­ing more ef­fec­tively than Opus 4.8. Developers can now build with Opus 5 in Kiro, ac­cess­ing its ad­vanced ca­pa­bil­i­ties to tackle am­bi­tious pro­jects.

01 /

24

Alignment and safety

Alignment. During pre-de­ploy­ment test­ing, our au­to­mated be­hav­ioral au­dit found Opus 5 to be our most aligned model to date (as shown in the graph be­low). It ad­heres to Claude’s Constitution bet­ter than Opus 4.8, Sonnet 5, or Fable 5; ex­hibits the low­est rates of de­cep­tive be­hav­ior; and is the least sus­cep­ti­ble to be­ing tricked into mis­use. It’s also our safest model yet in terms of avoid­ing reck­less ac­tions that could have hard-to-re­verse side ef­fects.

Safety. Opus 5 does not ad­vance the fron­tier in risky, dual-use ca­pa­bil­i­ties. In rig­or­ous eval­u­a­tions con­ducted along­side pri­vate-sec­tor and gov­ern­ment part­ners, we found it re­mains be­hind Mythos 5 in both bi­ol­ogy re­search and of­fen­sive cy­ber­se­cu­rity. More in­for­ma­tion about these eval­u­a­tions can be found in our System Card.

As with its pre­de­ces­sor, Opus 4.8, we’ve in­ten­tion­ally avoided train­ing Opus 5 on cy­ber tasks. The model has nev­er­the­less im­proved sub­stan­tially on these tasks as a re­sult of be­com­ing more gen­er­ally ca­pa­ble, and it comes close to Mythos 5 at find­ing cy­ber­se­cu­rity vul­ner­a­bil­i­ties. However, it re­mains sub­stan­tially be­hind Mythos 5 on the ex­ploita­tion of those vul­ner­a­bil­i­ties—that is, in turn­ing vul­ner­a­bil­i­ties into ma­te­r­ial cy­ber threats.

This is il­lus­trated by Opus 5’s per­for­mance on OSS-Fuzz, an eval­u­a­tion we’ve de­vel­oped to as­sess how well mod­els can find and then ex­ploit vul­ner­a­bil­i­ties with­out ex­ten­sive hu­man guid­ance. Although Mythos 5 and Opus 5 iden­tify vul­ner­a­bil­i­ties with sim­i­lar suc­cess, Opus 5’s score on the de­vel­op­ment of ex­ploits is far be­hind that of Mythos 5.

Safeguards for Opus 5

Claude Opus 5’s safe­guards are de­signed to al­low ben­e­fi­cial uses of the model in both cy­ber­se­cu­rity and bi­ol­ogy. They are sim­i­lar to those we ap­plied to Opus 4.8, with the ex­cep­tion of some stronger guardrails on a nar­row range of cy­ber tasks.

Cybersecurity. Opus 5’s cy­ber clas­si­fiers are pro­por­tion­ally less re­stric­tive than those on Fable 5. They al­low Opus 5 to find vul­ner­a­bil­i­ties in source code, but block binary-based” vul­ner­a­bil­ity scan­ning (a method more likely to be as­so­ci­ated with ma­li­cious ac­tors), pen­e­tra­tion test­ing, and ex­ploit gen­er­a­tion.

Based on our test­ing, we ex­pect the clas­si­fiers to in­ter­vene around 85% less of­ten than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged re­quests will fall back to Opus 4.8 by de­fault. Fallbacks to Opus 4.8 can also be en­abled on the API.

Our Cyber Verification Program (CVP) fa­cil­i­tates cy­ber­se­cu­rity work that would oth­er­wise be im­peded by the mod­el’s safe­guards. Enterprises and re­searchers who are al­ready part of the CVP have im­me­di­ate ac­cess to a ver­sion of Opus 5 with fewer se­cu­rity re­stric­tions.

Biology. Since Opus 5 has a sim­i­lar suite of safe­guards to Opus 4.8, it is now our most ca­pa­ble gen­er­ally avail­able model for sci­en­tific re­search. Nevertheless, the model still shows im­por­tant lim­i­ta­tions on long-run­ning, au­tonomous re­search tasks, which is where we ex­pect AI mod­els to pose the most sub­stan­tial bi­ol­ogy-re­lated risks. (Mythos 5 re­mains the stronger model for this type of bi­o­log­i­cal work.) As part of this launch, bi­ol­ogy-re­lated re­quests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.

Getting started

Claude Opus 5 is avail­able to­day on all plat­forms, priced at $5 per mil­lion in­put to­kens and $25 per mil­lion out­put to­kens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.

It’s also of­fered in Fast mode, where it runs around 2.5 times the de­fault speed. As with Opus 4.8, Fast mode is avail­able at twice Opus 5’s base price on the Claude Platform and through us­age cred­its in Claude Code.

Alongside Opus 5, we’re re­leas­ing two up­dates in beta:

Mid-conversation tool changes on the Claude Platform. Within a con­ver­sa­tion, de­vel­op­ers can now change which tools Claude can use with­out in­val­i­dat­ing the prompt cache.

Automatic fall­backs on the API. Users can now choose to have re­quests that are flagged by our safety clas­si­fiers on Opus 5 (or Fable 5) au­to­mat­i­cally route to an­other model. With au­to­matic fall­backs on, API re­quests al­ways route to the best avail­able model by de­fault rather than be­ing blocked.

Consistent with prior Opus mod­els, Opus 5 does not have data re­ten­tion re­quire­ments for gen­eral ac­cess.

For more guid­ance on how to get the best out of Opus 5, see our prompt­ing guide.

Footnotes

Frontier-Bench v0.1, Effort plot: These re­sults are from an in­ter­nal run of Frontier-Bench v0.1, on the mini-SWE-agent har­ness and a GKE back­end, mean re­ward over 5 at­tempts per task. Opus 4.8 served as fall­back on safety-clas­si­fier re­fusals for Opus 5 and Fable 5.

Related con­tent

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

Ask Claude about the Anthropic Economic Index

We’re launch­ing the Anthropic Economic Index con­nec­tor for Claude, which lets any­one ex­plore real data about AI and work.

Read more

Anthropic is do­nat­ing an­other $20 mil­lion to Public First Action

Anthropic is con­tribut­ing an ad­di­tional $20 mil­lion to Public First Action, bring­ing our to­tal sup­port to $40 mil­lion.

Read more

openai.com

Writing by Hand is Good for your Brain - Here's how to do it

nealstephenson.substack.com

Because I am known to write us­ing a foun­tain pen on pa­per, a num­ber of peo­ple have pointed me to this post and its un­der­ly­ing re­search. I won’t re­hash what is said in those sources, but the gist of it is that when you write things down by hand you’re re­cruit­ing more of your brain, which is a good thing.

I’m not an ex­pert on how the brain works, but I can say that, when writ­ing by hand, one is con­tin­u­ally solv­ing a se­ries of small prob­lems hav­ing to do with the spac­ing of words, how let­ters are con­nected, the cross­ing of the let­ter t (sometimes more than one in the same word) and the dot­ting of the let­ters i and j, and how to ac­com­plish all of those things through co­or­di­nated move­ments not just of the fin­gers but of the whole arm. All of that has to be in­te­grated in real time with what­ever is hap­pen­ing on a more ab­stract level in the brain’s pro­cess­ing of ideas and im­agery.

Concurrently I have been fol­low­ing dis­course on Reddit and other sources about how wide­spread use of AI has forced ed­u­ca­tors to re­turn to the long-aban­doned prac­tice of hav­ing their stu­dents take ex­ams in per­son by writ­ing things out long­hand in blue books. This has cre­ated new chal­lenges for stu­dents who never re­ally learned how to write by hand, and for teach­ers who can’t make sense of their stu­dents’ ter­ri­ble hand­writ­ing.

About twenty-five years ago I stopped com­pos­ing at the key­board and switched over to foun­tain pen on pa­per. Since then I have writ­ten many thou­sands of pages that way. The man­u­script of The Baroque Cycle was a stack of hand­writ­ten pages 42 inches high, which for a time was on dis­play at the Museum of Science Fiction in Seattle. With the ex­cep­tion of The Rise and Fall of D.O.D.O., which I co-wrote with Nicole Galland by email­ing Word files back and forth, every book I’ve writ­ten since then has been com­posed with foun­tain pen on pa­per.

Every so of­ten, when I’m sign­ing books at a book tour ap­pear­ance, some­one will come up to me and say some­thing like you must have writer’s cramp!” or is your hand sore yet?” I never have the time to pro­vide a full an­swer. If I did, how­ever, my an­swer would be that never, at any time dur­ing a quar­ter of a cen­tury dur­ing which I have spent a sub­stan­tial frac­tion of each work­ing day writ­ing by hand, have I ex­pe­ri­enced even the faintest traces of so-called writer’s cramp” or any other such hob­gob­lins.

Yet I can re­mem­ber get­ting a sore hand when I was a kid writ­ing out as­sign­ments in school. Many peo­ple prob­a­bly re­mem­ber such ex­pe­ri­ences and as­sume, rea­son­ably enough, that it’s a nat­ural con­se­quence of writ­ing by hand for any length of time. This is not the case.

Here are some fairly sim­ple dos and don’ts for peo­ple who want to reap the ben­e­fits of writ­ing by hand.

It’s pretty ob­vi­ous that you’re go­ing to get tired faster if your mus­cles have to ex­ert more force. Writing with a pen­cil re­quires sig­nif­i­cantly more force than writ­ing with a good pen. Old-school ball­points with thick ink are no bet­ter. You can see vi­sual ev­i­dence of this if you flip over a sheet of pa­per on which you’ve been writ­ing with a pen­cil or an old ball­point. The pa­per will bear a vis­i­ble im­print where it was pressed down by the writ­ing in­stru­ment. Often that will con­tinue down into the stack of pa­per be­neath. That’s be­cause you had to push hard. This does­n’t hap­pen with a foun­tain pen. If the nib is work­ing prop­erly you need to ex­ert very lit­tle force. The nib is ba­si­cally skat­ing on the lit­tle lake of ink that it has just laid down.

Pains me to say it, but roller­ball gel pens are about as good as foun­tain pens on this front.

It might then seem rea­son­able to think that writ­ing with a sty­lus on an iPad or sim­i­lar would be best, since no force is needed and fric­tion is min­i­mized. I don’t think this is true. A small amount of fric­tion is ac­tu­ally de­sir­able. You don’t want the tip of the writ­ing in­stru­ment to skid out of con­trol. Your brain and your lit­tle hand mus­cles are re­ly­ing on a lit­tle bit of fric­tion. Since I’m writ­ing this dur­ing the World Cup, I’ll make a soc­cer anal­ogy. Soccer play­ers have spent many hours drib­bling balls across play­ing fields, and they’ve in­ter­nal­ized the physics—they know about how far the ball is go­ing to travel when they kick it a cer­tain way, and how of­ten they need to give it an­other kick to keep it mov­ing. If you put them on a gi­ant, fric­tion­less air hockey table, all of that knowl­edge would be­come use­less. Every touch on the ball would send it out of con­trol. Dribbling the ball down the field would be­come more tir­ing be­cause they’d have to be mak­ing con­tin­ual ef­forts to con­trol the bal­l’s move­ment. Relying on a lit­tle bit of fric­tion re­duces the amount of men­tal and phys­i­cal ef­fort.

The com­bi­na­tion of foun­tain pens and pa­per em­bod­ies a bal­ance that has been worked out over a long span of time by peo­ple who write a lot. This phe­nom­e­non is called tooth” by afi­ciona­dos. Removing fric­tion by us­ing a hard sty­lus on glass will ac­tu­ally make the process more tir­ing.

Too much fric­tion, and too lit­tle fric­tion, are both more tir­ing than just a lit­tle bit of fric­tion, and that’s the bal­ance that is re­flected in the foun­tain pen/​pa­per tech­nol­ogy.

Rresults vary when you use var­i­ous pens on var­i­ous kinds of pa­per. Generally I get the worst re­sults on cheap printer pa­per, be­cause it wicks ink out of the nib too fast, and so cre­ates fat, blurry lines. Often I have the same prob­lem with yel­low le­gal pads. But al­most any pa­per in a blank note­book, or higher-grade printer pa­per with at least 25% cot­ton con­tent, works fine. I’ve learned over time that some of my foun­tain pens work bet­ter with cer­tain kinds of pa­per than oth­ers, so I match them up with­out hav­ing to think about it too hard.

Here’s a 300 dpi scan of tests I did with three dif­fer­ent pens on var­i­ous types of pa­per. You might have to zoom in to see much dif­fer­ence.

The pen on the left is a Jorg Hysek with a wide nib, and you can see that the cheap printer pa­per soaked up a lot of ink and left a thicker, fuzzier line. The le­gal pad was­n’t much bet­ter. Everything else ba­si­cally worked. The 100% cot­ton pa­per is from a box I pur­chased a long time ago - it was mar­keted for print­ing re­sumes, back in the days when peo­ple printed re­sumes. It is the tooth­iest of all these pa­pers and felt no­tice­ably scratch­ier. I guess it goes with­out say­ing that fancy Italian pa­per is the best, but the comp book and mole­sk­ine work per­fectly well with just about any pen.

(For those scor­ing at home, the mid­dle pen is a Diplomat Aero and the one on the right is a Monteverde Invincia)

If the pa­per is thin, writ­ing on one side can bleed through to the other, so the re­sults can be slightly harder to read if you write on both sides. Which leads me to:

The ecosys­tem is­n’t go­ing to col­lapse if you use more pa­per. It’s cheap. Focus on what’s im­por­tant here: your brain and your time. Write on one side. Trying to cram more words into a sheet will take you out of your nat­ural and com­fort­able writ­ing style and make you tired. Just buy a shit­load of pa­per or note­books or what­ever it is you want to use, and use it.

There’s a rea­son cur­sive was in­vented. Don’t even think about not us­ing it. It is far less tir­ing than print­ing one let­ter at a time. I learned cur­sive as a child. Then I went for many years with­out us­ing it much, and for­got some of it. Later I re-learned it by sit­ting in my kid’s el­e­men­tary school class­room dur­ing a par­ent-teacher con­fer­ence and ex­am­in­ing the forms printed on a long strip above the chalk­board (I still re­mem­bered how to do the lower-case let­ters, but I had for­got­ten some of the cap­i­tals).

Legibility was more im­por­tant back in the day when writ­ten doc­u­ments had to be read by other peo­ple. Hence the need for ex­act­ing pen­man­ship, taught in schools to long-suf­fer­ing chil­dren. This is prob­a­bly the source of a lot of angst around writer’s cramp and ink dis­as­ters. Today, if you’re writ­ing things down with ink on pa­per, you’re prob­a­bly writ­ing just for your­self, or per­haps for fam­ily mem­bers who can learn to rec­og­nize your hand­writ­ing.

To judge from the way peo­ple talk, a lot of them have mem­o­ries of foun­tain pen dis­as­ters where ink got all over the place for some rea­son. Or per­haps it’s just gen­er­a­tional trauma, handed down in an oral tra­di­tion. If the pen is work­ing cor­rectly, ink can only come out of it so fast. A cou­ple of rare ex­cep­tions:

If the pen’s ink reser­voir is partly empty, so that it con­tains an air bub­ble, and if it’s po­si­tioned nib down, then, when you go up in an air­plane, the bub­ble will ex­pand as the am­bi­ent pres­sure drops, forc­ing ink out the nib. Once I fig­ured that out, I got in the habit of mak­ing sure my pens were po­si­tioned nib up when tak­ing off in an air­plane. If I have time I’ll also re­fill the pen be­fore de­par­ture, to min­i­mize the size of the air bub­ble.

Sometimes if a pen gets dirty, or if the nib is some­how dam­aged, the ink will stop com­ing out and you can restart it by giv­ing it a lit­tle shake. If you do it just right, the ink flow restarts with­out in­ci­dent, but if you overdo it, a few drops of ink might shoot out onto the page and be­come blots. This sce­nario hap­pens a few times of year for me, only with one pen that has this prob­lem. I blot it with a piece of scrap pa­per and move on.

Just have note­books ly­ing around, or on your per­son. Write gro­cery lists, doo­dles, notes on meet­ings, to-do lists, or stray ideas. Journal. Copy out good lines from books. Anything that has your men­tal fo­cus will have a more en­dur­ing pres­ence in your brain if you write it down.

I am left handed. I have never had any trou­ble with my hand smear­ing the ink. Yet every con­ver­sa­tion I have about foun­tain pens leads to some­one claim­ing that it can never work for them be­cause they are left handed. I have no idea what they’re talk­ing about. When I was a child, writ­ing at length with pen­cil, the side of my hand some­times be­came gray from graphite picked up as my hand rubbed across the page. And some­times I have got ink on my hand when us­ing a ball­point pen that left an ink glob on the pa­per. But with foun­tain pens it’s easy to find a pen/​pa­per com­bi­na­tion such that the ink soaks into the pa­per and dries quickly enough that it does­n’t smudge when you’re writ­ing the next line. Here’s a sim­ple demon­stra­tion of dry­ing time and how it works with two pens: first a foun­tain pen and then a Pilot G-2 gel pen.

Obviously the Pilot gel pen ink dries faster, and so that might be a bet­ter choice for peo­ple who are re­ally wor­ried about smudg­ing.

Most mod­ern pens al­low you to choose be­tween us­ing pre­loaded plas­tic ink car­tridges and a plunger that en­ables you to draw ink up out of a bot­tle by hand. I use both. Start with the ink car­tridges, es­pe­cially if you travel. There’s no need to com­pli­cate mat­ters by mess­ing around with bot­tles. Since I do a lot of work from one lo­ca­tion, I have a cor­ner of a table­top set up there with ink bot­tles and a folded-up pa­per towel for wip­ing off the nib af­ter it’s filled (I have been us­ing the same pa­per towel for about twenty years). In the­ory this works bet­ter in the long term be­cause it al­lows you to flush the nib by forc­ing ink in and out of it a cou­ple of times when­ever you re­fill. In prac­tice I see no dif­fer­ence at all - pens that I re­fill with car­tridges don’t get clogged.

Even if every­thing works per­fectly you’ll end up with the oc­ca­sional ink-smudged fin­ger. It will wash off quickly - the ink is wa­ter-sol­u­ble. Until then, con­sider it a mark of dis­tinc­tion.

If you’re new to this I think it makes most sense to start by con­sid­er­ing what kind of pa­per is go­ing to work best in your life. Are you writ­ing loose­leaf, or in note­books? Legal pads? Blue books? Remember, it’s okay to use lots of pa­per, so pick some­thing that is­n’t too pre­cious and that is easy to re­plen­ish. I use a lot of mole­sk­ine note­books and Mead comp books, which I can buy in bulk on­line. For com­pos­ing fic­tion I use fancy loose­leaf pa­per.

If you have ac­cess to a store where they sell foun­tain pens, take some of that pa­per there and see what works best. If you’re work­ing with cheaper, thin­ner pa­per, start with finer nibs and work up to fat­ter ones un­til you start to see bleed-through.

Buy cheaper pens un­til you know what you like. I doubt there’s much of a dif­fer­ence be­tween cheaper and more ex­pen­sive foun­tain pens in terms of their ac­tual per­for­mance. What you’re pay­ing for, in an ex­pen­sive pen, is fancy ma­te­ri­als and styling. For ex­am­ple, if you look at the Pilot Vanishing Point line of pens - an in­ge­nious foun­tain pen that you can click, like an old-fash­ioned ball­point, to re­tract the nib in­side the bar­rel - fancier ver­sions cost five times as much as the base model.

In all hon­esty, the Pilot G-2 gel pens are go­ing to give you 80% of what you could ex­pect from a foun­tain pen for min­i­mal cost.

On the other hand, a ten-pack of Pilot G-2 gel pens goes for about twenty bucks. For the same amount you can buy a sim­ple but com­pletely ser­vice­able foun­tain pen that will last longer than you will.

No posts

moonshotai/Kimi-K3 · Hugging Face

huggingface.co

📰  Tech Blog |     📄  Full Report

1. Model Introduction

Kimi K3 is an open-weight, na­tive mul­ti­modal agen­tic model and our most ca­pa­ble model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with na­tive vi­sion ca­pa­bil­i­ties and a 1-million-token con­text win­dow. It is the world’s first open 3T-class model, de­signed for fron­tier in­tel­li­gence across long-hori­zon cod­ing, knowl­edge work, and rea­son­ing.

Key Features

New Architecture: Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE spar­sity with a Stable LatentMoE frame­work that ac­ti­vates 16 out of 896 ex­perts — yield­ing an ap­prox­i­mate 2.5× im­prove­ment in over­all scal­ing ef­fi­ciency over Kimi K2.

Long-Horizon Coding: Operating with min­i­mal hu­man over­sight, Kimi K3 sus­tains long en­gi­neer­ing ses­sions, nav­i­gates mas­sive repos­i­to­ries, and or­ches­trates ter­mi­nal tools — from GPU ker­nel op­ti­miza­tion and com­piler de­vel­op­ment to vi­sion-in-the-loop game dev, CAD, and even chip de­sign.

Agentic Knowledge Work: Kimi K3 ad­vances end-to-end knowl­edge work, pro­duc­ing deep re­search with in­ter­ac­tive vi­su­al­iza­tions, wid­gets and dash­boards, and mo­tion de­sign and video edit­ing, pow­ered by its na­tive mul­ti­modal ar­chi­tec­ture.

Native Multimodality & Long Context: Kimi K3 un­der­stands text, im­ages, and video within the same model, and sup­ports a 1-million-token con­text win­dow.

Open Frontier Weights: We re­lease the full Kimi K3 model weights un­der the Kimi K3 License, mak­ing fron­tier in­tel­li­gence openly avail­able for re­search, de­ploy­ment, and fur­ther in­no­va­tion.

2. Model Summary

3. Evaluation Results

All Kimi K3 re­sults are ob­tained with rea­son­ing ef­fort set to max’ and tem­per­a­ture = 1.0. For sin­gle-step tasks, such as GPQA Diamond, HLE-Full, and vi­sion bench­marks with­out tools, we set top-p = 0.95; for agen­tic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, and ZeroBench, each cell re­ports the scores with­out and with tool aug­men­ta­tion (general tools for HLE-Full, Python for the vi­sion bench­marks), in that or­der.

Reasoning & knowl­edge bench­marks CritPt and AA-LCR. Scores are cited from Artificial Analysis as of July 23, 2026.

CritPt and AA-LCR. Scores are cited from Artificial Analysis as of July 23, 2026.

Coding bench­marks DeepSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is taken from the GLM-5.2 re­lease blog; all re­main­ing scores are from the of­fi­cial DeepSWE leader­board, un­der which Kimi K3 at­tains 67.3 with the mini-SWE-agent har­ness. We re­port the DeepSWE v1.1 tasks. Terminal-Bench 2.1. Kimi K3 is eval­u­ated with the Kimi Code har­ness. For all other mod­els, we re­port the best score across har­nesses: GLM-5.2 with Claude Code (GLM-5.2 re­lease blog); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 (Artificial Analysis); GPT-5.5 and GPT-5.6 Sol with Codex (OpenAI). ProgramBench. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog; all other scores are from Vals AI. SWE-Marathon. Kimi K3, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.6 Sol is eval­u­ated with the Codex har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog. Our eval­u­a­tion is based on an H20-calibrated branch of the of­fi­cial tasks as of July 9, 2026, prior to the fi­nal v1.1 re­lease: the Docker im­ages, per­for­mance gates, and ref­er­ence or­a­cles for the GPU tasks have been re­cal­i­brated for H20, while the cor­rect­ness and anti-cheat val­ida­tors re­main un­changed. Additionally, Claude Fable 5 hit fall­backs on 35% of the tasks in our eval­u­a­tion, which may have neg­a­tively im­pacted its mea­sured per­for­mance. FrontierSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness and GPT-5.6 Sol with the Codex har­ness; all other re­sults are from FrontierSWE. Dominance scores are re­com­puted from the raw scores us­ing the of­fi­cial eval­u­a­tion script and are cur­rent as of July 16, 2026. PostTrainBench. Scores for GLM-5.2, GPT-5.5, and Claude Opus 4.8 are adopted from the of­fi­cial PostTrainBench re­sults. Kimi K3, Claude Fable 5, and GPT-5.6 Sol are eval­u­ated with the of­fi­cial Harbor im­ple­men­ta­tion at max­i­mum rea­son­ing ef­fort, av­er­aged over three runs on H20 GPUs (instead of H100 in the of­fi­cial set­ting) — Kimi K3 and Claude Fable 5 with the Claude Code har­ness, and GPT-5.6 Sol with the Codex har­ness. MLS-Bench-Lite. Kimi K3 is eval­u­ated with the Kimi Code har­ness; GLM-5.2 and the Claude mod­els with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness. SciCode. Scores are cited from Artificial Analysis as of July 23, 2026. Kimi Code Bench 2.0 (in-house). Kimi K3 is eval­u­ated with the Kimi Code har­ness (it at­tains 73.7 with the Claude Code har­ness); GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness. All mod­els are eval­u­ated at max­i­mum rea­son­ing ef­fort, ex­cept GPT-5.5, which uses the xhigh” set­ting. As the bench­mark in­cludes cy­ber­se­cu­rity and safety-re­lated tasks, we also dis­close the frac­tion of re­fused or fall­back tasks: Claude Fable 5 hit 13 fall­backs and 1 re­fusal out of 80 tasks; 10 re­fusals out of 80 tasks en­tered GPT-5.6 Sol’s cy­ber guard; GPT-5.5 had 3 re­fusals out of 80 tasks.

DeepSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is taken from the GLM-5.2 re­lease blog; all re­main­ing scores are from the of­fi­cial DeepSWE leader­board, un­der which Kimi K3 at­tains 67.3 with the mini-SWE-agent har­ness. We re­port the DeepSWE v1.1 tasks.

Terminal-Bench 2.1. Kimi K3 is eval­u­ated with the Kimi Code har­ness. For all other mod­els, we re­port the best score across har­nesses: GLM-5.2 with Claude Code (GLM-5.2 re­lease blog); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 (Artificial Analysis); GPT-5.5 and GPT-5.6 Sol with Codex (OpenAI).

ProgramBench. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog; all other scores are from Vals AI.

SWE-Marathon. Kimi K3, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.6 Sol is eval­u­ated with the Codex har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog. Our eval­u­a­tion is based on an H20-calibrated branch of the of­fi­cial tasks as of July 9, 2026, prior to the fi­nal v1.1 re­lease: the Docker im­ages, per­for­mance gates, and ref­er­ence or­a­cles for the GPU tasks have been re­cal­i­brated for H20, while the cor­rect­ness and anti-cheat val­ida­tors re­main un­changed. Additionally, Claude Fable 5 hit fall­backs on 35% of the tasks in our eval­u­a­tion, which may have neg­a­tively im­pacted its mea­sured per­for­mance.

FrontierSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness and GPT-5.6 Sol with the Codex har­ness; all other re­sults are from FrontierSWE. Dominance scores are re­com­puted from the raw scores us­ing the of­fi­cial eval­u­a­tion script and are cur­rent as of July 16, 2026.

PostTrainBench. Scores for GLM-5.2, GPT-5.5, and Claude Opus 4.8 are adopted from the of­fi­cial PostTrainBench re­sults. Kimi K3, Claude Fable 5, and GPT-5.6 Sol are eval­u­ated with the of­fi­cial Harbor im­ple­men­ta­tion at max­i­mum rea­son­ing ef­fort, av­er­aged over three runs on H20 GPUs (instead of H100 in the of­fi­cial set­ting) — Kimi K3 and Claude Fable 5 with the Claude Code har­ness, and GPT-5.6 Sol with the Codex har­ness.

MLS-Bench-Lite. Kimi K3 is eval­u­ated with the Kimi Code har­ness; GLM-5.2 and the Claude mod­els with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness.

SciCode. Scores are cited from Artificial Analysis as of July 23, 2026.

Kimi Code Bench 2.0 (in-house). Kimi K3 is eval­u­ated with the Kimi Code har­ness (it at­tains 73.7 with the Claude Code har­ness); GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness. All mod­els are eval­u­ated at max­i­mum rea­son­ing ef­fort, ex­cept GPT-5.5, which uses the xhigh” set­ting. As the bench­mark in­cludes cy­ber­se­cu­rity and safety-re­lated tasks, we also dis­close the frac­tion of re­fused or fall­back tasks: Claude Fable 5 hit 13 fall­backs and 1 re­fusal out of 80 tasks; 10 re­fusals out of 80 tasks en­tered GPT-5.6 Sol’s cy­ber guard; GPT-5.5 had 3 re­fusals out of 80 tasks.

Agentic bench­marks OfficeQA Pro. Each test case pro­vides the agent with the en­tire PDF cor­pus, with all PDFs ren­dered as im­ages and no ma­chine-read­able text avail­able. OfficeQA Pro and SpreadsheetBench 2. Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol are eval­u­ated with the Codex har­ness. MCP-Atlas. All mod­els are eval­u­ated on the 500-task pub­lic sub­set with a 100-turn limit, us­ing Gemini 3.1 Pro as the judge. AutomationBench. All mod­els are eval­u­ated on the 600-task pub­lic sub­set, fol­low­ing the of­fi­cial GitHub setup in all other re­spects. BrowseComp. We adopt a con­text-com­paction strat­egy trig­gered at 300K to­kens. When eval­u­ated with the full 1M-token con­text win­dow and no con­text man­age­ment, Kimi K3 achieves a score of 90.4. The re­sults of Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 are cited from Anthropic and OpenAI. GDPval-AA v2, AA-Briefcase, τ³-Bank­ing, Harvey Lab-AA, and APEX-Agents. Scores are cited from Artificial Analysis and the APEX-Agents leader­board as of July 23, 2026. For Harvey Lab-AA, we re­port the cri­te­rion pass rate. CorpFin v2, Finance Agent v2, and Legal Research Bench. Scores are cited from Vals AI. Agents’ Last Exam. Scores are cited from the of­fi­cial leader­board as of July 23, 2026; we re­port the leader­board’s pri­mary pass-rate met­ric. On the leader­board, each model is paired with a spe­cific har­ness: Kimi K3 with Kimi Code; GPT-5.6 Sol and GPT-5.5 with Codex; Claude Fable 5, Claude Opus 4.8, and GLM-5.2 with Claude Code. † The Claude Fable 5 en­try runs at xhigh ef­fort with 40% of tasks an­no­tated as down­graded.

OfficeQA Pro. Each test case pro­vides the agent with the en­tire PDF cor­pus, with all PDFs ren­dered as im­ages and no ma­chine-read­able text avail­able.

OfficeQA Pro and SpreadsheetBench 2. Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol are eval­u­ated with the Codex har­ness.

MCP-Atlas. All mod­els are eval­u­ated on the 500-task pub­lic sub­set with a 100-turn limit, us­ing Gemini 3.1 Pro as the judge.

AutomationBench. All mod­els are eval­u­ated on the 600-task pub­lic sub­set, fol­low­ing the of­fi­cial GitHub setup in all other re­spects.

BrowseComp. We adopt a con­text-com­paction strat­egy trig­gered at 300K to­kens. When eval­u­ated with the full 1M-token con­text win­dow and no con­text man­age­ment, Kimi K3 achieves a score of 90.4. The re­sults of Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 are cited from Anthropic and OpenAI.

GDPval-AA v2, AA-Briefcase, τ³-Bank­ing, Harvey Lab-AA, and APEX-Agents. Scores are cited from Artificial Analysis and the APEX-Agents leader­board as of July 23, 2026. For Harvey Lab-AA, we re­port the cri­te­rion pass rate.

CorpFin v2, Finance Agent v2, and Legal Research Bench. Scores are cited from Vals AI.

Agents’ Last Exam. Scores are cited from the of­fi­cial leader­board as of July 23, 2026; we re­port the leader­board’s pri­mary pass-rate met­ric. On the leader­board, each model is paired with a spe­cific har­ness: Kimi K3 with Kimi Code; GPT-5.6 Sol and GPT-5.5 with Codex; Claude Fable 5, Claude Opus 4.8, and GLM-5.2 with Claude Code. † The Claude Fable 5 en­try runs at xhigh ef­fort with 40% of tasks an­no­tated as down­graded.

Multimodal bench­marks Except for ZeroBench, which fol­lows the of­fi­cial set­ting and is run five times, all mul­ti­modal scores are av­er­aged over three runs. MMMU-Pro is eval­u­ated fol­low­ing the of­fi­cial pro­to­col, pre­serv­ing the orig­i­nal in­put or­der and prepend­ing im­ages to the text in­put. PerceptionBench is an in-house bench­mark that fo­cuses on atomic vi­sual per­cep­tion ca­pa­bil­i­ties.

Except for ZeroBench, which fol­lows the of­fi­cial set­ting and is run five times, all mul­ti­modal scores are av­er­aged over three runs. MMMU-Pro is eval­u­ated fol­low­ing the of­fi­cial pro­to­col, pre­serv­ing the orig­i­nal in­put or­der and prepend­ing im­ages to the text in­put.

PerceptionBench is an in-house bench­mark that fo­cuses on atomic vi­sual per­cep­tion ca­pa­bil­i­ties.

4. Native MXFP4 Quantization

Kimi K3 ap­plies quan­ti­za­tion-aware train­ing from the SFT stage on­ward, us­ing MXFP4 weights with MXFP8 ac­ti­va­tions for broad hard­ware com­pat­i­bil­ity.

5. Deployment

You can ac­cess Kimi K3′s API on https://​plat­form.kimi.ai by se­lect­ing kimi-k3, and we pro­vide OpenAI/Anthropic-compatible API for you. Currently, Kimi K3 is rec­om­mended to run on the fol­low­ing in­fer­ence en­gines:

You can ac­cess Kimi K3′s API on https://​plat­form.kimi.ai by se­lect­ing kimi-k3, and we pro­vide OpenAI/Anthropic-compatible API for you. Currently, Kimi K3 is rec­om­mended to run on the fol­low­ing in­fer­ence en­gines:

vLLM — see recipes

SGLang — see cook­book

TokenSpeed — see recipes

6. Model Usage

Kimi K3 al­ways has think­ing en­abled, and will re­turn rea­son­ing_­con­tent. Thinking ef­fort is con­fig­ured with the top-level rea­son­ing_­ef­fort re­quest field, which sup­ports low”, high”, and max” (default max”).

Kimi K3 was trained in the pre­served think­ing his­tory mode. For multi-turn con­ver­sa­tions and tool calls, Kimi K3 re­quires the com­plete as­sis­tant mes­sage re­turned by the API to be passed back to mes­sages as-is — in­clud­ing rea­son­ing_­con­tent and tool_­calls, not just con­tent:

im­port ope­nai

def chat_with­_p­re­served_­think­ing(client: ope­nai.Ope­nAI, mod­el_­name: str): mes­sages = [ { role”: user”, content”: Tell me three ran­dom num­bers.” }, { role”: assistant”, reasoning_content”: I’ll start by list­ing five num­bers: 473, 921, 235, 215, 222, and I’ll tell you the first three.”, content”: 473, 921, 235″ }, { role”: user”, content”: What are the other two num­bers you have in mind?” } ]

re­sponse = client.chat.com­ple­tions.cre­ate( model=mod­el_­name, mes­sages=mes­sages, stream=False, max_­to­kens=4096, rea­son­ing_­ef­fort=“max”, ) # the as­sis­tant should men­tion 215 and 222 that ap­pear in the prior rea­son­ing con­tent print(f”re­sponse: {response.choices[0].message.reasoning}“) re­turn re­sponse.choices[0].mes­sage.con­tent

For full guides and ex­am­ples (vision in­put, struc­tured out­put, par­tial mode, tool choice, dy­namic tool load­ing, con­text caching), see the Kimi K3 Quickstart and Thinking Effort.

Coding Agent Framework

Kimi K3 works best with Kimi Code CLI as its agent frame­work. We warmly in­vite you to give it a try — run Kimi Code in your ter­mi­nal and se­lect Kimi K3 us­ing the /model com­mand. We hope you en­joy build­ing with Kimi K3, and we would love to hear your feed­back!

7. License

Both the code repos­i­tory and the model weights are re­leased un­der the Kimi K3 License.

8. Contact Us

If you have any ques­tions, please reach out at sup­port@moon­shot.ai.

Safetensors

Model tree for moon­shotai/​Kimi-K3

Spaces us­ing moon­shotai/​Kimi-K3 5

Collection in­clud­ing moon­shotai/​Kimi-K3

US prosecutors charge Atlanta man after GrapheneOS phone wipes itself during airport search

www.techspot.com

Serving tech en­thu­si­asts for over 25 years. TechSpot means tech analy­sis and ad­vice you can trust.

A hot potato: A fed­eral case in Atlanta is rais­ing ques­tions about a pri­vacy-fo­cused mo­bile op­er­at­ing sys­tem, with pros­e­cu­tors ar­gu­ing that its fea­tures were used to erase ev­i­dence. The US Department of Justice is at­tempt­ing to pros­e­cute Atlanta res­i­dent Sam Tunick un­der a fed­eral statute that makes it a crime to de­stroy prop­erty in an ef­fort to pre­vent it from be­ing seized.

The case cen­ters on Tunick’s use of GrapheneOS, an open-source op­er­at­ing sys­tem that works on Google Pixel phones and lets users en­ter a pass­code to wipe a de­vice clean.

Experts said the le­gal ap­proach is un­usual and may be the first time the law has been aimed at an op­er­at­ing sys­tem. It’s con­cern­ing — and sends the mes­sage that [GrapheneOS] is crim­i­nal by de­fault,” said Christophe Boutry, a cy­ber­se­cu­rity and sur­veil­lance ex­pert. Boutry and Bill Buddington, se­nior staff tech­nol­o­gist at the Electronic Frontier Foundation, both said they had not seen a sim­i­lar case.

The in­ci­dent be­gan at Hartsfield-Jackson Atlanta International Airport on January 24 of last year. Tunick had just re­turned from a trip to the Dominican Republic when he was stopped for ques­tion­ing. According to court tes­ti­mony, fed­eral agents had al­ready cir­cu­lated his name and photo in­ter­nally, say­ing he was un­der in­ves­ti­ga­tion for suspected ter­ror­ism ac­tiv­i­ties” be­cause of his al­leged as­so­ci­a­tion with the move­ment against Cop City.

Tunick was taken to a sec­ondary screen­ing room, where mul­ti­ple agents ques­tioned him. A mo­tion filed by his de­fense ar­gues the in­ter­ro­ga­tion fo­cused on child sex­ual abuse ma­te­r­ial as a pre­text for in­ves­ti­gat­ing his con­nec­tions to the protest move­ment. The mo­tion also states that Tunick asked four times to speak with a lawyer and was de­nied each time. According to the same fil­ing, agents did not pre­sent a war­rant or read him his rights.

Government at­tor­neys and agents pushed back dur­ing Monday’s hear­ing. They de­scribed the en­counter as a rou­tine air­port in­spec­tion. Larry Findley, a Customs and Border Protection of­fi­cer, said agents were looking for any­thing that’s pro­hib­ited.”

During the ques­tion­ing, agents re­peat­edly asked Tunick to un­lock his phone and warned they would seize it if he re­fused. When he fi­nally pro­vided a pass­code, the phone ap­peared to restart. The de­fense mo­tion states that the screen went blank, flashed sev­eral times, and the phone ap­peared to restart,” re­sult­ing in the loss of data.

The wipe is now cen­tral to the case. Prosecutors are treat­ing it as an in­ten­tional act to de­stroy ev­i­dence, while the de­fense ar­gues that the search vi­o­lated Tunick’s con­sti­tu­tional rights and that the ev­i­dence should be sup­pressed.

The case raises ques­tions about which con­sti­tu­tional rights ap­ply at US bor­ders, in­clud­ing in­ter­na­tional air­ports, where au­thor­i­ties have broader search pow­ers. A judge is not ex­pected to rule on the de­fense mo­tion un­til at least late October.

GrapheneOS is de­signed to im­prove pri­vacy and se­cu­rity on Pixel phones. Supporters say those tools are le­git­i­mate se­cu­rity pro­tec­tions, not ev­i­dence of crim­i­nal in­tent. Boutry pointed to France and Spain, where au­thor­i­ties have strug­gled to gain ac­cess to se­cured de­vices. He said au­thor­i­ties have treated the use of GrapheneOS it­self as sus­pi­cious. In Catalonia, Spain, po­lice have been pro­fil­ing peo­ple car­ry­ing Pixel phones, as­sum­ing they have GrapheneOS in­stalled and are drug deal­ers or gang mem­bers.

The main goal [of the op­er­at­ing sys­tem] is pro­tec­tion of pri­vacy,” Boutry said. They’re our phones and the state can’t tell us how to use them.”

The case is tied to on­go­ing op­po­si­tion to Cop City, a $109 mil­lion po­lice train­ing fa­cil­ity that opened last spring. The pro­ject has drawn op­po­si­tion from ac­tivists con­cerned about po­lice mil­i­ta­riza­tion and en­vi­ron­men­tal im­pacts. Law en­force­ment of­fi­cials have de­fended it as nec­es­sary for train­ing and re­cruit­ment.

Previous at­tempts to pros­e­cute pro­test­ers at the state level have foundered, while fed­eral au­thor­i­ties have more re­cently stepped in, in­clud­ing a sep­a­rate in­dict­ment an­nounced last month.

Kill the Cookie Banner!

killthecookiebanner.eu

Stop the track­ing cir­cus.

Tired of mis­lead­ing cookie ban­ners? The EU Commission has fi­nally pro­posed a so­lu­tion: set your pri­vacy pref­er­ences in the browser once, and never see an­other ban­ner. Unfortunately, the track­ing in­dus­try is push­ing back — and so far, they’ve been suc­cess­ful. We need YOUR help to #KillTheCookieBanner!

Cookie ban­ners are made to trick you into waiv­ing your rights

You may think that EU pri­vacy law re­quires cookie ban­ners. But the law is clear: on­line track­ing is pro­hib­ited by de­fault.

Therefore, the track­ing in­dus­try needs you to waive your rights. That’s why they in­vented cookie ban­ners, which of­ten are de­lib­er­ately mis­lead­ing as well as an­noy­ing.

This re­sults in up to 90% of peo­ple say­ing YES — even though only around 3% ac­tu­ally want to be tracked on­line. This sys­tem is bro­ken by de­sign.

The so­lu­tion: au­to­mat­i­cally com­mu­ni­cate your pri­vacy pref­er­ence

In Autumn 2025, as part of a big­ger le­gal re­form*, the EU Commission fi­nally pro­posed a so­lu­tion to the cookie ban­ner prob­lem: au­to­mated sig­nals that would com­mu­ni­cate your pri­vacy pref­er­ences be­tween your de­vice and web­sites or apps. You could then choose whether you want to ac­cept, refuse, or limit track­ing.

This idea is nei­ther new nor com­plex. Your browser al­ready au­to­mat­i­cally sig­nals other pref­er­ences to web­sites, for ex­am­ple your pre­ferred lan­guage. In some US states, such sig­nals for pri­vacy pref­er­ences are al­ready legally sup­ported.

The track­ing lobby fights to keep the cookie ban­ner

This would be a sim­ple so­lu­tion. But the track­ing in­dus­try seems afraid that if you can ex­press your pref­er­ences ef­fi­ciently, it could re­sult in lower con­sent rates for track­ing.

Following lob­by­ing ef­forts spear­headed by Google and the track­ing in­dus­try, sev­eral Member States are now block­ing the EU Commission’s pro­posal to get rid of the cookie ban­ner.

But not only that: in­dus­try groups are also lob­by­ing the European Parliament to re­ject the pro­posal.

We need YOUR help!

That’s where you come in: the fight to kill the cookie ban­ner is far from over. The Member States and the European Parliament have not yet de­cided on their po­si­tion on the is­sue yet.

You can take ac­tion by con­tact­ing your rep­re­sen­ta­tive in the European Parliament or in your Member State and ex­press your frus­tra­tion.

*This pro­posal for pri­vacy sig­nals is part of an EU law re­form called the Digital Omnibus. Most other parts of this re­form are prob­lem­atic and would weaken peo­ple’s rights. We want to make clear that we do not sup­port these other as­pects of the pro­posed re­form.

*This pro­posal for pri­vacy sig­nals is part of an EU law re­form called the Digital Omnibus. Most other parts of this re­form are prob­lem­atic and would weaken peo­ple’s rights. We want to make clear that we do not sup­port these other as­pects of the pro­posed re­form.

Check out this chat

chatgpt.com

Get re­sponses tai­lored to you

Log in to get an­swers based on saved chats, plus cre­ate im­ages and up­load files.

Just a moment...

ads.openai.com

Just a moment...

www.politico.com

Bento Slides

bento.page

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.