10 interesting stories served every morning and every evening.

openai.com

Writing by Hand is Good for your Brain - Here's how to do it

nealstephenson.substack.com

Because I am known to write us­ing a foun­tain pen on pa­per, a num­ber of peo­ple have pointed me to this post and its un­der­ly­ing re­search. I won’t re­hash what is said in those sources, but the gist of it is that when you write things down by hand you’re re­cruit­ing more of your brain, which is a good thing.

I’m not an ex­pert on how the brain works, but I can say that, when writ­ing by hand, one is con­tin­u­ally solv­ing a se­ries of small prob­lems hav­ing to do with the spac­ing of words, how let­ters are con­nected, the cross­ing of the let­ter t (sometimes more than one in the same word) and the dot­ting of the let­ters i and j, and how to ac­com­plish all of those things through co­or­di­nated move­ments not just of the fin­gers but of the whole arm. All of that has to be in­te­grated in real time with what­ever is hap­pen­ing on a more ab­stract level in the brain’s pro­cess­ing of ideas and im­agery.

Concurrently I have been fol­low­ing dis­course on Reddit and other sources about how wide­spread use of AI has forced ed­u­ca­tors to re­turn to the long-aban­doned prac­tice of hav­ing their stu­dents take ex­ams in per­son by writ­ing things out long­hand in blue books. This has cre­ated new chal­lenges for stu­dents who never re­ally learned how to write by hand, and for teach­ers who can’t make sense of their stu­dents’ ter­ri­ble hand­writ­ing.

About twenty-five years ago I stopped com­pos­ing at the key­board and switched over to foun­tain pen on pa­per. Since then I have writ­ten many thou­sands of pages that way. The man­u­script of The Baroque Cycle was a stack of hand­writ­ten pages 42 inches high, which for a time was on dis­play at the Museum of Science Fiction in Seattle. With the ex­cep­tion of The Rise and Fall of D.O.D.O., which I co-wrote with Nicole Galland by email­ing Word files back and forth, every book I’ve writ­ten since then has been com­posed with foun­tain pen on pa­per.

Every so of­ten, when I’m sign­ing books at a book tour ap­pear­ance, some­one will come up to me and say some­thing like you must have writer’s cramp!” or is your hand sore yet?” I never have the time to pro­vide a full an­swer. If I did, how­ever, my an­swer would be that never, at any time dur­ing a quar­ter of a cen­tury dur­ing which I have spent a sub­stan­tial frac­tion of each work­ing day writ­ing by hand, have I ex­pe­ri­enced even the faintest traces of so-called writer’s cramp” or any other such hob­gob­lins.

Yet I can re­mem­ber get­ting a sore hand when I was a kid writ­ing out as­sign­ments in school. Many peo­ple prob­a­bly re­mem­ber such ex­pe­ri­ences and as­sume, rea­son­ably enough, that it’s a nat­ural con­se­quence of writ­ing by hand for any length of time. This is not the case.

Here are some fairly sim­ple dos and don’ts for peo­ple who want to reap the ben­e­fits of writ­ing by hand.

It’s pretty ob­vi­ous that you’re go­ing to get tired faster if your mus­cles have to ex­ert more force. Writing with a pen­cil re­quires sig­nif­i­cantly more force than writ­ing with a good pen. Old-school ball­points with thick ink are no bet­ter. You can see vi­sual ev­i­dence of this if you flip over a sheet of pa­per on which you’ve been writ­ing with a pen­cil or an old ball­point. The pa­per will bear a vis­i­ble im­print where it was pressed down by the writ­ing in­stru­ment. Often that will con­tinue down into the stack of pa­per be­neath. That’s be­cause you had to push hard. This does­n’t hap­pen with a foun­tain pen. If the nib is work­ing prop­erly you need to ex­ert very lit­tle force. The nib is ba­si­cally skat­ing on the lit­tle lake of ink that it has just laid down.

Pains me to say it, but roller­ball gel pens are about as good as foun­tain pens on this front.

It might then seem rea­son­able to think that writ­ing with a sty­lus on an iPad or sim­i­lar would be best, since no force is needed and fric­tion is min­i­mized. I don’t think this is true. A small amount of fric­tion is ac­tu­ally de­sir­able. You don’t want the tip of the writ­ing in­stru­ment to skid out of con­trol. Your brain and your lit­tle hand mus­cles are re­ly­ing on a lit­tle bit of fric­tion. Since I’m writ­ing this dur­ing the World Cup, I’ll make a soc­cer anal­ogy. Soccer play­ers have spent many hours drib­bling balls across play­ing fields, and they’ve in­ter­nal­ized the physics—they know about how far the ball is go­ing to travel when they kick it a cer­tain way, and how of­ten they need to give it an­other kick to keep it mov­ing. If you put them on a gi­ant, fric­tion­less air hockey table, all of that knowl­edge would be­come use­less. Every touch on the ball would send it out of con­trol. Dribbling the ball down the field would be­come more tir­ing be­cause they’d have to be mak­ing con­tin­ual ef­forts to con­trol the bal­l’s move­ment. Relying on a lit­tle bit of fric­tion re­duces the amount of men­tal and phys­i­cal ef­fort.

The com­bi­na­tion of foun­tain pens and pa­per em­bod­ies a bal­ance that has been worked out over a long span of time by peo­ple who write a lot. This phe­nom­e­non is called tooth” by afi­ciona­dos. Removing fric­tion by us­ing a hard sty­lus on glass will ac­tu­ally make the process more tir­ing.

Too much fric­tion, and too lit­tle fric­tion, are both more tir­ing than just a lit­tle bit of fric­tion, and that’s the bal­ance that is re­flected in the foun­tain pen/​pa­per tech­nol­ogy.

Rresults vary when you use var­i­ous pens on var­i­ous kinds of pa­per. Generally I get the worst re­sults on cheap printer pa­per, be­cause it wicks ink out of the nib too fast, and so cre­ates fat, blurry lines. Often I have the same prob­lem with yel­low le­gal pads. But al­most any pa­per in a blank note­book, or higher-grade printer pa­per with at least 25% cot­ton con­tent, works fine. I’ve learned over time that some of my foun­tain pens work bet­ter with cer­tain kinds of pa­per than oth­ers, so I match them up with­out hav­ing to think about it too hard.

Here’s a 300 dpi scan of tests I did with three dif­fer­ent pens on var­i­ous types of pa­per. You might have to zoom in to see much dif­fer­ence.

The pen on the left is a Jorg Hysek with a wide nib, and you can see that the cheap printer pa­per soaked up a lot of ink and left a thicker, fuzzier line. The le­gal pad was­n’t much bet­ter. Everything else ba­si­cally worked. The 100% cot­ton pa­per is from a box I pur­chased a long time ago - it was mar­keted for print­ing re­sumes, back in the days when peo­ple printed re­sumes. It is the tooth­iest of all these pa­pers and felt no­tice­ably scratch­ier. I guess it goes with­out say­ing that fancy Italian pa­per is the best, but the comp book and mole­sk­ine work per­fectly well with just about any pen.

(For those scor­ing at home, the mid­dle pen is a Diplomat Aero and the one on the right is a Monteverde Invincia)

If the pa­per is thin, writ­ing on one side can bleed through to the other, so the re­sults can be slightly harder to read if you write on both sides. Which leads me to:

The ecosys­tem is­n’t go­ing to col­lapse if you use more pa­per. It’s cheap. Focus on what’s im­por­tant here: your brain and your time. Write on one side. Trying to cram more words into a sheet will take you out of your nat­ural and com­fort­able writ­ing style and make you tired. Just buy a shit­load of pa­per or note­books or what­ever it is you want to use, and use it.

There’s a rea­son cur­sive was in­vented. Don’t even think about not us­ing it. It is far less tir­ing than print­ing one let­ter at a time. I learned cur­sive as a child. Then I went for many years with­out us­ing it much, and for­got some of it. Later I re-learned it by sit­ting in my kid’s el­e­men­tary school class­room dur­ing a par­ent-teacher con­fer­ence and ex­am­in­ing the forms printed on a long strip above the chalk­board (I still re­mem­bered how to do the lower-case let­ters, but I had for­got­ten some of the cap­i­tals).

Legibility was more im­por­tant back in the day when writ­ten doc­u­ments had to be read by other peo­ple. Hence the need for ex­act­ing pen­man­ship, taught in schools to long-suf­fer­ing chil­dren. This is prob­a­bly the source of a lot of angst around writer’s cramp and ink dis­as­ters. Today, if you’re writ­ing things down with ink on pa­per, you’re prob­a­bly writ­ing just for your­self, or per­haps for fam­ily mem­bers who can learn to rec­og­nize your hand­writ­ing.

To judge from the way peo­ple talk, a lot of them have mem­o­ries of foun­tain pen dis­as­ters where ink got all over the place for some rea­son. Or per­haps it’s just gen­er­a­tional trauma, handed down in an oral tra­di­tion. If the pen is work­ing cor­rectly, ink can only come out of it so fast. A cou­ple of rare ex­cep­tions:

If the pen’s ink reser­voir is partly empty, so that it con­tains an air bub­ble, and if it’s po­si­tioned nib down, then, when you go up in an air­plane, the bub­ble will ex­pand as the am­bi­ent pres­sure drops, forc­ing ink out the nib. Once I fig­ured that out, I got in the habit of mak­ing sure my pens were po­si­tioned nib up when tak­ing off in an air­plane. If I have time I’ll also re­fill the pen be­fore de­par­ture, to min­i­mize the size of the air bub­ble.

Sometimes if a pen gets dirty, or if the nib is some­how dam­aged, the ink will stop com­ing out and you can restart it by giv­ing it a lit­tle shake. If you do it just right, the ink flow restarts with­out in­ci­dent, but if you overdo it, a few drops of ink might shoot out onto the page and be­come blots. This sce­nario hap­pens a few times of year for me, only with one pen that has this prob­lem. I blot it with a piece of scrap pa­per and move on.

Just have note­books ly­ing around, or on your per­son. Write gro­cery lists, doo­dles, notes on meet­ings, to-do lists, or stray ideas. Journal. Copy out good lines from books. Anything that has your men­tal fo­cus will have a more en­dur­ing pres­ence in your brain if you write it down.

I am left handed. I have never had any trou­ble with my hand smear­ing the ink. Yet every con­ver­sa­tion I have about foun­tain pens leads to some­one claim­ing that it can never work for them be­cause they are left handed. I have no idea what they’re talk­ing about. When I was a child, writ­ing at length with pen­cil, the side of my hand some­times be­came gray from graphite picked up as my hand rubbed across the page. And some­times I have got ink on my hand when us­ing a ball­point pen that left an ink glob on the pa­per. But with foun­tain pens it’s easy to find a pen/​pa­per com­bi­na­tion such that the ink soaks into the pa­per and dries quickly enough that it does­n’t smudge when you’re writ­ing the next line. Here’s a sim­ple demon­stra­tion of dry­ing time and how it works with two pens: first a foun­tain pen and then a Pilot G-2 gel pen.

Obviously the Pilot gel pen ink dries faster, and so that might be a bet­ter choice for peo­ple who are re­ally wor­ried about smudg­ing.

Most mod­ern pens al­low you to choose be­tween us­ing pre­loaded plas­tic ink car­tridges and a plunger that en­ables you to draw ink up out of a bot­tle by hand. I use both. Start with the ink car­tridges, es­pe­cially if you travel. There’s no need to com­pli­cate mat­ters by mess­ing around with bot­tles. Since I do a lot of work from one lo­ca­tion, I have a cor­ner of a table­top set up there with ink bot­tles and a folded-up pa­per towel for wip­ing off the nib af­ter it’s filled (I have been us­ing the same pa­per towel for about twenty years). In the­ory this works bet­ter in the long term be­cause it al­lows you to flush the nib by forc­ing ink in and out of it a cou­ple of times when­ever you re­fill. In prac­tice I see no dif­fer­ence at all - pens that I re­fill with car­tridges don’t get clogged.

Even if every­thing works per­fectly you’ll end up with the oc­ca­sional ink-smudged fin­ger. It will wash off quickly - the ink is wa­ter-sol­u­ble. Until then, con­sider it a mark of dis­tinc­tion.

If you’re new to this I think it makes most sense to start by con­sid­er­ing what kind of pa­per is go­ing to work best in your life. Are you writ­ing loose­leaf, or in note­books? Legal pads? Blue books? Remember, it’s okay to use lots of pa­per, so pick some­thing that is­n’t too pre­cious and that is easy to re­plen­ish. I use a lot of mole­sk­ine note­books and Mead comp books, which I can buy in bulk on­line. For com­pos­ing fic­tion I use fancy loose­leaf pa­per.

If you have ac­cess to a store where they sell foun­tain pens, take some of that pa­per there and see what works best. If you’re work­ing with cheaper, thin­ner pa­per, start with finer nibs and work up to fat­ter ones un­til you start to see bleed-through.

Buy cheaper pens un­til you know what you like. I doubt there’s much of a dif­fer­ence be­tween cheaper and more ex­pen­sive foun­tain pens in terms of their ac­tual per­for­mance. What you’re pay­ing for, in an ex­pen­sive pen, is fancy ma­te­ri­als and styling. For ex­am­ple, if you look at the Pilot Vanishing Point line of pens - an in­ge­nious foun­tain pen that you can click, like an old-fash­ioned ball­point, to re­tract the nib in­side the bar­rel - fancier ver­sions cost five times as much as the base model.

In all hon­esty, the Pilot G-2 gel pens are go­ing to give you 80% of what you could ex­pect from a foun­tain pen for min­i­mal cost.

On the other hand, a ten-pack of Pilot G-2 gel pens goes for about twenty bucks. For the same amount you can buy a sim­ple but com­pletely ser­vice­able foun­tain pen that will last longer than you will.

No posts

Introducing Claude Opus 5

www.anthropic.com

Claude Opus 5 is avail­able to­day. It’s a thought­ful and proac­tive model that comes close to the fron­tier in­tel­li­gence of Claude Fable 5 at half the price.

On cod­ing and knowl­edge work eval­u­a­tions like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it re­mains be­hind Mythos 5 on cy­ber­se­cu­rity tasks.

Opus 5 is de­signed to be used every day: it works more ef­fi­ciently than other mod­els. It’s the new de­fault model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-ef­fec­tive­ness

Claude Opus 5 pro­vides greatly im­proved per­for­mance for the same cost as its pre­de­ces­sor, Opus 4.8. The charts in this sec­tion show how per­for­mance changes ac­cord­ing to the mod­el’s ef­fort set­ting, which cus­tomers can use to op­ti­mize for in­tel­li­gence or con­serve to­kens for faster and cheaper re­sults.

Opus 5 ex­cels on valu­able soft­ware en­gi­neer­ing tasks. For ex­am­ple, on Frontier-Bench v0.1, Opus 5 sur­passes all other mod­els, and more than dou­bles Opus 4.8’s per­for­mance at a lower cost per task. On CursorBench 3.2, at max ef­fort, the model per­forms within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater per­for­mance at a given cost than all other mod­els on high, xhigh, and max ef­fort.

We see sim­i­lar re­sults on knowl­edge work and prob­lem-solv­ing tasks. For ex­am­ple:

On ARC-AGI 3, an eval­u­a­tion where the model has to solve novel prob­lems, Opus 5’s score is three times as high as the next-best model.

On Zapier AutomationBench, which mea­sures whether mod­els can com­plete busi­ness tasks from start to fin­ish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its low­est ef­fort set­ting, Opus 5 passes more tasks than any other model.

On OSWorld 2.0, a com­puter use bench­mark, Opus 5 out­per­forms every other model at any given cost, sur­pass­ing Fable 5’s best re­sult at just over a third of the cost.

It’s also our best and most cost-ef­fi­cient model on sev­eral re­lated eval­u­a­tions:

Opus 5 is a mean­ing­ful im­prove­ment over Opus 4.8 for sci­en­tific re­search. It shows bet­ter per­for­mance than Opus 4.8 on every one of our life sci­ences eval­u­a­tions, which cover top­ics in­clud­ing struc­tural bi­ol­ogy, or­ganic chem­istry, and bioin­for­mat­ics. Its im­prove­ments are most no­table on or­ganic chem­istry tasks, like in­fer­ring mol­e­c­u­lar struc­tures from spec­troscopy data (it scores 10.2 per­cent­age points higher than Opus 4.8 on our in­ter­nal bench­mark), and on pro­tein-re­lated tasks like pre­dict­ing how vari­a­tions in a pro­tein’s se­quence af­fect how it func­tions (here, it scores 7.7 per­cent­age points higher).

Finally, Opus 5 is ca­pa­ble of pro­duc­ing much stronger vi­sual out­puts:

Working with Claude Opus 5

Claude Opus 5 is much stronger at ver­i­fy­ing its work and it­er­at­ing care­fully un­til it suc­ceeds. In eval­u­a­tions and early-ac­cess test­ing, we and our users found many ex­am­ples of Opus 5’s agency and thor­ough­ness:

On one Frontier-Bench task, Opus 5 was given a draw­ing of a ma­chine part and asked to write code to re­build it as a 3D FreeCAD model. However, in this task, the model was in­ten­tion­ally given no way to di­rectly view the draw­ing. Opus 5 re­sponded by writ­ing its own com­puter vi­sion pipeline to pull the geom­e­try from the raw pix­els, then re­con­structed the full ma­chine part. It suc­ceeded in do­ing so re­peat­edly; no com­pet­ing model with the same setup could solve it af­ter five at­tempts.

Given a real bug in a pop­u­lar open-source pack­age man­ager, Opus 5 found the root cause and fixed an edge case that the com­mu­ni­ty’s patch had missed. A com­pet­ing model fixed only the sur­face symp­tom (not the un­der­ly­ing cause), then re­ported the bug re­solved.

An en­gi­neer at a trad­ing firm used Opus 5 to build a mar­ket data feed for a new ex­change in a sin­gle ses­sion. Previous mod­els could not com­plete this task at all, even given ex­ten­sive plans from the en­gi­neer. Finding no live feed to val­i­date against, Opus 5 even built its own test har­ness to check that its code parsed the ex­change’s data cor­rectly.

Below are fur­ther re­ports from our early-ac­cess cus­tomers on their ex­pe­ri­ence of work­ing with Opus 5:

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

01 /

22

Alignment and safety

Alignment. During pre-de­ploy­ment test­ing, our au­to­mated be­hav­ioral au­dit found Opus 5 to be our most aligned model to date (as shown in the graph be­low). It ad­heres to Claude’s Constitution bet­ter than Opus 4.8, Sonnet 5, or Fable 5; ex­hibits the low­est rates of de­cep­tive be­hav­ior; and is the least sus­cep­ti­ble to be­ing tricked into mis­use. It’s also our safest model yet in terms of avoid­ing reck­less ac­tions that could have hard-to-re­verse side ef­fects.

Safety. Opus 5 does not ad­vance the fron­tier in risky, dual-use ca­pa­bil­i­ties. In rig­or­ous eval­u­a­tions con­ducted along­side pri­vate-sec­tor and gov­ern­ment part­ners, we found it re­mains be­hind Mythos 5 in both bi­ol­ogy re­search and of­fen­sive cy­ber­se­cu­rity. More in­for­ma­tion about these eval­u­a­tions can be found in our System Card.

As with its pre­de­ces­sor, Opus 4.8, we’ve in­ten­tion­ally avoided train­ing Opus 5 on cy­ber tasks. The model has nev­er­the­less im­proved sub­stan­tially on these tasks as a re­sult of be­com­ing more gen­er­ally ca­pa­ble, and it comes close to Mythos 5 at find­ing cy­ber­se­cu­rity vul­ner­a­bil­i­ties. However, it re­mains sub­stan­tially be­hind Mythos 5 on the ex­ploita­tion of those vul­ner­a­bil­i­ties—that is, in turn­ing vul­ner­a­bil­i­ties into ma­te­r­ial cy­ber threats.

This is il­lus­trated by Opus 5’s per­for­mance on OSS-Fuzz, an eval­u­a­tion we’ve de­vel­oped to as­sess how well mod­els can find and then ex­ploit vul­ner­a­bil­i­ties with­out ex­ten­sive hu­man guid­ance. Although Mythos 5 and Opus 5 iden­tify vul­ner­a­bil­i­ties with sim­i­lar suc­cess, Opus 5’s score on the de­vel­op­ment of ex­ploits is far be­hind that of Mythos 5.

Safeguards for Opus 5

Claude Opus 5’s safe­guards are de­signed to al­low ben­e­fi­cial uses of the model in both cy­ber­se­cu­rity and bi­ol­ogy. They are sim­i­lar to those we ap­plied to Opus 4.8, with the ex­cep­tion of some stronger guardrails on a nar­row range of cy­ber tasks.

Cybersecurity. Opus 5’s cy­ber clas­si­fiers are pro­por­tion­ally less re­stric­tive than those on Fable 5. They al­low Opus 5 to find vul­ner­a­bil­i­ties in source code, but block binary-based” vul­ner­a­bil­ity scan­ning (a method more likely to be as­so­ci­ated with ma­li­cious ac­tors), pen­e­tra­tion test­ing, and ex­ploit gen­er­a­tion.

Based on our test­ing, we ex­pect the clas­si­fiers to in­ter­vene around 85% less of­ten than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged re­quests will fall back to Opus 4.8 by de­fault. Fallbacks to Opus 4.8 can also be en­abled on the API.

Our Cyber Verification Program (CVP) fa­cil­i­tates cy­ber­se­cu­rity work that would oth­er­wise be im­peded by the mod­el’s safe­guards. Enterprises and re­searchers who are al­ready part of the CVP have im­me­di­ate ac­cess to a ver­sion of Opus 5 with fewer se­cu­rity re­stric­tions.

Biology. Since Opus 5 has a sim­i­lar suite of safe­guards to Opus 4.8, it is now our most ca­pa­ble gen­er­ally avail­able model for sci­en­tific re­search. Nevertheless, the model still shows im­por­tant lim­i­ta­tions on long-run­ning, au­tonomous re­search tasks, which is where we ex­pect AI mod­els to pose the most sub­stan­tial bi­ol­ogy-re­lated risks. (Mythos 5 re­mains the stronger model for this type of bi­o­log­i­cal work.) As part of this launch, bi­ol­ogy-re­lated re­quests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.

Getting started

Claude Opus 5 is avail­able to­day on all plat­forms, priced at $5 per mil­lion in­put to­kens and $25 per mil­lion out­put to­kens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.

It’s also of­fered in Fast mode, where it runs around 2.5 times the de­fault speed. As with Opus 4.8, Fast mode is avail­able at twice Opus 5’s base price on the Claude Platform and through us­age cred­its in Claude Code.

Alongside Opus 5, we’re re­leas­ing two up­dates in beta:

Mid-conversation tool changes on the Claude Platform. Within a con­ver­sa­tion, de­vel­op­ers can now change which tools Claude can use with­out in­val­i­dat­ing the prompt cache.

Automatic fall­backs on the API. Users can now choose to have re­quests that are flagged by our safety clas­si­fiers on Opus 5 (or Fable 5) au­to­mat­i­cally route to an­other model. With au­to­matic fall­backs on, API re­quests al­ways route to the best avail­able model by de­fault rather than be­ing blocked.

Consistent with prior Opus mod­els, Opus 5 does not have data re­ten­tion re­quire­ments for gen­eral ac­cess.

For more guid­ance on how to get the best out of Opus 5, see our prompt­ing guide.

Footnotes

Frontier-Bench v0.1, Effort plot: These re­sults are from an in­ter­nal run of Frontier-Bench v0.1, on the mini-SWE-agent har­ness and a GKE back­end, mean re­ward over 5 at­tempts per task. Opus 4.8 served as fall­back on safety-clas­si­fier re­fusals for Opus 5 and Fable 5.

Related con­tent

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

Ask Claude about the Anthropic Economic Index

We’re launch­ing the Anthropic Economic Index con­nec­tor for Claude, which lets any­one ex­plore real data about AI and work.

Read more

Anthropic is do­nat­ing an­other $20 mil­lion to Public First Action

Anthropic is con­tribut­ing an ad­di­tional $20 mil­lion to Public First Action, bring­ing our to­tal sup­port to $40 mil­lion.

Read more

American AI is locked down and proprietary. It's losing.

werd.io

China’s open-weights AI strat­egy is win­ning: its com­pa­nies are tak­ing the lead. America’s closed-first, locked-down strat­egy is doomed to fail­ure - and it could take the US econ­omy down with it.

Link: China de­liv­ers a one-two punch to America’s AI dom­i­nance, by Robert Hart in The Verge

AI mod­els, as a prod­uct in them­selves, have very lit­tle moat be­yond what amounts to brand loy­alty and su­per­fi­cial switch­ing costs. Instead, the moat is in the en­ter­prise ser­vices that sit around them: the deals and con­tracts, con­nec­tiv­ity with en­ter­prise sys­tems, and qual­ity of life fea­tures in an en­ter­prise con­text.

If we con­sider the mod­els them­selves, it’s easy to switch be­tween them: some­one could be us­ing ChatGPT to­day and Claude to­mor­row, with very lit­tle im­pact on their work­flows. This is par­tic­u­larly true in the en­gi­neer­ing world, where mod­els are ac­cessed via API: you can swap out the API and use the same prompt.

Those com­pa­nies can make deals to lock their cus­tomers in, but in prac­tice there’s very lit­tle long-term tech­ni­cal in­cen­tive to use one ven­dor over an­other. You pick the best model for your needs and change mod­els and ven­dors if an­other one be­comes bet­ter.

The US gov­ern­ment has placed ex­port con­trols on GPUs. There are also strong reg­u­la­tions that (reasonably) pre­vent shar­ing cer­tain kinds of data with Chinese servers. The re­sult is that while Chinese com­pa­nies have enough com­pute to train mod­els, they can’t re­ally pro­vide the kinds of global-scale cen­tral­ized ser­vices that we see from OpenAI and Anthropic — at least, not in the same way.

And open al­most al­ways wins when it comes to in­fra­struc­ture adop­tion. Open tech­nolo­gies can be used per­mis­sion­lessly and there­fore can be at the cen­ter of more in­no­va­tion. You can host them where you want, ex­per­i­ment with them, al­ter them, and tweak to fit your use case. Open weights mod­els are not open source, but they are portable and per­mis­sion­less.

With all this in mind, it makes sense for China to re­lease its AI mod­els openly. It turns a US-created com­pute dis­ad­van­tage into a dis­tri­b­u­tion ad­van­tage; it com­modi­tizes the layer where American com­pa­nies make money; and it cre­ates a far more ef­fec­tive global ecosys­tem than could be es­tab­lished through locked-in, cen­tral­ized ser­vices. It’s ob­vi­ous to me that there are ecosys­tem ben­e­fits through­out China, from man­u­fac­tur­ing to sci­en­tific re­search; every sec­tor can just plug in these mod­els.

The sav­ing grace for American com­pa­nies has been that US fron­tier mod­els have out­per­formed open ones. That gap is now clos­ing:

Moonshot and Alibaba un­veiled mod­els they claim can go toe-to-toe with the best from OpenAI and Anthropic at a frac­tion of the cost. The rapid-fire re­leases sug­gest America’s lead at the AI fron­tier is in­creas­ingly tight, just as the tech­nol­ogy is be­com­ing cen­tral to na­tional se­cu­rity, eco­nomic power, and geopo­lit­i­cal in­flu­ence.”

Even with­out these new ca­pa­bil­i­ties, the strat­egy has al­ready been work­ing. a16z part­ner Martin Casado noted in the Economist that there’s an 80% chance that any given startup is us­ing Chinese mod­els, and Chinese mod­els are poised to take the lead.

It’s worth tak­ing a step back and con­sid­er­ing the sur­pris­ing un­der­ly­ing dy­nam­ics. We think of China as be­ing a locked-down so­ci­ety — and it is in many ways. I have se­ri­ous con­cerns about how these mod­els might re­flect Chinese gov­ern­ment per­spec­tives (try ask­ing them about Tiananmen Square). But it’s American com­pa­nies that are keep­ing tight con­trol of their tech­nol­ogy rather than re­leas­ing it as openly as pos­si­ble. This is in stark con­trast to the strat­egy be­hind US gov­ern­ment sup­port for the open in­ter­net, for ex­am­ple.

Locked-down busi­ness prac­tices for a tech­nol­ogy with no real moat but sig­nif­i­cant po­ten­tial ecosys­tem ben­e­fits is an ob­vi­ously los­ing strat­egy; per­mis­sively re­leas­ing it with an open, col­lab­o­ra­tive ap­proach is ob­vi­ously a win­ning one. But the in­cen­tives in the US aren’t there: in­stead, these com­pa­nies are forced to chase first-or­der prof­its rather than ecosys­tem ben­e­fits, and the gov­ern­ment tries to put its fin­ger on the scale through forcible mea­sures like tight ex­port con­trols. We should con­sider what would need to change to make those in­cen­tives more aligned. That’s par­tic­u­larly im­por­tant given how much of the US econ­omy is cur­rently dri­ven by AI spend­ing. If the bot­tom falls out of that spend­ing — and I think it clearly will, given the dy­nam­ics — the out­come could be se­vere.

I care about hav­ing open tech­nol­ogy that can be run in the pub­lic in­ter­est, aligned with the pub­lic’s val­ues. Threads like pub­lic AI, fed­er­ated ser­vices, and open re­search have trac­tion but need back­ing. Getting there in the US needs more nu­anced strat­egy and sup­port than we’re see­ing to­day.

Access denied | videocardz.com used Cloudflare to restrict access | videocardz.com

videocardz.com

Please en­able cook­ies.

Error 1005

What hap­pened?

The owner of this web­site (videocardz.com) has banned the au­tonomous sys­tem num­ber (ASN) your IP ad­dress is in (14061) from ac­cess­ing this web­site.

Please see https://​de­vel­op­ers.cloud­flare.com/​sup­port/​trou­bleshoot­ing/​http-sta­tus-codes/​cloud­flare-1xxx-er­rors/​er­ror-1005/ for more de­tails.

Was this page help­ful?

Thank you for your feed­back!

Cloudflare Ray ID: a1e17eec­c9ce8153 •

Your IP:

167.99.127.20 •

Performance & se­cu­rity by Cloudflare

Check out this chat

chatgpt.com

Get re­sponses tai­lored to you

Log in to get an­swers based on saved chats, plus cre­ate im­ages and up­load files.

Just a moment...

ads.openai.com

Just a moment...

www.politico.com

Bento Slides

bento.page

Who’s Afraid of Chinese Models?

stratechery.com

Listen to this post:

There’s a story I tell about my first day in STRT-431 at Kellogg School of Management, the in­tro­duc­tory class that every first-year MBA was re­quired to take; I leafed through the read­ings and case stud­ies and was dis­mayed that there weren’t any tech com­pa­nies on the docket. Me be­ing me, I spoke to the pro­fes­sor af­ter class won­der­ing why, and was told that the goal of the course was not to nec­es­sar­ily learn about spe­cific in­dus­tries, but rather to un­cover broadly ap­plic­a­ble uni­ver­sal prin­ci­ples that could be ap­plied to any com­pany in any in­dus­try.

I did not, as I usu­ally tell the story, find this very sat­is­fac­tory: to me the na­ture of tech, par­tic­u­larly the fact that soft­ware and dis­tri­b­u­tion had zero mar­ginal costs (and zero trans­ac­tion costs), was some­thing fun­da­men­tally dif­fer­ent; putting in ze­roes in for­mu­las tends to wreak havoc! I soon re­al­ized, how­ever, that that was my op­por­tu­nity. The fun­da­men­tal in­sight un­der­gird­ing Aggregation Theory is that zero mar­ginal costs leads to fun­da­men­tally dif­fer­ent value chains than peo­ple once ex­pected from the Internet: cen­tral­iza­tion and scale in a world where con­trol­ling de­mand mat­tered more than dis­trib­ut­ing sup­ply.

What is fas­ci­nat­ing about AI, how­ever, is the ex­tent to which those old uni­ver­sal prin­ci­ples are com­ing back to the fore­front. That was never more ap­par­ent than this past week­end, when ar­gu­ments raged on X about the im­pli­ca­tions of Kimi K3, an­other open weights model out of China, ap­proach­ing the state-of-the-art in terms of ca­pa­bil­i­ties. The long and short of it is this: mar­ginal costs are back in a big way, both in terms of short-term im­pli­ca­tions of state-of-the-art free mod­els, and in terms of the long-term struc­ture of the in­dus­try.

COGS Versus R&D

One of the most com­mon mis­con­cep­tions un­der­gird­ing dis­cus­sion of open weights mod­els is that they are cheaper — free, even. After all, you can just down­load the weights, and skip the time and ex­pense and ca­pa­bil­i­ties nec­es­sary to cre­ate your own model. That is, of course, true, but the free” in this case is a ref­er­ence to the amount you need to spend on re­search and de­vel­op­ment; R&D is a fixed ex­pense that is in­de­pen­dent of the rev­enue you gen­er­ate. If you spend $1 mil­lion in R&D, it does­n’t mat­ter if you do $100 thou­sand in rev­enue or $100 mil­lion; you still spent $1 mil­lion on R&D (it does, of course, im­pact your prof­itabil­ity).

What is re­lated to rev­enue is COGS — cost of goods sold — and COGS is real for AI in a way it has­n’t been for soft­ware for a very long time. Specifically, run­ning in­fer­ence on a model — whether that model be Kimi or Fable — costs money, and the amount of money an AI provider spends on in­fer­ence is, at least in most busi­ness mod­els, di­rectly cor­re­lated to rev­enue. To reuse the above ex­am­ple, gen­er­at­ing $100 mil­lion ver­sus $100 thou­sand in rev­enue will likely re­quire 1,000x COGS. In con­crete terms, if it costs 50 cents to gen­er­ate the to­kens that drive $1 in rev­enue, then $100 mil­lion in rev­enue will have $50 mil­lion in COGS; $100 thou­sand in rev­enue will only have $50 thou­sand in COGS.

The point in terms of open weight mod­els is that they are not free to serve. Kimi K3 costs $3 per mil­lion in­put to­kens, and $15 per mil­lion out­put to­kens; that is cheaper than Sol’s $5 per mil­lion in­put to­kens and $30 per mil­lion out­put to­kens, but that might not even be the right mea­sure­ment.

Tokens Versus Intelligence

Nvidia CEO Jensen Huang has de­scribed what Nvidia is build­ing as token fac­to­ries”, and from Nvidia’s per­spec­tive that fram­ing makes sense. Nvidia GPUs are model ag­nos­tic: they gen­er­ate to­kens, and do so in the fastest and most ef­fi­cient way pos­si­ble. That leads to mea­sure­ments like to­kens-per-sec­ond, time-to-first-to­ken, to­kens-per-watt, to­ken cost, etc., and Huang ar­gues that these met­rics will be the ba­sis for de­ci­sion-mak­ing.

This is a fram­ing that def­i­nitely made sense dur­ing the first par­a­digm of AI, the ChatGPT era, when to­kens were de­liv­ered straight to the end user. The sec­ond par­a­digm of AI, how­ever, the rea­son­ing era, con­founds this mea­sure­ment. Reasoning en­tails an ex­plo­sion in chain-of-thought to­kens, and dif­fer­ent mod­els need dif­fer­ent amounts of rea­son­ing to­kens to ar­rive at the right an­swer. Kimi, for ex­am­ple, re­port­edly uses sig­nif­i­cantly more to­kens than Sol, ren­der­ing its price ad­van­tage moot. Agents in­tro­duce a sim­i­lar dy­namic: some mod­els are more ef­fi­cient than oth­ers in terms of the num­ber of to­kens they need to ex­e­cute agen­tic work­flows.

What this means is that to­kens are not a com­mod­ity. The defin­ing char­ac­ter­is­tic of a com­mod­ity is that it is fun­gi­ble: a gal­lon of oil is a gal­lon of oil; a ton of cop­per is a ton of cop­per; a bushel of wheat is a bushel of wheat. A to­ken from one model, how­ever, is not the same as a to­ken from an­other model. What is fun­gi­ble is what is con­structed from to­kens, which is to say in­tel­li­gence. In other words, if both Kimi and Sol gen­er­ated the right an­swer, then that an­swer is fun­gi­ble; the dif­fer­ence in to­kens gen­er­ated to get to that right an­swer is a con­trib­u­tor to a dif­fer­ence in COGS.

The COGS for in­tel­li­gence is a func­tion of a few dif­fer­ent fac­tors:

Model foot­print: The weights and run­time state de­ter­mine how much ex­pen­sive mem­ory and how many ac­cel­er­a­tors are re­quired to host each serv­ing replica.

Inference ef­fi­ciency: Architectural choices (e.g. Mixture-of-Experts) re­duce com­pu­ta­tion per gen­er­ated to­ken.

Memory ef­fi­ciency: Architectural choices can re­duce KV cache re­quire­ments, al­low­ing more con­cur­rent re­quests and bet­ter GPU uti­liza­tion.

Serving ef­fi­ciency: Batching, sched­ul­ing, pre­fix caching, and other in­fer­ence op­ti­miza­tions max­i­mize uti­liza­tion and share work across re­quests.

Token ef­fi­ciency: The fewer to­kens re­quired to reach a cor­rect an­swer, the lower the in­fer­ence cost.

The rea­son this mat­ters is that we are rapidly ap­proach­ing a state in which in­tel­li­gence for many eco­nom­i­cally ben­e­fi­cial tasks is in fact a com­mod­ity. Anyone build­ing a ba­sic CRUD app, for ex­am­ple, can likely do so us­ing mod­els from mul­ti­ple providers. And, in a com­mod­ity mar­ket, the route to prof­itabil­ity is not through charg­ing higher prices — again, you can (or will soon be able to) make the ex­act same app us­ing mul­ti­ple mod­els — but rather through hav­ing a su­pe­rior cost struc­ture.

Understanding Commodity Markets

It’s worth step­ping through the me­chan­ics here, be­cause, as I noted a few months ago in Amazon’s Durability, the dy­nam­ics of com­mod­ity mar­kets are not some­thing peo­ple in tech are gen­er­ally fa­mil­iar with:

In com­mod­ity mar­kets, every­one charges the same price, be­cause every­one is sell­ing the same thing; that price is de­ter­mined by sup­ply and de­mand.

The de­mand for a com­mod­ity is a func­tion of price elas­tic­ity: the cheaper the com­mod­ity, the more de­mand there is for it, and vice-versa.

The sup­ply for a com­mod­ity is a func­tion of the mar­ginal cost of pro­duc­ing the com­mod­ity.

The key thing to un­der­stand is that the mar­ginal cost of pro­duc­ing the com­mod­ity dif­fers by sup­plier. What this means in prac­tice is that the sup­plier with the worst cost struc­ture ends up sell­ing the com­mod­ity at their mar­ginal cost (if they can pro­duce at all); the prof­its of every­one else de­pend on the ex­tent to which their cost struc­ture is bet­ter than the mar­ginal sup­plier.

As an ex­am­ple:

Supplier A can pro­duce 10 units of the com­mod­ity for $10 each

Supplier B can pro­duce 10 units of the com­mod­ity for $15 each

Supplier C can pro­duce 10 units of the com­mod­ity for $20 each

Let’s as­sume the price elas­tic­ity is such that there is de­mand for 25 units of the com­mod­ity at $20. That means:

Supplier A will sell 10 units of the com­mod­ity for $20, earn­ing $10/unit

Supplier B will sell 10 units of the com­mod­ity for $20, earn­ing $5/unit

Supplier C will sell 5 units of the com­mod­ity for $20, earn­ing $0/unit

This is­n’t pre­cisely right: the rea­son why Supplier C will bear the short­fall is be­cause Suppliers A and B will be able to slightly un­der­cut them in price, which will of course af­fect de­mand (which is elas­tic), but it makes the point. Supplier A has a great busi­ness, Supplier B has a good busi­ness, and Supplier C is go­ing to go bank­rupt.

Bankruptcy risk is where fixed costs come back to the fore­front: Supplier C has both fixed costs (like po­ten­tially R&D spend) and also may have taken on debt to fi­nance the equip­ment nec­es­sary to pro­duce the com­mod­ity. It can’t price its com­mod­ity with these costs in mind — re­mem­ber, the mar­ket-clear­ing price ap­prox­i­mates the mar­ginal cost of the high­est-cost unit needed to sat­isfy de­mand — but those costs can ab­solutely drive the sup­plier out of busi­ness. And, if that sup­plier goes out of busi­ness, then prices go up, un­til an­other sup­plier de­cides to en­ter (or the other sup­pli­ers ex­pand).

The Intelligence Market

Let’s bring this back to mod­els. Right now, none of the above analy­sis ap­plies be­cause de­mand ex­ceeds sup­ply for fron­tier mod­els, and sup­ply is lim­ited by a lack of com­pute. This com­pute short­age does­n’t just mean that a com­pute sup­plier like Nvidia makes very large mar­gins, but also that Nvidia’s cus­tomers, like SpaceXAI, can turn around and re­sell com­pute at high mar­gins as well to a com­pany like Anthropic. Anthropic, mean­while, can pay the markup be­cause they can sell to­kens with a higher markup still.

It’s not just ex­cess de­mand that gives Anthropic great mar­gins, how­ever: Anthropic and OpenAI likely have among the low­est costs per unit of fron­tier-qual­ity in­tel­li­gence, thanks to model ca­pa­bil­ity, serv­ing scale, and to­ken ef­fi­ciency. They are serv­ing mod­els at a par­tic­u­lar ca­pa­bil­ity level for months be­fore their com­peti­tors, and are si­mul­ta­ne­ously ap­ply­ing the best mod­els to op­ti­miz­ing those costs.

It’s also worth not­ing that the mar­ket is not yet treat­ing in­tel­li­gence like a com­mod­ity: de­mand is for Anthropic and OpenAI specif­i­cally, and much less for mod­els that aren’t as good (thus SpaceXAI and Meta sell­ing ca­pac­ity to Anthropic); one way to think about the push for op­ti­miz­ing cost is that that is a func­tion of defin­ing jobs-to-be-done by in­tel­li­gence level, such that in­tel­li­gence buy­ers can cre­ate a mar­ket where in­tel­li­gence is com­modi­tized. In the long run, how­ever, who­ever is on the fron­tier is the best placed to dom­i­nate non-fron­tier mar­kets as well, which are just the fron­tier mi­nus n-months, i.e. months in which the fron­tier model mak­ers have been op­ti­miz­ing their cost of serv­ing.

All of this is to say that I think the re­ac­tion to Kimi and Chinese mod­els gen­er­ally is pretty over-blown, at least from an eco­nomic per­spec­tive. Right now there is a price um­brella that is down­stream of the lack of com­pute; I highly doubt that Chinese mod­els are cheaper to serve on a mar­ginal cost ba­sis, they just seem cheaper be­cause Anthropic and OpenAI are so sup­ply con­strained that they are charg­ing far more than they would if there were suf­fi­cient sup­ply to meet the de­mand for in­tel­li­gence.

Frontier Lab Paranoia

Why, then, do the model mak­ers in par­tic­u­lar seem so pan­icked about Chinese mod­els?

First, I think the fron­tier labs are an­chored in a world where train­ing costs dom­i­nated their fi­nan­cial mod­el­ing. As long as train­ing con­sumed more GPUs than in­fer­ence, it was crit­i­cal to max­i­mize in­fer­ence rev­enue to help fund the next train­ing run, which meant charg­ing very high prices for in­fer­ence.

Going for­ward, how­ever, I ex­pect the in­fer­ence mar­ket to grow much faster than train­ing costs (and that in­cludes the as­sump­tion that train­ing costs will con­tinue to sky­rocket), which means they re­ally can make it up in vol­ume. It was­n’t clear this would be the case as re­cently as eight months ago, but the agent par­a­digm un­lock is so mas­sive that fron­tier labs should have more con­fi­dence that they can not just sur­vive but thrive with lower prices (once they have suf­fi­cient com­pute).

Second, in­tel­li­gence is­n’t in fact a per­fect com­mod­ity, in part be­cause ap­plied in­tel­li­gence makes it­self smarter. Specifically, who­ever is run­ning in­fer­ence is also col­lect­ing data, and that data goes into mak­ing the next it­er­a­tion of the model bet­ter. This is, on one hand, all the more rea­son for the fron­tier labs to lower prices and in­crease us­age as more com­pute comes on­line; on the other hand, this is why com­pa­nies like Microsoft are in­creas­ingly ob­sessed with help­ing com­pa­nies run their own mod­els. That is much more vi­able if Chinese mod­els are a vi­able al­ter­na­tive.

Third, the other way that fron­tier labs can not only dif­fer­en­ti­ate from Chinese mod­els but also from each other is by con­tin­u­ing to in­te­grate up into the cus­tomer ex­pe­ri­ence. It’s strik­ing the ex­tent to which Claude Code and Codex are prov­ing to be quite sticky; whichever har­ness you start work­ing with is likely to be the one you stick with, and that fig­ures to be even more the case with non-tech­ni­cal users. And, in the long run, this im­per­a­tive to move up the stack does mean that fron­tier mod­els are ab­solutely a threat to soft­ware providers, in­clud­ing Microsoft. On the flip­side, the ex­tent to which soft­ware com­pa­nies who cur­rently own the cus­tomer ex­pe­ri­ence have ac­cess to com­pet­i­tive mod­els is the ex­tent to which they may be able to re­sist the en­croach­ment of the fron­tier labs.

Finally, the ide­o­log­i­cal an­gle of Anthropic in par­tic­u­lar is im­pos­si­ble to ig­nore. This is a com­pany that be­lieves only it can be en­trusted with AI, and the ex­is­tence of open weights al­ter­na­tives strikes a fa­tal blow to that pre­sump­tion.

China’s Motivation

Kimi is­n’t the only new Chinese model; from Bloomberg:

Alibaba Group Holding Ltd. shares rose as much as 5.4% on Monday af­ter the com­pany launched a pre­view ver­sion of its flag­ship Qwen3.8 Max model, de­scrib­ing it as sec­ond only to Anthropic PBCs Fable 5. The Sunday re­lease came only days af­ter startup Moonshot AI un­veiled a pow­er­ful new of­fer­ing that’s roiled mar­kets and trig­gered con­cern in the US about China clos­ing the gap on global lead­ers like Anthropic and OpenAI. Qwen3.8 Max has 2.4 tril­lion pa­ra­me­ters, join­ing Moonshot’s Kimi K3 in the heavy­weight class. With 2.8 tril­lion pa­ra­me­ters, K3 ri­vals top of­fer­ings and Alibaba is set­ting sim­i­larly high ex­pec­ta­tions.

Developers can now ac­cess Qwen3.8 Max through Alibaba’s cod­ing plat­forms, in­clud­ing Qoder. Alibaba plans to make the model open-weight soon, ex­pand­ing ac­cess be­yond the pre­view re­lease. Interest in these made-in-China ar­ti­fi­cial in­tel­li­gence sys­tems and mod­els is so high that Moonshot was forced to pause tak­ing on new sub­scrip­tions late on Sunday to man­age over­whelm­ing de­mand.

Alibaba Group Holding Ltd. shares rose as much as 5.4% on Monday af­ter the com­pany launched a pre­view ver­sion of its flag­ship Qwen3.8 Max model, de­scrib­ing it as sec­ond only to Anthropic PBCs Fable 5. The Sunday re­lease came only days af­ter startup Moonshot AI un­veiled a pow­er­ful new of­fer­ing that’s roiled mar­kets and trig­gered con­cern in the US about China clos­ing the gap on global lead­ers like Anthropic and OpenAI. Qwen3.8 Max has 2.4 tril­lion pa­ra­me­ters, join­ing Moonshot’s Kimi K3 in the heavy­weight class. With 2.8 tril­lion pa­ra­me­ters, K3 ri­vals top of­fer­ings and Alibaba is set­ting sim­i­larly high ex­pec­ta­tions.

Developers can now ac­cess Qwen3.8 Max through Alibaba’s cod­ing plat­forms, in­clud­ing Qoder. Alibaba plans to make the model open-weight soon, ex­pand­ing ac­cess be­yond the pre­view re­lease. Interest in these made-in-China ar­ti­fi­cial in­tel­li­gence sys­tems and mod­els is so high that Moonshot was forced to pause tak­ing on new sub­scrip­tions late on Sunday to man­age over­whelm­ing de­mand.

The fact that Qwen3.8 Max will also have open weights is no­table. Alibaba stopped re­leas­ing weights for its lead­ing edge mod­els ear­lier this year, but ap­pears to have re­verted that change; I sus­pect that shift was re­lated to last week’s Xi Jinping speech about AI that dou­bled down on the open weights ap­proach:

We should ad­here to the prin­ci­ple of open­ness and win-win and boost in­no­va­tion-dri­ven de­vel­op­ment. As a new en­gine of world eco­nomic growth and an ac­cel­er­a­tor for the shift of growth dri­vers, AI is mov­ing from the dig­i­tal world into the phys­i­cal world. We should seize this rare, his­toric op­por­tu­nity to en­cour­age open source, open­ness, col­lab­o­ra­tion and shar­ing. We should fa­cil­i­tate tech­no­log­i­cal in­no­va­tion, in­dus­trial de­vel­op­ment and sce­nario-based ap­pli­ca­tion of AI. We should make co­or­di­nated ad­vances in the trans­for­ma­tion and up­grade of tra­di­tional in­dus­tries, the cul­ti­va­tion and growth of emerg­ing in­dus­tries and for­ward-look­ing plan­ning for fu­ture in­dus­tries, so that all sec­tors and busi­nesses can ben­e­fit from AI.

We should ad­here to the prin­ci­ple of open­ness and win-win and boost in­no­va­tion-dri­ven de­vel­op­ment. As a new en­gine of world eco­nomic growth and an ac­cel­er­a­tor for the shift of growth dri­vers, AI is mov­ing from the dig­i­tal world into the phys­i­cal world. We should seize this rare, his­toric op­por­tu­nity to en­cour­age open source, open­ness, col­lab­o­ra­tion and shar­ing. We should fa­cil­i­tate tech­no­log­i­cal in­no­va­tion, in­dus­trial de­vel­op­ment and sce­nario-based ap­pli­ca­tion of AI. We should make co­or­di­nated ad­vances in the trans­for­ma­tion and up­grade of tra­di­tional in­dus­tries, the cul­ti­va­tion and growth of emerg­ing in­dus­tries and for­ward-look­ing plan­ning for fu­ture in­dus­tries, so that all sec­tors and busi­nesses can ben­e­fit from AI.

The strat­egy for China is ob­vi­ous: com­modi­tize your com­ple­ments. Note that Xi ex­plic­itly ties open­ness to AI moving from the dig­i­tal world into the phys­i­cal world”; the phys­i­cal world is the world dom­i­nated by China, and the coun­try’s lead in ar­eas like ro­bot­ics is go­ing to mas­sively ben­e­fit from widely avail­able AI mod­els.

Along the same lines, China does not want the U.S. to gain an asym­met­ric ad­van­tage in AI; to the ex­tent that China can weaken the U.S. fron­tier labs while strength­en­ing any and all po­ten­tial U.S. ad­ver­saries so much the bet­ter, and it can ben­e­fit from the in­no­va­tion that will at­tach it­self to an open ecosys­tem.

The Distillation Question

By the same to­ken, don’t ex­pect China to do any­thing about dis­til­la­tion at­tacks on the fron­tier labs. I think it is mis­taken to at­tribute all of the suc­cess of Chinese labs to dis­til­la­tion, but it’s just as much of a mis­take to pre­tend like dis­til­la­tion does­n’t give Chinese labs a big ad­van­tage. That ad­van­tage has re­ally come to bear in the last year as post-train­ing re­in­force­ment learn­ing has be­come in­creas­ingly cru­cial to model per­for­mance. Instead of hav­ing to fash­ion re­in­force­ment learn­ing en­vi­ron­ments from scratch, Chinese labs can sim­ply use fron­tier labs mod­els as teach­ers, al­low­ing for rapid im­prove­ment at much lower costs (this is not the only rea­son why Chinese mod­els are cheaper to de­velop, but it’s a big one).

What is in­ter­est­ing is that one of the most im­por­tant use cases for Chinese mod­els in the West is it­self dis­til­la­tion. Thinking Machines, for ex­am­ple, which just re­leased an open-weight model, re­lies on Chinese mod­els to solve the cold start prob­lem for re­in­force­ment learn­ing. Dean Meyer and Konstantine Buhler wrote an ex­cel­lent ar­ti­cle on X ex­plain­ing that dis­til­la­tion means that Western open weight mod­els are fun­da­men­tally dis­ad­van­taged rel­a­tive to China:

Distillation does not ex­plain China’s en­tire open-model lead. Chinese labs have world-class re­searchers, sub­stan­tial com­pute, strong pre-trained mod­els, soft­ware-hard­ware code­sign, and rapidly im­prov­ing post-train­ing ca­pa­bil­i­ties. But dis­til­la­tion com­presses the costly fi­nal gap be­tween a strong base and a near-fron­tier sys­tem. Even if dis­til­la­tion rep­re­sents a smaller share of a Chinese mod­el’s to­tal ca­pa­bil­ity, it rep­re­sents a mean­ing­ful share of its ad­van­tage over American open mod­els.

New en­force­ment mech­a­nisms will make large-scale dis­til­la­tion harder, slower, and more ex­pen­sive for Chinese com­pa­nies. However, en­force­ment will not elim­i­nate dis­til­la­tion backed by state ac­tors. Every Western fron­tier ad­vance there­fore cre­ates an­other teacher for Chinese labs. Western builders must ei­ther re­pro­duce those ca­pa­bil­i­ties in­de­pen­dently or wait to learn from Chinese mod­els. This gap gives Chinese labs a re­cur­ring struc­tural ad­van­tage over Western com­pa­nies.

Distillation does not ex­plain China’s en­tire open-model lead. Chinese labs have world-class re­searchers, sub­stan­tial com­pute, strong pre-trained mod­els, soft­ware-hard­ware code­sign, and rapidly im­prov­ing post-train­ing ca­pa­bil­i­ties. But dis­til­la­tion com­presses the costly fi­nal gap be­tween a strong base and a near-fron­tier sys­tem. Even if dis­til­la­tion rep­re­sents a smaller share of a Chinese mod­el’s to­tal ca­pa­bil­ity, it rep­re­sents a mean­ing­ful share of its ad­van­tage over American open mod­els.

New en­force­ment mech­a­nisms will make large-scale dis­til­la­tion harder, slower, and more ex­pen­sive for Chinese com­pa­nies. However, en­force­ment will not elim­i­nate dis­til­la­tion backed by state ac­tors. Every Western fron­tier ad­vance there­fore cre­ates an­other teacher for Chinese labs. Western builders must ei­ther re­pro­duce those ca­pa­bil­i­ties in­de­pen­dently or wait to learn from Chinese mod­els. This gap gives Chinese labs a re­cur­ring struc­tural ad­van­tage over Western com­pa­nies.

This is a point that bears re­peat­ing: be­cause U.S. open weight model mak­ers must fol­low the fron­tier labs’ terms of ser­vice, they (1) are worse than Chinese al­ter­na­tives and (2) end up dis­till­ing the dis­til­la­tion, just with a de­tour through Chinese labs. Wouldn’t it be bet­ter if west­ern open weight model mak­ers could go to the source?

To that end, here’s an even more in­ter­est­ing ques­tion around dis­til­la­tion: why ex­actly is it bad? After all, what are large lan­guage mod­els but the dis­til­la­tion of all of the knowl­edge on the open Internet, scraped by the fron­tier labs and dis­tilled into the mod­els that are them­selves be­ing dis­tilled? Who is ex­actly be­ing wronged here?

In fact, this para­dox is the so­lu­tion. I be­lieve that open weight mod­els are good for in­no­va­tion (and, per the above, I think that labs on the fron­tier will be fine), but it’s a prob­lem to be de­pen­dent on China. The U.S. should pass a law that (1) makes ex­plicit that col­lect­ing data for train­ing mod­els is fair use, and (2) bars terms of ser­vice that for­bid dis­til­la­tion, for U.S. com­pa­nies at a min­i­mum. Stopping dis­til­la­tion — which is lit­er­ally just query­ing the API — is nearly im­pos­si­ble; the U.S. should go the other way and lean into a new copy­right pol­icy that both in­dem­ni­fies the labs and also guar­an­tees that what they learned fu­els fur­ther in­no­va­tion for every­one else.

The Reason to Be Afraid

This en­tire Article has been an ex­er­cise in de­fus­ing over­re­ac­tion to Kimi K3 specif­i­cally and Chinese open weight mod­els gen­er­ally; how­ever, there is one rea­son to be con­cerned, and that is cy­ber­se­cu­rity. Consider this story from The Stack:

Hugging Face said its pro­duc­tion in­fra­struc­ture was breached by an autonomous” AI agent sys­tem early last week. The plat­for­m’s se­cu­rity team were ini­tially stymied in their in­ci­dent re­sponse (IR) by un­named US LLM fron­tier model guardrails which can­not dis­tin­guish an in­ci­dent re­spon­der from an at­tacker,” they said. So Hugging Face’s de­fend­ers turned in­stead to the open-source GLM 5.2 model from China’s Z.ai lab — run­ning it on their own in­fra­struc­ture to analyse the 17,000+ logs, or foot­prints, that the at­tack­ers left be­hind.

That’s a strik­ing pub­lic ad­mis­sion for the New York-headquartered Hugging Face, which lets users col­lab­o­rate on mod­els, datasets and ap­pli­ca­tions, and which this sum­mer hit the $100 mil­lion ARR mark. In an in­ci­dent re­port, the com­pany rec­om­mended that de­fend­ers have a ca­pa­ble model you can run on your own in­fra­struc­ture [our ital­ics] vet­ted and ready be­fore an in­ci­dent, both to avoid guardrail lock­out and to keep at­tacker data and cre­den­tials from leav­ing your en­vi­ron­ment.”

Hugging Face said its pro­duc­tion in­fra­struc­ture was breached by an autonomous” AI agent sys­tem early last week. The plat­for­m’s se­cu­rity team were ini­tially stymied in their in­ci­dent re­sponse (IR) by un­named US LLM fron­tier model guardrails which can­not dis­tin­guish an in­ci­dent re­spon­der from an at­tacker,” they said. So Hugging Face’s de­fend­ers turned in­stead to the open-source GLM 5.2 model from China’s Z.ai lab — run­ning it on their own in­fra­struc­ture to analyse the 17,000+ logs, or foot­prints, that the at­tack­ers left be­hind.

That’s a strik­ing pub­lic ad­mis­sion for the New York-headquartered Hugging Face, which lets users col­lab­o­rate on mod­els, datasets and ap­pli­ca­tions, and which this sum­mer hit the $100 mil­lion ARR mark. In an in­ci­dent re­port, the com­pany rec­om­mended that de­fend­ers have a ca­pa­ble model you can run on your own in­fra­struc­ture [our ital­ics] vet­ted and ready be­fore an in­ci­dent, both to avoid guardrail lock­out and to keep at­tacker data and cre­den­tials from leav­ing your en­vi­ron­ment.”

It’s dif­fi­cult to over­state how wrong-headed the Trump ad­min­is­tra­tion’s pan­icked re­sponse to Anthropic’s re­lease of Fable was, par­tic­u­larly since it ex­ac­er­bated Anthropic’s worst ten­den­cies in terms of as­sum­ing only they can be trusted with pow­er­ful AI. In a world with only one AI, it might make sense to re­serve the most pow­er­ful cy­ber­se­cu­rity ca­pa­bil­i­ties for the U.S. gov­ern­ment and trusted al­lies; how­ever, that’s not the world we live in.

There are and will be mod­els em­i­nently ca­pa­ble of mount­ing cy­ber­se­cu­rity at­tacks on ex­ist­ing in­fra­struc­ture, and those mod­els will be — al­ready are — widely avail­able. The best de­fense — the only vi­able de­fense, in fact — will be to make sure de­fend­ers have ac­cess to the best mod­els as well. Right now de­fend­ers are ef­fec­tively banned from us­ing Fable or Sol for cy­ber­se­cu­rity be­cause of Trump ad­min­is­tra­tion di­rec­tives; that means the best al­ter­na­tive is us­ing mod­els from a coun­try which has been try­ing to weaken our cy­ber de­fenses for years. This is in­sane!

The bet­ter course is clear: first, loosen Fable and Sol re­stric­tions on cy­ber­se­cu­rity, and sec­ond, en­sure that U.S. open weight model mak­ers are on an equal play­ing field with China. Yes, the fron­tier labs will kick and scream about this, but the Administration should re­al­ize that lis­ten­ing to their histri­on­ics has led the U.S. to a po­si­tion where U.S. com­pa­nies are de­pen­dent on China for their de­fenses. Let the fron­tier labs win by be­ing bet­ter; don’t let them de­fine safety or se­cu­rity, or pull up the lad­der of hu­man­i­ty’s col­lec­tive knowl­edge. China is al­ready hard enough to com­pete with; let­ting them carry the stan­dard for open­ness and in­no­va­tion is sim­ply giv­ing away our biggest ad­van­tage.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.