10 interesting stories served every morning and every evening.

Introducing Claude Opus 5

www.anthropic.com

Claude Opus 5 is avail­able to­day. It’s a thought­ful and proac­tive model that comes close to the fron­tier in­tel­li­gence of Claude Fable 5 at half the price.

On cod­ing and knowl­edge work eval­u­a­tions like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it re­mains be­hind Mythos 5 on cy­ber­se­cu­rity tasks.

Opus 5 is de­signed to be used every day: it works more ef­fi­ciently than other mod­els. It’s the new de­fault model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-ef­fec­tive­ness

Claude Opus 5 pro­vides greatly im­proved per­for­mance for the same cost as its pre­de­ces­sor, Opus 4.8. The charts in this sec­tion show how per­for­mance changes ac­cord­ing to the mod­el’s ef­fort set­ting, which cus­tomers can use to op­ti­mize for in­tel­li­gence or con­serve to­kens for faster and cheaper re­sults.

Opus 5 ex­cels on valu­able soft­ware en­gi­neer­ing tasks. For ex­am­ple, on Frontier-Bench v0.1, Opus 5 sur­passes all other mod­els, and more than dou­bles Opus 4.8’s per­for­mance at a lower cost per task. On CursorBench 3.2, at max ef­fort, the model per­forms within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater per­for­mance at a given cost than all other mod­els on high, xhigh, and max ef­fort.

We see sim­i­lar re­sults on knowl­edge work and prob­lem-solv­ing tasks. For ex­am­ple:

On ARC-AGI 3, an eval­u­a­tion where the model has to solve novel prob­lems, Opus 5’s score is three times as high as the next-best model.

On Zapier AutomationBench, which mea­sures whether mod­els can com­plete busi­ness tasks from start to fin­ish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its low­est ef­fort set­ting, Opus 5 passes more tasks than any other model.

On OSWorld 2.0, a com­puter use bench­mark, Opus 5 out­per­forms every other model at any given cost, sur­pass­ing Fable 5’s best re­sult at just over a third of the cost.

It’s also our best and most cost-ef­fi­cient model on sev­eral re­lated eval­u­a­tions:

Opus 5 is a mean­ing­ful im­prove­ment over Opus 4.8 for sci­en­tific re­search. It shows bet­ter per­for­mance than Opus 4.8 on every one of our life sci­ences eval­u­a­tions, which cover top­ics in­clud­ing struc­tural bi­ol­ogy, or­ganic chem­istry, and bioin­for­mat­ics. Its im­prove­ments are most no­table on or­ganic chem­istry tasks, like in­fer­ring mol­e­c­u­lar struc­tures from spec­troscopy data (it scores 10.2 per­cent­age points higher than Opus 4.8 on our in­ter­nal bench­mark), and on pro­tein-re­lated tasks like pre­dict­ing how vari­a­tions in a pro­tein’s se­quence af­fect how it func­tions (here, it scores 7.7 per­cent­age points higher).

Finally, Opus 5 is ca­pa­ble of pro­duc­ing much stronger vi­sual out­puts:

Working with Claude Opus 5

Claude Opus 5 is much stronger at ver­i­fy­ing its work and it­er­at­ing care­fully un­til it suc­ceeds. In eval­u­a­tions and early-ac­cess test­ing, we and our users found many ex­am­ples of Opus 5’s agency and thor­ough­ness:

On one Frontier-Bench task, Opus 5 was given a draw­ing of a ma­chine part and asked to write code to re­build it as a 3D FreeCAD model. However, in this task, the model was in­ten­tion­ally given no way to di­rectly view the draw­ing. Opus 5 re­sponded by writ­ing its own com­puter vi­sion pipeline to pull the geom­e­try from the raw pix­els, then re­con­structed the full ma­chine part. It suc­ceeded in do­ing so re­peat­edly; no com­pet­ing model with the same setup could solve it af­ter five at­tempts.

Given a real bug in a pop­u­lar open-source pack­age man­ager, Opus 5 found the root cause and fixed an edge case that the com­mu­ni­ty’s patch had missed. A com­pet­ing model fixed only the sur­face symp­tom (not the un­der­ly­ing cause), then re­ported the bug re­solved.

An en­gi­neer at a trad­ing firm used Opus 5 to build a mar­ket data feed for a new ex­change in a sin­gle ses­sion. Previous mod­els could not com­plete this task at all, even given ex­ten­sive plans from the en­gi­neer. Finding no live feed to val­i­date against, Opus 5 even built its own test har­ness to check that its code parsed the ex­change’s data cor­rectly.

Below are fur­ther re­ports from our early-ac­cess cus­tomers on their ex­pe­ri­ence of work­ing with Opus 5:

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 lets mon­i­tor­ing agents man­age parts of their own mem­ory in pro­duc­tion, mak­ing them more au­tonomous and re­li­able over longer hori­zons. The agent treats its con­text as a liv­ing doc­u­ment: af­ter flag­ging a po­ten­tial anom­aly in one of our ser­vices, it re-checked its own as­sump­tion against pro­duc­tion, found the sig­nal was be­nign, wrote the cor­rec­tion into its mem­ory, and re­tired its mon­i­tor­ing queries on its own.

Claude Opus 5 lets mon­i­tor­ing agents man­age parts of their own mem­ory in pro­duc­tion, mak­ing them more au­tonomous and re­li­able over longer hori­zons. The agent treats its con­text as a liv­ing doc­u­ment: af­ter flag­ging a po­ten­tial anom­aly in one of our ser­vices, it re-checked its own as­sump­tion against pro­duc­tion, found the sig­nal was be­nign, wrote the cor­rec­tion into its mem­ory, and re­tired its mon­i­tor­ing queries on its own.

Claude Opus 5 is a strong agen­tic cod­ing model built for long-run­ning, multi-step work. It deeply un­der­stands your code­base, holds the thread across com­plex tasks, and pins down re­quire­ments for fea­ture de­vel­op­ment and bug-fix­ing more ef­fec­tively than Opus 4.8. Developers can now build with Opus 5 in Kiro, ac­cess­ing its ad­vanced ca­pa­bil­i­ties to tackle am­bi­tious pro­jects.

Claude Opus 5 is a strong agen­tic cod­ing model built for long-run­ning, multi-step work. It deeply un­der­stands your code­base, holds the thread across com­plex tasks, and pins down re­quire­ments for fea­ture de­vel­op­ment and bug-fix­ing more ef­fec­tively than Opus 4.8. Developers can now build with Opus 5 in Kiro, ac­cess­ing its ad­vanced ca­pa­bil­i­ties to tackle am­bi­tious pro­jects.

01 /

24

Alignment and safety

Alignment. During pre-de­ploy­ment test­ing, our au­to­mated be­hav­ioral au­dit found Opus 5 to be our most aligned model to date (as shown in the graph be­low). It ad­heres to Claude’s Constitution bet­ter than Opus 4.8, Sonnet 5, or Fable 5; ex­hibits the low­est rates of de­cep­tive be­hav­ior; and is the least sus­cep­ti­ble to be­ing tricked into mis­use. It’s also our safest model yet in terms of avoid­ing reck­less ac­tions that could have hard-to-re­verse side ef­fects.

Safety. Opus 5 does not ad­vance the fron­tier in risky, dual-use ca­pa­bil­i­ties. In rig­or­ous eval­u­a­tions con­ducted along­side pri­vate-sec­tor and gov­ern­ment part­ners, we found it re­mains be­hind Mythos 5 in both bi­ol­ogy re­search and of­fen­sive cy­ber­se­cu­rity. More in­for­ma­tion about these eval­u­a­tions can be found in our System Card.

As with its pre­de­ces­sor, Opus 4.8, we’ve in­ten­tion­ally avoided train­ing Opus 5 on cy­ber tasks. The model has nev­er­the­less im­proved sub­stan­tially on these tasks as a re­sult of be­com­ing more gen­er­ally ca­pa­ble, and it comes close to Mythos 5 at find­ing cy­ber­se­cu­rity vul­ner­a­bil­i­ties. However, it re­mains sub­stan­tially be­hind Mythos 5 on the ex­ploita­tion of those vul­ner­a­bil­i­ties—that is, in turn­ing vul­ner­a­bil­i­ties into ma­te­r­ial cy­ber threats.

This is il­lus­trated by Opus 5’s per­for­mance on OSS-Fuzz, an eval­u­a­tion we’ve de­vel­oped to as­sess how well mod­els can find and then ex­ploit vul­ner­a­bil­i­ties with­out ex­ten­sive hu­man guid­ance. Although Mythos 5 and Opus 5 iden­tify vul­ner­a­bil­i­ties with sim­i­lar suc­cess, Opus 5’s score on the de­vel­op­ment of ex­ploits is far be­hind that of Mythos 5.

Safeguards for Opus 5

Claude Opus 5’s safe­guards are de­signed to al­low ben­e­fi­cial uses of the model in both cy­ber­se­cu­rity and bi­ol­ogy. They are sim­i­lar to those we ap­plied to Opus 4.8, with the ex­cep­tion of some stronger guardrails on a nar­row range of cy­ber tasks.

Cybersecurity. Opus 5’s cy­ber clas­si­fiers are pro­por­tion­ally less re­stric­tive than those on Fable 5. They al­low Opus 5 to find vul­ner­a­bil­i­ties in source code, but block binary-based” vul­ner­a­bil­ity scan­ning (a method more likely to be as­so­ci­ated with ma­li­cious ac­tors), pen­e­tra­tion test­ing, and ex­ploit gen­er­a­tion.

Based on our test­ing, we ex­pect the clas­si­fiers to in­ter­vene around 85% less of­ten than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged re­quests will fall back to Opus 4.8 by de­fault. Fallbacks to Opus 4.8 can also be en­abled on the API.

Our Cyber Verification Program (CVP) fa­cil­i­tates cy­ber­se­cu­rity work that would oth­er­wise be im­peded by the mod­el’s safe­guards. Enterprises and re­searchers who are al­ready part of the CVP have im­me­di­ate ac­cess to a ver­sion of Opus 5 with fewer se­cu­rity re­stric­tions.

Biology. Since Opus 5 has a sim­i­lar suite of safe­guards to Opus 4.8, it is now our most ca­pa­ble gen­er­ally avail­able model for sci­en­tific re­search. Nevertheless, the model still shows im­por­tant lim­i­ta­tions on long-run­ning, au­tonomous re­search tasks, which is where we ex­pect AI mod­els to pose the most sub­stan­tial bi­ol­ogy-re­lated risks. (Mythos 5 re­mains the stronger model for this type of bi­o­log­i­cal work.) As part of this launch, bi­ol­ogy-re­lated re­quests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.

Getting started

Claude Opus 5 is avail­able to­day on all plat­forms, priced at $5 per mil­lion in­put to­kens and $25 per mil­lion out­put to­kens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.

It’s also of­fered in Fast mode, where it runs around 2.5 times the de­fault speed. As with Opus 4.8, Fast mode is avail­able at twice Opus 5’s base price on the Claude Platform and through us­age cred­its in Claude Code.

Alongside Opus 5, we’re re­leas­ing two up­dates in beta:

Mid-conversation tool changes on the Claude Platform. Within a con­ver­sa­tion, de­vel­op­ers can now change which tools Claude can use with­out in­val­i­dat­ing the prompt cache.

Automatic fall­backs on the API. Users can now choose to have re­quests that are flagged by our safety clas­si­fiers on Opus 5 (or Fable 5) au­to­mat­i­cally route to an­other model. With au­to­matic fall­backs on, API re­quests al­ways route to the best avail­able model by de­fault rather than be­ing blocked.

Consistent with prior Opus mod­els, Opus 5 does not have data re­ten­tion re­quire­ments for gen­eral ac­cess.

For more guid­ance on how to get the best out of Opus 5, see our prompt­ing guide.

Footnotes

Frontier-Bench v0.1, Effort plot: These re­sults are from an in­ter­nal run of Frontier-Bench v0.1, on the mini-SWE-agent har­ness and a GKE back­end, mean re­ward over 5 at­tempts per task. Opus 4.8 served as fall­back on safety-clas­si­fier re­fusals for Opus 5 and Fable 5.

Related con­tent

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

Ask Claude about the Anthropic Economic Index

We’re launch­ing the Anthropic Economic Index con­nec­tor for Claude, which lets any­one ex­plore real data about AI and work.

Read more

Anthropic is do­nat­ing an­other $20 mil­lion to Public First Action

Anthropic is con­tribut­ing an ad­di­tional $20 mil­lion to Public First Action, bring­ing our to­tal sup­port to $40 mil­lion.

Read more

moonshotai/Kimi-K3 · Hugging Face

huggingface.co

📰  Tech Blog |     📄  Full Report

1. Model Introduction

Kimi K3 is an open-weight, na­tive mul­ti­modal agen­tic model and our most ca­pa­ble model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with na­tive vi­sion ca­pa­bil­i­ties and a 1-million-token con­text win­dow. It is the world’s first open 3T-class model, de­signed for fron­tier in­tel­li­gence across long-hori­zon cod­ing, knowl­edge work, and rea­son­ing.

Key Features

New Architecture: Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE spar­sity with a Stable LatentMoE frame­work that ac­ti­vates 16 out of 896 ex­perts — yield­ing an ap­prox­i­mate 2.5× im­prove­ment in over­all scal­ing ef­fi­ciency over Kimi K2.

Long-Horizon Coding: Operating with min­i­mal hu­man over­sight, Kimi K3 sus­tains long en­gi­neer­ing ses­sions, nav­i­gates mas­sive repos­i­to­ries, and or­ches­trates ter­mi­nal tools — from GPU ker­nel op­ti­miza­tion and com­piler de­vel­op­ment to vi­sion-in-the-loop game dev, CAD, and even chip de­sign.

Agentic Knowledge Work: Kimi K3 ad­vances end-to-end knowl­edge work, pro­duc­ing deep re­search with in­ter­ac­tive vi­su­al­iza­tions, wid­gets and dash­boards, and mo­tion de­sign and video edit­ing, pow­ered by its na­tive mul­ti­modal ar­chi­tec­ture.

Native Multimodality & Long Context: Kimi K3 un­der­stands text, im­ages, and video within the same model, and sup­ports a 1-million-token con­text win­dow.

Open Frontier Weights: We re­lease the full Kimi K3 model weights un­der the Kimi K3 License, mak­ing fron­tier in­tel­li­gence openly avail­able for re­search, de­ploy­ment, and fur­ther in­no­va­tion.

2. Model Summary

3. Evaluation Results

All Kimi K3 re­sults are ob­tained with rea­son­ing ef­fort set to max’ and tem­per­a­ture = 1.0. For sin­gle-step tasks, such as GPQA Diamond, HLE-Full, and vi­sion bench­marks with­out tools, we set top-p = 0.95; for agen­tic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, and ZeroBench, each cell re­ports the scores with­out and with tool aug­men­ta­tion (general tools for HLE-Full, Python for the vi­sion bench­marks), in that or­der.

Reasoning & knowl­edge bench­marks CritPt and AA-LCR. Scores are cited from Artificial Analysis as of July 23, 2026.

CritPt and AA-LCR. Scores are cited from Artificial Analysis as of July 23, 2026.

Coding bench­marks DeepSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is taken from the GLM-5.2 re­lease blog; all re­main­ing scores are from the of­fi­cial DeepSWE leader­board, un­der which Kimi K3 at­tains 67.3 with the mini-SWE-agent har­ness. We re­port the DeepSWE v1.1 tasks. Terminal-Bench 2.1. Kimi K3 is eval­u­ated with the Kimi Code har­ness. For all other mod­els, we re­port the best score across har­nesses: GLM-5.2 with Claude Code (GLM-5.2 re­lease blog); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 (Artificial Analysis); GPT-5.5 and GPT-5.6 Sol with Codex (OpenAI). ProgramBench. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog; all other scores are from Vals AI. SWE-Marathon. Kimi K3, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.6 Sol is eval­u­ated with the Codex har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog. Our eval­u­a­tion is based on an H20-calibrated branch of the of­fi­cial tasks as of July 9, 2026, prior to the fi­nal v1.1 re­lease: the Docker im­ages, per­for­mance gates, and ref­er­ence or­a­cles for the GPU tasks have been re­cal­i­brated for H20, while the cor­rect­ness and anti-cheat val­ida­tors re­main un­changed. Additionally, Claude Fable 5 hit fall­backs on 35% of the tasks in our eval­u­a­tion, which may have neg­a­tively im­pacted its mea­sured per­for­mance. FrontierSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness and GPT-5.6 Sol with the Codex har­ness; all other re­sults are from FrontierSWE. Dominance scores are re­com­puted from the raw scores us­ing the of­fi­cial eval­u­a­tion script and are cur­rent as of July 16, 2026. PostTrainBench. Scores for GLM-5.2, GPT-5.5, and Claude Opus 4.8 are adopted from the of­fi­cial PostTrainBench re­sults. Kimi K3, Claude Fable 5, and GPT-5.6 Sol are eval­u­ated with the of­fi­cial Harbor im­ple­men­ta­tion at max­i­mum rea­son­ing ef­fort, av­er­aged over three runs on H20 GPUs (instead of H100 in the of­fi­cial set­ting) — Kimi K3 and Claude Fable 5 with the Claude Code har­ness, and GPT-5.6 Sol with the Codex har­ness. MLS-Bench-Lite. Kimi K3 is eval­u­ated with the Kimi Code har­ness; GLM-5.2 and the Claude mod­els with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness. SciCode. Scores are cited from Artificial Analysis as of July 23, 2026. Kimi Code Bench 2.0 (in-house). Kimi K3 is eval­u­ated with the Kimi Code har­ness (it at­tains 73.7 with the Claude Code har­ness); GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness. All mod­els are eval­u­ated at max­i­mum rea­son­ing ef­fort, ex­cept GPT-5.5, which uses the xhigh” set­ting. As the bench­mark in­cludes cy­ber­se­cu­rity and safety-re­lated tasks, we also dis­close the frac­tion of re­fused or fall­back tasks: Claude Fable 5 hit 13 fall­backs and 1 re­fusal out of 80 tasks; 10 re­fusals out of 80 tasks en­tered GPT-5.6 Sol’s cy­ber guard; GPT-5.5 had 3 re­fusals out of 80 tasks.

DeepSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is taken from the GLM-5.2 re­lease blog; all re­main­ing scores are from the of­fi­cial DeepSWE leader­board, un­der which Kimi K3 at­tains 67.3 with the mini-SWE-agent har­ness. We re­port the DeepSWE v1.1 tasks.

Terminal-Bench 2.1. Kimi K3 is eval­u­ated with the Kimi Code har­ness. For all other mod­els, we re­port the best score across har­nesses: GLM-5.2 with Claude Code (GLM-5.2 re­lease blog); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 (Artificial Analysis); GPT-5.5 and GPT-5.6 Sol with Codex (OpenAI).

ProgramBench. Kimi K3 is eval­u­ated with the Kimi Code har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog; all other scores are from Vals AI.

SWE-Marathon. Kimi K3, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.6 Sol is eval­u­ated with the Codex har­ness. The GLM-5.2 score is from the GLM-5.2 re­lease blog. Our eval­u­a­tion is based on an H20-calibrated branch of the of­fi­cial tasks as of July 9, 2026, prior to the fi­nal v1.1 re­lease: the Docker im­ages, per­for­mance gates, and ref­er­ence or­a­cles for the GPU tasks have been re­cal­i­brated for H20, while the cor­rect­ness and anti-cheat val­ida­tors re­main un­changed. Additionally, Claude Fable 5 hit fall­backs on 35% of the tasks in our eval­u­a­tion, which may have neg­a­tively im­pacted its mea­sured per­for­mance.

FrontierSWE. Kimi K3 is eval­u­ated with the Kimi Code har­ness and GPT-5.6 Sol with the Codex har­ness; all other re­sults are from FrontierSWE. Dominance scores are re­com­puted from the raw scores us­ing the of­fi­cial eval­u­a­tion script and are cur­rent as of July 16, 2026.

PostTrainBench. Scores for GLM-5.2, GPT-5.5, and Claude Opus 4.8 are adopted from the of­fi­cial PostTrainBench re­sults. Kimi K3, Claude Fable 5, and GPT-5.6 Sol are eval­u­ated with the of­fi­cial Harbor im­ple­men­ta­tion at max­i­mum rea­son­ing ef­fort, av­er­aged over three runs on H20 GPUs (instead of H100 in the of­fi­cial set­ting) — Kimi K3 and Claude Fable 5 with the Claude Code har­ness, and GPT-5.6 Sol with the Codex har­ness.

MLS-Bench-Lite. Kimi K3 is eval­u­ated with the Kimi Code har­ness; GLM-5.2 and the Claude mod­els with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness.

SciCode. Scores are cited from Artificial Analysis as of July 23, 2026.

Kimi Code Bench 2.0 (in-house). Kimi K3 is eval­u­ated with the Kimi Code har­ness (it at­tains 73.7 with the Claude Code har­ness); GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol with the Codex har­ness. All mod­els are eval­u­ated at max­i­mum rea­son­ing ef­fort, ex­cept GPT-5.5, which uses the xhigh” set­ting. As the bench­mark in­cludes cy­ber­se­cu­rity and safety-re­lated tasks, we also dis­close the frac­tion of re­fused or fall­back tasks: Claude Fable 5 hit 13 fall­backs and 1 re­fusal out of 80 tasks; 10 re­fusals out of 80 tasks en­tered GPT-5.6 Sol’s cy­ber guard; GPT-5.5 had 3 re­fusals out of 80 tasks.

Agentic bench­marks OfficeQA Pro. Each test case pro­vides the agent with the en­tire PDF cor­pus, with all PDFs ren­dered as im­ages and no ma­chine-read­able text avail­able. OfficeQA Pro and SpreadsheetBench 2. Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol are eval­u­ated with the Codex har­ness. MCP-Atlas. All mod­els are eval­u­ated on the 500-task pub­lic sub­set with a 100-turn limit, us­ing Gemini 3.1 Pro as the judge. AutomationBench. All mod­els are eval­u­ated on the 600-task pub­lic sub­set, fol­low­ing the of­fi­cial GitHub setup in all other re­spects. BrowseComp. We adopt a con­text-com­paction strat­egy trig­gered at 300K to­kens. When eval­u­ated with the full 1M-token con­text win­dow and no con­text man­age­ment, Kimi K3 achieves a score of 90.4. The re­sults of Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 are cited from Anthropic and OpenAI. GDPval-AA v2, AA-Briefcase, τ³-Bank­ing, Harvey Lab-AA, and APEX-Agents. Scores are cited from Artificial Analysis and the APEX-Agents leader­board as of July 23, 2026. For Harvey Lab-AA, we re­port the cri­te­rion pass rate. CorpFin v2, Finance Agent v2, and Legal Research Bench. Scores are cited from Vals AI. Agents’ Last Exam. Scores are cited from the of­fi­cial leader­board as of July 23, 2026; we re­port the leader­board’s pri­mary pass-rate met­ric. On the leader­board, each model is paired with a spe­cific har­ness: Kimi K3 with Kimi Code; GPT-5.6 Sol and GPT-5.5 with Codex; Claude Fable 5, Claude Opus 4.8, and GLM-5.2 with Claude Code. † The Claude Fable 5 en­try runs at xhigh ef­fort with 40% of tasks an­no­tated as down­graded.

OfficeQA Pro. Each test case pro­vides the agent with the en­tire PDF cor­pus, with all PDFs ren­dered as im­ages and no ma­chine-read­able text avail­able.

OfficeQA Pro and SpreadsheetBench 2. Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are eval­u­ated with the Claude Code har­ness; GPT-5.5 and GPT-5.6 Sol are eval­u­ated with the Codex har­ness.

MCP-Atlas. All mod­els are eval­u­ated on the 500-task pub­lic sub­set with a 100-turn limit, us­ing Gemini 3.1 Pro as the judge.

AutomationBench. All mod­els are eval­u­ated on the 600-task pub­lic sub­set, fol­low­ing the of­fi­cial GitHub setup in all other re­spects.

BrowseComp. We adopt a con­text-com­paction strat­egy trig­gered at 300K to­kens. When eval­u­ated with the full 1M-token con­text win­dow and no con­text man­age­ment, Kimi K3 achieves a score of 90.4. The re­sults of Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 are cited from Anthropic and OpenAI.

GDPval-AA v2, AA-Briefcase, τ³-Bank­ing, Harvey Lab-AA, and APEX-Agents. Scores are cited from Artificial Analysis and the APEX-Agents leader­board as of July 23, 2026. For Harvey Lab-AA, we re­port the cri­te­rion pass rate.

CorpFin v2, Finance Agent v2, and Legal Research Bench. Scores are cited from Vals AI.

Agents’ Last Exam. Scores are cited from the of­fi­cial leader­board as of July 23, 2026; we re­port the leader­board’s pri­mary pass-rate met­ric. On the leader­board, each model is paired with a spe­cific har­ness: Kimi K3 with Kimi Code; GPT-5.6 Sol and GPT-5.5 with Codex; Claude Fable 5, Claude Opus 4.8, and GLM-5.2 with Claude Code. † The Claude Fable 5 en­try runs at xhigh ef­fort with 40% of tasks an­no­tated as down­graded.

Multimodal bench­marks Except for ZeroBench, which fol­lows the of­fi­cial set­ting and is run five times, all mul­ti­modal scores are av­er­aged over three runs. MMMU-Pro is eval­u­ated fol­low­ing the of­fi­cial pro­to­col, pre­serv­ing the orig­i­nal in­put or­der and prepend­ing im­ages to the text in­put. PerceptionBench is an in-house bench­mark that fo­cuses on atomic vi­sual per­cep­tion ca­pa­bil­i­ties.

Except for ZeroBench, which fol­lows the of­fi­cial set­ting and is run five times, all mul­ti­modal scores are av­er­aged over three runs. MMMU-Pro is eval­u­ated fol­low­ing the of­fi­cial pro­to­col, pre­serv­ing the orig­i­nal in­put or­der and prepend­ing im­ages to the text in­put.

PerceptionBench is an in-house bench­mark that fo­cuses on atomic vi­sual per­cep­tion ca­pa­bil­i­ties.

4. Native MXFP4 Quantization

Kimi K3 ap­plies quan­ti­za­tion-aware train­ing from the SFT stage on­ward, us­ing MXFP4 weights with MXFP8 ac­ti­va­tions for broad hard­ware com­pat­i­bil­ity.

5. Deployment

You can ac­cess Kimi K3′s API on https://​plat­form.kimi.ai by se­lect­ing kimi-k3, and we pro­vide OpenAI/Anthropic-compatible API for you. Currently, Kimi K3 is rec­om­mended to run on the fol­low­ing in­fer­ence en­gines:

You can ac­cess Kimi K3′s API on https://​plat­form.kimi.ai by se­lect­ing kimi-k3, and we pro­vide OpenAI/Anthropic-compatible API for you. Currently, Kimi K3 is rec­om­mended to run on the fol­low­ing in­fer­ence en­gines:

vLLM — see recipes

SGLang — see cook­book

TokenSpeed — see recipes

6. Model Usage

Kimi K3 al­ways has think­ing en­abled, and will re­turn rea­son­ing_­con­tent. Thinking ef­fort is con­fig­ured with the top-level rea­son­ing_­ef­fort re­quest field, which sup­ports low”, high”, and max” (default max”).

Kimi K3 was trained in the pre­served think­ing his­tory mode. For multi-turn con­ver­sa­tions and tool calls, Kimi K3 re­quires the com­plete as­sis­tant mes­sage re­turned by the API to be passed back to mes­sages as-is — in­clud­ing rea­son­ing_­con­tent and tool_­calls, not just con­tent:

im­port ope­nai

def chat_with­_p­re­served_­think­ing(client: ope­nai.Ope­nAI, mod­el_­name: str): mes­sages = [ { role”: user”, content”: Tell me three ran­dom num­bers.” }, { role”: assistant”, reasoning_content”: I’ll start by list­ing five num­bers: 473, 921, 235, 215, 222, and I’ll tell you the first three.”, content”: 473, 921, 235″ }, { role”: user”, content”: What are the other two num­bers you have in mind?” } ]

re­sponse = client.chat.com­ple­tions.cre­ate( model=mod­el_­name, mes­sages=mes­sages, stream=False, max_­to­kens=4096, rea­son­ing_­ef­fort=“max”, ) # the as­sis­tant should men­tion 215 and 222 that ap­pear in the prior rea­son­ing con­tent print(f”re­sponse: {response.choices[0].message.reasoning}“) re­turn re­sponse.choices[0].mes­sage.con­tent

For full guides and ex­am­ples (vision in­put, struc­tured out­put, par­tial mode, tool choice, dy­namic tool load­ing, con­text caching), see the Kimi K3 Quickstart and Thinking Effort.

Coding Agent Framework

Kimi K3 works best with Kimi Code CLI as its agent frame­work. We warmly in­vite you to give it a try — run Kimi Code in your ter­mi­nal and se­lect Kimi K3 us­ing the /model com­mand. We hope you en­joy build­ing with Kimi K3, and we would love to hear your feed­back!

7. License

Both the code repos­i­tory and the model weights are re­leased un­der the Kimi K3 License.

8. Contact Us

If you have any ques­tions, please reach out at sup­port@moon­shot.ai.

Safetensors

Model tree for moon­shotai/​Kimi-K3

Spaces us­ing moon­shotai/​Kimi-K3 8

Collection in­clud­ing moon­shotai/​Kimi-K3

US prosecutors charge Atlanta man after GrapheneOS phone wipes itself during airport search

www.techspot.com

Serving tech en­thu­si­asts for over 25 years. TechSpot means tech analy­sis and ad­vice you can trust.

A hot potato: A fed­eral case in Atlanta is rais­ing ques­tions about a pri­vacy-fo­cused mo­bile op­er­at­ing sys­tem, with pros­e­cu­tors ar­gu­ing that its fea­tures were used to erase ev­i­dence. The US Department of Justice is at­tempt­ing to pros­e­cute Atlanta res­i­dent Sam Tunick un­der a fed­eral statute that makes it a crime to de­stroy prop­erty in an ef­fort to pre­vent it from be­ing seized.

The case cen­ters on Tunick’s use of GrapheneOS, an open-source op­er­at­ing sys­tem that works on Google Pixel phones and lets users en­ter a pass­code to wipe a de­vice clean.

Experts said the le­gal ap­proach is un­usual and may be the first time the law has been aimed at an op­er­at­ing sys­tem. It’s con­cern­ing — and sends the mes­sage that [GrapheneOS] is crim­i­nal by de­fault,” said Christophe Boutry, a cy­ber­se­cu­rity and sur­veil­lance ex­pert. Boutry and Bill Buddington, se­nior staff tech­nol­o­gist at the Electronic Frontier Foundation, both said they had not seen a sim­i­lar case.

The in­ci­dent be­gan at Hartsfield-Jackson Atlanta International Airport on January 24 of last year. Tunick had just re­turned from a trip to the Dominican Republic when he was stopped for ques­tion­ing. According to court tes­ti­mony, fed­eral agents had al­ready cir­cu­lated his name and photo in­ter­nally, say­ing he was un­der in­ves­ti­ga­tion for suspected ter­ror­ism ac­tiv­i­ties” be­cause of his al­leged as­so­ci­a­tion with the move­ment against Cop City.

Tunick was taken to a sec­ondary screen­ing room, where mul­ti­ple agents ques­tioned him. A mo­tion filed by his de­fense ar­gues the in­ter­ro­ga­tion fo­cused on child sex­ual abuse ma­te­r­ial as a pre­text for in­ves­ti­gat­ing his con­nec­tions to the protest move­ment. The mo­tion also states that Tunick asked four times to speak with a lawyer and was de­nied each time. According to the same fil­ing, agents did not pre­sent a war­rant or read him his rights.

Government at­tor­neys and agents pushed back dur­ing Monday’s hear­ing. They de­scribed the en­counter as a rou­tine air­port in­spec­tion. Larry Findley, a Customs and Border Protection of­fi­cer, said agents were looking for any­thing that’s pro­hib­ited.”

During the ques­tion­ing, agents re­peat­edly asked Tunick to un­lock his phone and warned they would seize it if he re­fused. When he fi­nally pro­vided a pass­code, the phone ap­peared to restart. The de­fense mo­tion states that the screen went blank, flashed sev­eral times, and the phone ap­peared to restart,” re­sult­ing in the loss of data.

The wipe is now cen­tral to the case. Prosecutors are treat­ing it as an in­ten­tional act to de­stroy ev­i­dence, while the de­fense ar­gues that the search vi­o­lated Tunick’s con­sti­tu­tional rights and that the ev­i­dence should be sup­pressed.

The case raises ques­tions about which con­sti­tu­tional rights ap­ply at US bor­ders, in­clud­ing in­ter­na­tional air­ports, where au­thor­i­ties have broader search pow­ers. A judge is not ex­pected to rule on the de­fense mo­tion un­til at least late October.

GrapheneOS is de­signed to im­prove pri­vacy and se­cu­rity on Pixel phones. Supporters say those tools are le­git­i­mate se­cu­rity pro­tec­tions, not ev­i­dence of crim­i­nal in­tent. Boutry pointed to France and Spain, where au­thor­i­ties have strug­gled to gain ac­cess to se­cured de­vices. He said au­thor­i­ties have treated the use of GrapheneOS it­self as sus­pi­cious. In Catalonia, Spain, po­lice have been pro­fil­ing peo­ple car­ry­ing Pixel phones, as­sum­ing they have GrapheneOS in­stalled and are drug deal­ers or gang mem­bers.

The main goal [of the op­er­at­ing sys­tem] is pro­tec­tion of pri­vacy,” Boutry said. They’re our phones and the state can’t tell us how to use them.”

The case is tied to on­go­ing op­po­si­tion to Cop City, a $109 mil­lion po­lice train­ing fa­cil­ity that opened last spring. The pro­ject has drawn op­po­si­tion from ac­tivists con­cerned about po­lice mil­i­ta­riza­tion and en­vi­ron­men­tal im­pacts. Law en­force­ment of­fi­cials have de­fended it as nec­es­sary for train­ing and re­cruit­ment.

Previous at­tempts to pros­e­cute pro­test­ers at the state level have foundered, while fed­eral au­thor­i­ties have more re­cently stepped in, in­clud­ing a sep­a­rate in­dict­ment an­nounced last month.

Our position on open-weights models

www.anthropic.com

A post by Dario Amodei, Anthropic CEO

Over the last few days there has been a lot of dis­cus­sion about open-weights mod­els, es­pe­cially those from China. Reports sug­gest that some US of­fi­cials are con­sid­er­ing ban­ning the use of Chinese open-weights mod­els by US com­pa­nies. In re­sponse, many tech com­pa­nies have signed a let­ter sup­port­ing open-weights mod­els, and some peo­ple have even ac­cused Anthropic of want­ing to ban open-weights mod­els as a means of pro­tect­ing our busi­ness. Anyone who has read my past writ­ing should know that I don’t re­gard such bans as a use­ful mea­sure, but let me state it clearly so that there is no doubt: Anthropic has never ad­vo­cated for a ban on open-weights mod­els.

Open-weights mod­els that don’t have dan­ger­ous ca­pa­bil­i­ties are a pub­lic good: they don’t cost any­thing be­sides the com­pute needed to run them, and they pro­vide value to busi­nesses, de­vel­op­ers, and re­searchers.

Protectionist bans would not ad­dress my most se­ri­ous na­tional se­cu­rity con­cerns. Specifically, I am wor­ried about two night­mare sce­nar­ios. I laid these out in my es­say The Adolescence of Technology six months ago1, and have held these po­si­tions con­sis­tently for many years:

My pri­mary con­cern is the risk that au­thor­i­tar­ian gov­ern­ments—not solely the Chinese Communist Party (CCP), al­though the CCP is clearly the most ca­pa­ble threat—build AI mod­els that are more pow­er­ful than those built by the US, and use them to achieve per­ma­nent mil­i­tary su­pe­ri­or­ity or per­pe­trate in­cred­i­bly deep re­pres­sion of their own peo­ple. This con­cern is widely shared within the US gov­ern­ment: Vice President Vance warned in Paris last year that authoritarian regimes have stolen and used AI to strengthen their mil­i­tary, in­tel­li­gence, and sur­veil­lance ca­pa­bil­i­ties,” and the Intelligence Community’s 2026 Annual Threat Assessment found that other global pow­ers’ ro­bust progress in AI is chal­leng­ing US eco­nomic com­pet­i­tive­ness and na­tional se­cu­rity ad­van­tages.” It is ir­rel­e­vant whether these mod­els are re­leased with open weights, and cer­tainly ir­rel­e­vant whether they are used by US busi­nesses. In fact, the most dan­ger­ous model may be one that is trained in se­cret and handed only to the People’s Liberation Army for use in drones and the Ministry of State Security for sur­veil­lance and re­pres­sion.

My sec­ondary con­cern is the risk that pow­er­ful AI mod­els may be mis­used to carry out cy­ber­at­tacks or bi­o­log­i­cal at­tacks, and may have se­ri­ous align­ment prob­lems. Open-weights mod­els—it does not mat­ter whether they come from China or any­where else—do po­ten­tially pre­sent a higher risk than closed mod­els, be­cause it is very dif­fi­cult to ap­ply guardrails to them or mon­i­tor their us­age, and once weights are re­leased they can­not be with­drawn2. But ban­ning the use of these mod­els by US busi­nesses does noth­ing to ad­dress this risk, be­cause bad ac­tors are un­likely to be le­git­i­mate US busi­nesses. It would pro­tect US AI com­pa­nies from com­pe­ti­tion, but that has never been my goal.

To ad­dress these con­cerns, I do sup­port the fol­low­ing three mea­sures, which I and Anthropic have con­sis­tently ad­vo­cated for:

We should not sell pow­er­ful chips or chip­mak­ing equip­ment to China, and we should crack down on the ram­pant smug­gling3 and workarounds used to ob­tain ac­cess to such chips. China has lim­ited do­mes­tic pro­duc­tion ca­pac­ity, and there­fore, due to the scal­ing laws, can­not build more pow­er­ful mod­els than the US with­out US chips. This is the most ef­fi­cient and di­rect way to block threat #1, and by ham­per­ing the train­ing of mod­els that are out of reach of US law, it also in­di­rectly helps with threat #2.

We should crack down on in­dus­trial-scale dis­til­la­tion op­er­a­tions. Distillation is a much more com­pute-ef­fi­cient process than train­ing mod­els from scratch. It al­lows China to build much bet­ter mod­els than its num­ber of chips would or­di­nar­ily en­able, and thus par­tially evade chip bans. Distillation does not al­low the CCP to ob­tain equiv­a­lent or su­pe­rior AI ca­pa­bil­i­ties to the US, but it can bring the Chinese fron­tier to within a few months of the US fron­tier. It is true that many of the com­pa­nies car­ry­ing out these op­er­a­tions re­lease open-weights mod­els—but the open weights are far less rel­e­vant than the fact that the op­er­a­tions are backed by an au­thor­i­tar­ian state seek­ing to over­take the US at the fron­tier. We should have pol­icy in­ter­ven­tions to de­ter this be­hav­ior. A blan­ket ban on open-weights mod­els is nei­ther the cor­rect rem­edy nor some­thing we have called for4.

All suf­fi­ciently ca­pa­ble mod­els, open and closed, should go through manda­tory safety test­ing. The best way to ad­dress threat #2 is to just di­rectly test mod­els for cy­ber, bi­o­log­i­cal, and align­ment risks be­fore re­lease. I think this idea is ac­tu­ally close to a con­sen­sus: I have been heart­ened both that the Trump ad­min­is­tra­tion has moved in this di­rec­tion in re­cent months, and by re­cent in­dus­try pro­pos­als that would ap­ply such test­ing to the most ca­pa­ble mod­els re­gard­less of their coun­try of ori­gin or whether they are open or closed (while ex­empt­ing less ca­pa­ble mod­els, such as those from star­tups and acad­e­mia, en­tirely). Whether open mod­els do or don’t pose an in­creased risk, and whether that risk can be mit­i­gated, is some­thing that should emerge from test­ing, rather than be de­cided in ad­vance—and there may be promis­ing meth­ods for im­prov­ing the safety of open-weights mod­els, in­clud­ing re­cent re­search from AE Studio and Anthropic on mod­u­lar train­ing strate­gies. Note that to be ef­fec­tive, test­ing would need to be global, which means even the CCP would need to be on board. I think this may ac­tu­ally be pos­si­ble: as I wrote in The Adolescence of Technology, lim­ited co­op­er­a­tion around pre­vent­ing AI bi­o­log­i­cal weapons may be pos­si­ble be­cause it is in China’s in­ter­est too.

This brings me to the open let­ter. I agree with much of it: open weights ex­pand ac­cess to the AI econ­omy, they strengthen com­pe­ti­tion at least for some use cases, and they give cus­tomers greater con­trol. Concerns about dis­til­la­tion should be ad­dressed through tar­geted le­gal and com­mer­cial frame­works—the same mea­sure I de­scribed above. But I don’t agree with the let­ter’s as­ser­tions that open-weights mod­els nec­es­sar­ily make it eas­ier to de­velop safe­guards or that broad ac­cess to ca­pa­bil­i­ties nec­es­sar­ily helps de­fend­ers more than at­tack­ers. It seems at least as likely to me that the op­po­site will be true. For ex­am­ple, I worry that bi­ol­ogy will have a strong at­tacker-de­fender asym­me­try, where suf­fi­ciently ca­pa­ble mod­els may be able to quickly weaponize pan­demic-level viruses with widely avail­able ma­te­ri­als, whereas de­fense against these agents is a multi-year op­er­a­tional task in the best case (as we saw with Operation Warp Speed)5. Questions like this should be em­pir­i­cally an­swered by rig­or­ous pre-re­lease test­ing, not as­sumed in ad­vance.

To sum­ma­rize my and Anthropic’s po­si­tion, we have not and are not ad­vo­cat­ing for a ban on open-weights mod­els as a cat­e­gory. We should in­stead fo­cus on keep­ing pow­er­ful chips out of au­thor­i­tar­ian hands, stop­ping in­dus­trial-scale dis­til­la­tion, and re­quir­ing safety test­ing of all suf­fi­ciently ca­pa­ble mod­els, open and closed.

*Edit 28 July: Updated to note that the cited re­search on mod­u­lar train­ing strate­gies was a col­lab­o­ra­tion be­tween Anthropic and AE Studio.

Related con­tent

Cognizant and Anthropic ex­pand their part­ner­ship to bring Claude to en­ter­prise clients

Read more

Introducing Claude Opus 5

Opus 5 is a step change im­prove­ment for the Opus tier pow­er­ing long-run­ning agents while de­liv­er­ing im­prove­ments in cod­ing and pro­fes­sional work.

Read more

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

Kill the Cookie Banner!

killthecookiebanner.eu

Stop the track­ing cir­cus.

Tired of mis­lead­ing cookie ban­ners? The EU Commission has fi­nally pro­posed a so­lu­tion: set your pri­vacy pref­er­ences in the browser once, and never see an­other ban­ner. Unfortunately, the track­ing in­dus­try is push­ing back — and so far, they’ve been suc­cess­ful. We need YOUR help to #KillTheCookieBanner!

Cookie ban­ners are made to trick you into waiv­ing your rights

You may think that EU pri­vacy law re­quires cookie ban­ners. But the law is clear: on­line track­ing is pro­hib­ited by de­fault.

Therefore, the track­ing in­dus­try needs you to waive your rights. That’s why they in­vented cookie ban­ners, which of­ten are de­lib­er­ately mis­lead­ing as well as an­noy­ing.

This re­sults in up to 90% of peo­ple say­ing YES — even though only around 3% ac­tu­ally want to be tracked on­line. This sys­tem is bro­ken by de­sign.

The so­lu­tion: au­to­mat­i­cally com­mu­ni­cate your pri­vacy pref­er­ence

In Autumn 2025, as part of a big­ger le­gal re­form*, the EU Commission fi­nally pro­posed a so­lu­tion to the cookie ban­ner prob­lem: au­to­mated sig­nals that would com­mu­ni­cate your pri­vacy pref­er­ences be­tween your de­vice and web­sites or apps. You could then choose whether you want to ac­cept, refuse, or limit track­ing.

This idea is nei­ther new nor com­plex. Your browser al­ready au­to­mat­i­cally sig­nals other pref­er­ences to web­sites, for ex­am­ple your pre­ferred lan­guage. In some US states, such sig­nals for pri­vacy pref­er­ences are al­ready legally sup­ported.

The track­ing lobby fights to keep the cookie ban­ner

This would be a sim­ple so­lu­tion. But the track­ing in­dus­try seems afraid that if you can ex­press your pref­er­ences ef­fi­ciently, it could re­sult in lower con­sent rates for track­ing.

Following lob­by­ing ef­forts spear­headed by Google and the track­ing in­dus­try, sev­eral Member States are now block­ing the EU Commission’s pro­posal to get rid of the cookie ban­ner.

But not only that: in­dus­try groups are also lob­by­ing the European Parliament to re­ject the pro­posal.

We need YOUR help!

That’s where you come in: the fight to kill the cookie ban­ner is far from over. The Member States and the European Parliament have not yet de­cided on their po­si­tion on the is­sue yet.

You can take ac­tion by con­tact­ing your rep­re­sen­ta­tive in the European Parliament or in your Member State and ex­press your frus­tra­tion.

*This pro­posal for pri­vacy sig­nals is part of an EU law re­form called the Digital Omnibus. Most other parts of this re­form are prob­lem­atic and would weaken peo­ple’s rights. We want to make clear that we do not sup­port these other as­pects of the pro­posed re­form.

*This pro­posal for pri­vacy sig­nals is part of an EU law re­form called the Digital Omnibus. Most other parts of this re­form are prob­lem­atic and would weaken peo­ple’s rights. We want to make clear that we do not sup­port these other as­pects of the pro­posed re­form.

Android May Soon Restrict On-Device ADB, Affecting Shizuku, libadb and Developers

kitsumed.github.io

Early Warning

Before we dive in, please note that this is not an of­fi­cial Google an­nounce­ment. Instead, this is based on a re­cent, on­go­ing fea­ture re­quest on Google IssueTracker, in which a com­ment by one of the core ADB main­tain­ers (Google em­ployee) talked about re­strict­ing On-Device ADB con­nec­tions to pro­tect from bad ac­tors”.

Before you stop what you are do­ing to head over to that IssueTracker thread, please read this care­fully:

If you plan to visit the is­sue tracker just to post low-qual­ity com­ments (such as Hey, don’t do this, I need Shizuku!”), com­plaints about mo­nop­oly, or in­sults, I highly rec­om­mend that you re­frain from do­ing so. Spamming the thread will only cause Google de­vel­op­ers to lock the is­sue, ig­nore valu­able com­mu­nity feed­back, or stop shar­ing pub­lic up­dates about this change en­tirely.

I think a change like this could ben­e­fit Google as it goes well along their new Sideloading changes, but I do not ac­tively be­lieve this is what is go­ing on here. There is a true, valid rea­son be­hind this, and I think two dif­fer­ent ap­proach can be taken. We will talk about it in this blog post.

How You Can Help:

If you have a unique use case: If you are di­rectly af­fected and can write a de­tailed, con­struc­tive mes­sage ex­plain­ing your work­flow, pro­vid­ing links, or of­fer­ing tech­ni­cal so­lu­tions/​com­pro­mises, please, by all means, share your feed­back in Google is­sue.

If your use case has al­ready been men­tioned: You do not need to re­peat it. Instead, sim­ply click the +1 but­ton in the top right cor­ner of the Google IssueTracker to let Google know you are af­fected, and tog­gle no­ti­fi­ca­tions to stay up­dated on the dis­cus­sion.

I am hes­i­tant to make this blog post, as I fear it may over­load the few de­vel­op­ers that work on ADB. I am un­sure if I should wait more and see what hap­pens or hold it longer and see what ap­proach they take. Waiting too long could also be bad… As of writ­ing this, I am un­sure when/​if this post will re­lease. I saw some re­cent up­dates on as­sign­ments which were given to the main guy who pre­vi­ously worked on ADB, so we will see what hap­pens.

Introduction

Hi! I’m Kitsumed, de­vel­oper of ShizuCallRecorder, a Shizuku based ap­pli­ca­tion. As you may have guessed by now, I would be af­fected by this change. Obviously, I would re­ally like it if they do not pro­ceed with it in a way that pre­vent loop­back con­nec­tions.

To talk briefly about my­self, I made ShizuCallRecorder to help with some of my own dis­abil­i­ties. I can get by with­out it, but it’s much eas­ier with it.

You could say I have a very unique use case, and I keep dis­cov­er­ing other un­usual ones, like this per­son on Reddit, who used my ap­pli­ca­tion to pre­serve the voice­mail of a de­ceased loved one.

Call record­ing on Android is a com­pli­cated topic. There are count­less user re­quests, an of­fi­cial at­tempt to add the fea­ture in Android 11 that was later can­celed, and many closed-source, pri­vacy-in­va­sive ap­pli­ca­tions that use work-arounds.

I used to hear that many users with dis­abil­i­ties had to trade their pri­vacy for an eas­ier daily life. I guess that’s what peo­ple mean when they talk about those trade-offs.

Don’t even get me started on OEMs that force an au­dio warn­ing such as This call is be­ing recorded,” when it’s in places where it’s not legally re­quired. People of­ten don’t re­act well to that, even if you ex­plain why you’re record­ing the call, it gives a bad im­pres­sion. To be hon­est, I prob­a­bly would­n’t re­act well to it ei­ther.

I’m sure most of you have a more power-user” use of Shizuku or even use loop­back ADB for de­vel­oper tasks. I do those too, but I wanted to point out one of my unique use case.

Alright, back on the main topic. Before I ex­plain what the pro­posed change is, I’m go­ing to ex­plain what is ADB for less tech­ni­cal users.

What’s ADB?

ADB, also called Android Debug Bridge, is pro­to­col cre­ated by Google to let de­vel­op­ers do de­vel­oper things on Android de­vices…

Basically, it grant us a high level of priv­i­leges, give us ac­cess to a lot of sen­si­ble com­mands to tests how the phone and ap­pli­ca­tion be­have. Useful stuff for any de­vel­op­ers or power-users.

ADB was orig­i­nally de­signed to work over a USB con­nec­tion, but later ex­panded how it could works:

USB: The orig­i­nal way. ADB com­mu­ni­cates di­rectly over a USB ca­ble.

TCP/IP: Introduced as a way to run ADB over a net­work us­ing an IP ad­dress and port (typically port 5555). The con­nec­tion car­ries ADB traf­fic in plain text and pro­vides a YES/NO prompt as au­then­ti­ca­tion. Can only be en­abled once you al­ready have an ac­tive ADB con­nec­tion.

Wireless Debugging (Wifi 1.0/2.0): Introduced in Android 11, it aims to im­prove the legacy TCP/IP work­flow. It re­quires pair­ing the com­puter with the de­vice us­ing a pair­ing code or QR code, then es­tab­lishes an au­then­ti­cated and en­crypted con­nec­tion for sub­se­quent ADB ses­sions. It does not re­quire an ac­tive ADB con­nec­tion to be en­abled.

What’s a On-Device ADB con­nec­tion?

ADB was orig­i­nally in­tended to be used with two de­vices (simplified ex­pla­na­tion): the Android de­vice be­ing de­bugged run the ADB Daemon (ADBD) and a sep­a­rate de­vel­oper ma­chine run the ADB client. In prac­tice, how­ever, this setup is not al­ways con­ve­nient. Some de­vel­op­ers work di­rectly from their Android de­vice and do not have ac­cess to a sec­ond ma­chine.

This led to us­age of On-Device ADB (this is not a of­fi­cial term). By us­ing a ter­mi­nal em­u­la­tor such as Termux, de­vel­op­ers can run an ADB client di­rectly on their phone and es­tab­lish a con­nec­tion to the lo­cal dae­mon (ADBD) server us­ing ei­ther ADB TCP/IP or Wireless Debugging. Since both the client and server are run­ning on the same de­vice, the con­nec­tion is made through the loop­back ad­dress (127.0.0.1). That’s what I call On-Device ADB.

While this is a niche use case com­pared to how ADB in­tended to be used, it has led to the cre­ation of pro­jects such as libadb-an­droid by MuntashirAkon and Shizuku by RikkaApps. These pro­jects, along with many oth­ers, have cre­ated a large open-source com­mu­nity that made a wide range of tools for de­vel­op­ers and power-users.

The pro­posed change

A new fea­ture was made on Google IssuerTracker to al­low de­vel­op­ers to choose what in­ter­face ADBD (ADB server dae­mon) would lis­ten to.

This fea­ture was pro­posed fol­low­ing a ma­jor se­cu­rity is­sue iden­ti­fied as CVE-2026 – 0073, which al­lowed the Wireless ADB au­then­ti­ca­tion process to be fully by­passed. What is pro­posed in this is­sue is ac­tu­ally a nice idea.

Right now, ADBD makes it­self avail­able on every net­work your phone is con­nected to. This fea­ture re­quest asks to let de­vel­op­ers choose which in­ter­face is cho­sen, re­duc­ing the ex­po­sure.

The prob­lem

The prob­lem lies in the re­sponse from one of ADB core main­tain­ers:

sa…@google.com:

Connection to lo­cal­host has also been the source of ex­ploit where app are us­ing that socket to adbd to es­ca­late their priv­i­leges.What about we re­strict to al­ways only bind­ing to wifi in­ter­face wlan0 ?

Connection to lo­cal­host has also been the source of ex­ploit where app are us­ing that socket to adbd to es­ca­late their priv­i­leges.

What about we re­strict to al­ways only bind­ing to wifi in­ter­face wlan0 ?

Here, we can see that the em­ployee talk about only al­low­ing wlan0, the in­ter­face of the Wifi con­nec­tion. Doing so would break many things, On-Device ADB, ADB via VPN, ADB via Ethernet, and many other unique de­vel­op­ers se­tups.

Another is­sue I see here is the stance they seems to cur­rently have on On-Device ADB. Their com­ment sug­gests that On-Device ADB is viewed pri­mar­ily as an ex­ploit bad ac­tors can use to el­e­vate priv­i­leges, yet there a lot of le­git­i­mate us­ages of on-de­vice ADB. Developers them­self uses it when they can’t ac­cess a com­puter.

While it can in­deed be used to el­e­vate priv­i­leges, a malicious” ap­pli­ca­tion can­not do so alone. It re­quires mul­ti­ples ac­tions that MUST be per­fomed by a hu­man.

Why On-Device ADB Is Not Really Used by Bad Actors

A ma­li­cious ap­pli­ca­tion could use an on-de­vice ADB con­nec­tion to per­form priv­i­lege es­ca­la­tion. However, it can­not es­tab­lish one by it­self.

To il­lus­trate this, I’ve laid out the gen­eral lim­i­ta­tions a bad ac­tor would face in a few sce­nar­ios, sim­i­lar to my com­ment on IssueTracker.

Scenario 1: General Android Users

You in­stall a ma­li­cious ap­pli­ca­tion.ADB is dis­abled. ADBD is not run­ning, and the ap­pli­ca­tion does not have the WRITE_SECURE_SETTINGS per­mis­sion, as it must be man­u­ally granted via ADB. No ex­ploita­tion at­tempts are pos­si­ble.

ADB is dis­abled. ADBD is not run­ning, and the ap­pli­ca­tion does not have the WRITE_SECURE_SETTINGS per­mis­sion, as it must be man­u­ally granted via ADB. No ex­ploita­tion at­tempts are pos­si­ble.

Scenario 2: A Developer on Android 11+ Using Wireless ADB

You in­stall a ma­li­cious ap­pli­ca­tion.

You en­able USB de­bug­ging, start­ing ADBD.

You en­able Wireless ADB. ADBD is now run­ning on your phone and lis­ten­ing on all net­work in­ter­faces.

The ap­pli­ca­tion would need to per­form the one-time pair­ing process. This re­quires the user to man­u­ally re­trieve the one-time code from Settings and pro­vide it. No ex­ploita­tion at­tempts are pos­si­ble.

Scenario 3: A Developer Using ADB over TCP/IP

You in­stall a ma­li­cious ap­pli­ca­tion.

You en­able USB de­bug­ging, start­ing ADBD.

You con­nect via USB ADB to en­able TCP/IP, then dis­con­nect the USB ca­ble.

ADBD con­tin­ues run­ning and lis­tens on all net­work in­ter­faces.

The ap­pli­ca­tion ini­ti­ates a con­nec­tion, caus­ing an au­tho­riza­tion prompt to ap­pear on the screen. If the user se­lects No, the con­nec­tion is re­jected. No silent ex­ploita­tion at­tempts are pos­si­ble.

Conclusion

In a nor­mal sce­nario, a bad ac­tor can­not gain an ADB con­nec­tion. It is only pos­si­ble while the de­vel­oper is ac­tively us­ing ADB on the de­vice, be­cause a bad ac­tor can­not start ADBD by them­selves.

However, if we re­turn to a sce­nario such as CVE-2026 – 0073, ex­ploita­tion would be­come pos­si­ble in Scenarios 2 and 3, but only af­ter the user has man­u­ally en­abled USB de­bug­ging. As for sce­nario 3, it would also re­quire the de­vel­oper to MANUALLY en­able TCP/IP. I can un­der­stand the mo­ti­va­tion for pre­vent­ing loop­back con­nec­tions by de­fault, but not to fully pre­vents it.

There is a dif­fer­ence be­tween pre­vent­ing it by de­fault and per­ma­nently pre­vent­ing it. I think this should be some­thing users can dis­able through a per­sis­tent set­ting. By this, I mean a tog­gle that sur­vives a re­boot (else it would make tools like Shizuku im­prac­ti­cals) and, ide­ally, can­not be read by third-party ap­pli­ca­tions. Otherwise, de­vel­op­ers would have to re­peat­edly dis­able it when­ever bank­ing apps or games de­tect that on-de­vice ADB is avail­abl. That said, this part could be worked around once a ap­pli­ca­tion is man­u­ally granted WRITE_SECURE_SETTINGS.

I think that jus­ti­fy­ing block­ing it be­cause a hu­man could per­form ac­tions to al­low it would be far-fetched. A hu­man could also des­ig­nate a ma­li­cious ap­pli­ca­tion as a de­vice ad­min­is­tra­tor or grant it Accessibility per­mis­sions, and yet we would not get rid of those fea­tures.

I be­lieve it should be an as­sumed risk that dis­abling such a se­cu­rity fea­ture en­ables on-de­vice de­bug­ging while also ex­pos­ing the de­vice to the rare pos­si­bil­ity of a fu­ture vul­ner­a­bil­ity. To me, the fea­ture/​risk ra­tio make sense here, as there is a real, le­git­i­mate us­age.

Although it may have not been orig­i­nally in­tended, On-Device ADB has en­abled a niche ecosys­tem of de­vel­oper and power-user tools, in­clud­ing pro­jects such as App Manager, libadb-an­droid, Canta, aShell, ShizuWall, ShizuCallRecorder and Shizuku.

To con­clude, I would like to men­tion my con­clu­sion in What Is Shizuku? How Does It Work? Security Implications?, where I said as a joke:

Here’s my clos­ing thought: I can’t wait to see how they’ll jus­tify dis­abling loop­back ADB over TCP/IP con­nec­tions on non-em­u­la­tor de­vices and start re­quir­ing a Google ac­count…

Here’s my clos­ing thought: I can’t wait to see how they’ll jus­tify dis­abling loop­back ADB over TCP/IP con­nec­tions on non-em­u­la­tor de­vices and start re­quir­ing a Google ac­count…

Well I guess they may in­deed dis­able loop­back ADB. Huh… I was al­most 100% right on that one I guess. If you are a tech­ni­cal user, I in­vite you to give your feed­back in­side the is­sue here. Please don’t for­get the early warn­ing.

2026 – 07-26 UPDATE

Hi, quick? up­date:

As this is gain­ing trac­tion, I’ve seen a few users men­tion that bad ac­tors could au­to­mate some ac­tions once the user is tricked into giv­ing the ac­ces­si­bil­ity per­mis­sions.

This is true to­day for side­load­ing (not for apps on the Play Store due to Google rules and re­stric­tions). However, the pro­posed so­lu­tion would also ad­dress this is­sue. So if Google comes out and uses this as a jus­ti­fi­ca­tion, it’s in my own eyes most likely just an ex­cuse rather than a real lim­i­ta­tion that force them” to do it.

They could still re­strict this be­hav­ior by de­fault, as pro­posed in my blog post, while al­low­ing de­vel­op­ers to dis­able that re­stric­tion us­ing a sec­ond com­puter, at this point, it be­comes as­sumed risk that you con­nected via ADB via USB, ran the com­mands, and dis­abled the re­stric­tion. There is no more accessibility per­mit ex­ploita­tion” as it would by de­fault be blocked.

I also still be­lieve, as I men­tioned pre­vi­ously, that there needs to be a rea­son­able bal­ance be­tween func­tion­al­ity and se­cu­rity. Just be­cause a user could grant a ma­li­cious app the abil­ity to do some­thing harm­ful does­n’t mean we should block the func­tion­al­ity en­tirely.

If we fol­low that logic, no one should be al­lowed to choose their own key­board app be­cause it could be a key­log­ger. Accessibility ser­vices would be too dan­ger­ous for us be­cause they can view and in­ter­act with the screen. VPN apps could in­ter­cept or mon­i­tor all net­work traf­fic. Device Administrator ca­pa­bil­i­ties could be abused to lock a user out of their de­vice or make an app dif­fi­cult to re­move. Apps with the Draw over other apps” per­mis­sion could be abused for tap­jack­ing or phish­ing at­tacks by dis­play­ing de­cep­tive over­lays. The same rea­son­ing could be ap­plied to many other pow­er­ful Android per­mis­sions and ca­pa­bil­i­ties. At some point, there needs to be a bal­ance be­tween mit­i­gat­ing risks and pre­serv­ing le­git­i­mate func­tion­al­ity.

Granting ac­ces­si­bil­ity per­mis­sions is al­ready a sig­nif­i­cant hur­dle (8 to 12+ steps), and every se­cu­rity prompt shown dur­ing the process is very ex­plicit that do­ing so is dan­ger­ous. If some­one chooses to pro­ceed de­spite mul­ti­ple clear warn­ings ex­plain­ing the risks, even if they are bee­ing scammed, they likely would have fallen to many other types of scams. Justifying a com­plete lock­down is, in my opin­ion, not ap­pro­pri­ate when there are al­ter­na­tive so­lu­tions (as men­tioned in the blog) that ef­fec­tively ad­dress the same con­cerns while not im­ple­ment­ing break­ing changes.

PGSimCity · How PostgreSQL Works, in 3D

nikolays.github.io

PGSimCity is an in­de­pen­dent, non-com­mer­cial ed­u­ca­tional vi­su­al­iza­tion of PostgreSQL in­ter­nals. It is not af­fil­i­ated with, spon­sored, en­dorsed, or ap­proved by Electronic Arts Inc. SimCity is a trade­mark of Electronic Arts Inc.

A work­ing model of the PostgreSQL en­gine

load­ing the city code…

Early, un­re­viewed pro­to­type. It al­most cer­tainly con­tains in­ac­cu­ra­cies in both the model and ex­pla­na­tions. Found one? Open an is­sue or send a pull re­quest.

GitHub - drumih/turbo-fieldfare: Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

github.com

TurboFieldfare

Gemma 4 26B-A4B in­fer­ence in about 2 GB of RAM A cus­tom Swift + Metal run­time for any Apple Silicon Mac, even the 8 GB ones.

Quick start · Local server · Benchmarks · Contribute re­sults · How it works · Experiments · References

Memory got ex­pen­sive. So I gave a 26-billion-parameter model a ~2 GB bud­get.

TurboFieldfare runs the in­struc­tion-tuned Gemma 4 26B-A4B with­out load­ing the en­tire 14.3 GB model into mem­ory. It keeps the shared 1.35 GB core and FP16 KV cache in mem­ory, then streams only the ex­perts needed for each to­ken from SSD. This is what lets the model run on Macs with 8 GB of RAM.

The run­time, stream­ing in­staller, CLI, and na­tive Mac app are writ­ten in Swift and Metal. TurboFieldfare is model-spe­cific rather than a wrap­per around MLX or llama.cpp. The cu­rated ex­per­i­ment record sum­ma­rizes 103 mea­sured re­sults across ker­nels, caching, I/O, pre­fill, and de­code.

Try it

git clone https://​github.com/​dru­mih/​turbo-field­fare.git cd turbo-field­fare swift build -c re­lease .build/release/TurboFieldfareMac

On the first run, Swift Package Manager down­loads and builds the Swift pack­ages re­quired by the to­k­enizer. The com­plete re­lease build in­cludes the fore­ground Mac app and its sib­ling de­code-ser­vice ex­e­cutable.

When the app opens, choose Download and let TurboFieldfare fetch and repack the pinned model (about 15 GB). Once it is ready, choose Load Model, type your prompt, and press Generate.

At a glance

The mea­sured re­sult is a ref­er­ence point, not a per­for­mance ceil­ing. Prompt length, gen­er­ated length, page-cache state, and hard­ware all af­fect through­put. To help mea­sure an­other Apple Silicon Mac, fol­low the com­mu­nity bench­mark guide.

Using TurboFieldfare

TurboFieldfare pro­vides a na­tive Mac app, a com­mand-line in­ter­face, and an ex­per­i­men­tal loop­back OpenAI-compatible server. They use the same .gturbo model di­rec­tory, but only one model-own­ing prod­uct should run at a time.

The Swift pack­age ex­poses six prod­ucts:

Requirements

An Apple Silicon Mac; the val­i­dated tar­get is an 8 GB M2 MacBook Air

ma­cOS 26 with Metal 4

Xcode 26 and Swift 6.2 or newer

Enough free stor­age for the ~14.3 GB model in­stal­la­tion

An in­ter­net con­nec­tion for the first model in­stall

The pack­age is ar­m64-only. Older ma­cOS and Metal ver­sions are not sup­ported.

Prompting the model

The Mac app treats what you type as an in­struc­tion and han­dles Gemma’s chat for­mat­ting au­to­mat­i­cally. Just de­scribe the task and in­clude any con­text the model needs.

Generation de­faults to tem­per­a­ture 0.2, Top-K 64, and Top-P 0.95. Set tem­per­a­ture to 0 for de­ter­min­is­tic greedy out­put. The model can still re­peat it­self or give in­cor­rect an­swers, so check im­por­tant re­sults.

TurboFieldfare is text-only. The app and CLI sup­port user and model mes­sages plus op­tional sys­tem guid­ance; they do not ex­pose or ex­e­cute tools. The loop­back server ac­cepts func­tion-tool de­c­la­ra­tions and re­turns model-pro­duced tool calls for the client to au­tho­rize and ex­e­cute. Images, au­dio, and video are not sup­ported.

Mac app

Clone the repos­i­tory, then run the app from its root:

swift build -c re­lease .build/release/TurboFieldfareMac

Build the com­plete pack­age so the app and its sib­ling de­code ser­vice are both avail­able. When launched from this check­out, the app stores the model in scratch/​gem­ma4.gturbo.

Install the model

On first launch, the app checks the avail­able stor­age and shows the down­load and in­stalled sizes. Choose Download to be­gin.

The in­staller never ma­te­ri­al­izes the full source check­point. It streams the re­quired byte ranges from the pinned Hugging Face re­vi­sion and repacks them di­rectly into the .gturbo lay­out as they ar­rive. This avoids a sec­ond full check­point on disk and keeps scratch mem­ory bounded.

The first in­stal­la­tion trans­fers about 15 GB through bounded Hugging Face range re­quests. Network speed and Hugging Face re­sponse times vary, so it can take a while. The com­pleted .gturbo in­stal­la­tion oc­cu­pies about 14.3 GB and is ac­cepted only af­ter its man­i­fest and file hashes have been val­i­dated. Installation does not load the model into mem­ory.

Load and gen­er­ate

After in­stal­la­tion:

Choose Load Model.

Enter a prompt in the com­poser.

Choose Generate, or press Command+Return.

Use the stop but­ton or Escape to end gen­er­a­tion early.

The sta­tus bar shows gen­er­a­tion progress, de­code speed, and mem­ory use. Use the right pane to con­fig­ure sam­pling, con­text length, ex­pert-cache slots, and run­time op­tions. See Runtime con­trols for de­tails and de­faults.

Command-line in­ter­face

The CLI uses an ex­ist­ing .gturbo in­stal­la­tion. If you in­stalled the model through the Mac app, it is al­ready avail­able at scratch/​gem­ma4.gturbo. Otherwise, in­stall it from the com­mand line:

swift run -c re­lease TurboFieldfareRepack \ –output scratch/​gem­ma4.gturbo \ –overwrite

Continue a can­celled or in­ter­rupted down­load:

swift run -c re­lease TurboFieldfareRepack \ –output scratch/​gem­ma4.gturbo \ –overwrite \ –resume

Remove saved down­load state:

swift run -c re­lease TurboFieldfareRepack \ –discard-partial \ –output scratch/​gem­ma4.gturbo

The run­time ac­cepts only a com­pleted .gturbo di­rec­tory with a fi­nal man­i­fest.json.

Verify an ex­ist­ing in­stal­la­tion with­out load­ing the model:

swift run -c re­lease TurboFieldfareRepack \ –verify-install \ –input-gturbo scratch/​gem­ma4.gturbo

Instruction chat

Put chat mes­sages in a JSON ar­ray and pass it with –messages-file:

[ {“role”: user”, content”: Explain why chun­ked pre­fill re­duces time to first to­ken while keep­ing mem­ory bounded.“} ]

swift run -c re­lease TurboFieldfareCLI \ –model scratch/​gem­ma4.gturbo \ –messages-file mes­sages.json

This for­mats mes­sages in the same way as the Mac app. The CLI re­sponse limit is set with –max-new, which de­faults to 1,024 to­kens. The Mac app can gen­er­ate un­til the se­lected con­text win­dow is full.

Raw com­ple­tion

–prompt is avail­able for raw com­ple­tion and re­pro­ducible com­par­isons. It passes the text di­rectly to the model with­out chat for­mat­ting. Use –messages-file for in­struc­tion-re­sponse con­ver­sa­tions.

swift run -c re­lease TurboFieldfareCLI \ –model scratch/​gem­ma4.gturbo \ –prompt The cap­i­tal of France is” \ –max-new 64 \ –temperature 0

This ex­am­ple de­lib­er­ately re­quests a short greedy com­ple­tion.

Common gen­er­a­tion op­tions in­clude –max-context, –temperature, –top-k, –top-p, –repetition-penalty, –seed, and re­peat­able –stop strings. The pub­lic CLI uses pro­duc­tion run­time de­faults. Run the fol­low­ing com­mand for the com­plete op­tion list:

swift run -c re­lease TurboFieldfareCLI –help

Generated text goes to stan­dard out­put. Timing sta­tis­tics go to stan­dard er­ror; add –quiet to sup­press that footer in scripts.

Local OpenAI-compatible server

Build the server and point it at an in­stalled model:

swift build -c re­lease –product TurboFieldfareServer .build/release/TurboFieldfareServer \ –model scratch/​gem­ma4.gturbo

It lis­tens on http://​127.0.0.1:8080/​v1 and sup­ports Chat Completions, stream­ing, func­tion tools, and sin­gle-pre­fix prompt reuse. The client must au­tho­rize and run every tool call. Keep the server on loop­back; it has no re­mote au­then­ti­ca­tion or TLS.

See Local server for a test re­quest, Python and OpenCode setup, prompt reuse, tool han­dling, and the sup­ported API sub­set.

Test and con­tribute

Run the pub­lic test suite se­ri­ally:

Scripts/test.sh

Before start­ing a model run, close mem­ory-heavy apps and check mem­o­ry_­pres­sure -Q. If it re­ports lit­tle free mem­ory, post­pone the run. Run only one TurboFieldfare app, de­code ser­vice, CLI, server, test, or other lo­cal-model process at a time.

To con­tribute a com­pa­ra­ble per­for­mance re­sult, fol­low the com­mu­nity bench­mark guide.

How the in­fer­ence en­gine works

At each trans­former layer, Metal com­putes at­ten­tion and the router from res­i­dent weights. The CPU uses the router’s top-8 ex­pert IDs to plan against the lay­er’s 16-slot LFU cache, then fills misses with bounded par­al­lel pread calls into Metal-visible buffers. Metal com­putes the res­i­dent shared-ex­pert branch while those reads run, then com­bines the shared and routed out­puts.

Prompt pre­fill uses chunks of up to 128 to­kens so one fetched ex­pert can serve mul­ti­ple rows. Generation re­peats the routed layer loop one to­ken at a time. The in­staller ap­plies the same bounded-mem­ory rule: it repacks re­mote ranges di­rectly into .gturbo with­out stag­ing a full shard or ten­sor.

For a vi­sual in­tro­duc­tion to the model ar­chi­tec­ture, see Maarten Grootendorst’s A Visual Guide to Gemma 4.

System de­sign ex­plains the .gturbo lay­out, mem­ory own­er­ship, pre­fill, router hand­off, cb1/​io/​cb2 phases, Metal ker­nels, and cor­rect­ness in­vari­ants.

Status and scope

TurboFieldfare cur­rently in­cludes:

Remote stream­ing repack into the .gturbo model for­mat

Instruction-tuned Gemma 4 26B-A4B with ver­i­fied text-only chat for­mat­ting

4-bit MLX affine em­bed­ding, at­ten­tion, shared-ex­pert, and routed-ex­pert weights, with an 8-bit router

Custom Metal ker­nels for quan­tized GEMV, at­ten­tion, MoE, nor­mal­iza­tion, RoPE, sam­pling, and pro­duc­tion fu­sions

SSD-backed routed-ex­pert stream­ing with a bounded ex­pert cache

Chunked sin­gle-prompt pre­fill and to­ken-by-to­ken gen­er­a­tion

FP16 KV stor­age with bounded cir­cu­lar stor­age for 25 slid­ing-win­dow lay­ers and lin­ear stor­age for 5 full-at­ten­tion lay­ers

Exact split-K/​V de­code at­ten­tion with dis­tinct nor­mal­ized K and V paths

A Swift li­brary, stream­ing in­staller, com­mand-line in­ter­face, loop­back OpenAI-compatible server, and na­tive SwiftUI/AppKit Mac app with a one-shot lo­cal de­code ser­vice

Current scope is text-only in­fer­ence from the pinned Gemma 4 26B-A4B in­struc­tion check­point on Apple Silicon Macs with at least 8 GB of RAM.

Future work

Build iPhone and iPad apps, then mea­sure in­fer­ence speed and mem­ory use on mo­bile hard­ware.

Benchmark more Apple Silicon Macs, es­pe­cially the base 16 GB M4 Mac mini and other 8 GB mod­els.

Experiments and tech­ni­cal doc­u­men­ta­tion

The ex­per­i­ments that shaped TurboFieldfare ex­plain the largest wins, the plau­si­ble ideas that failed, and the early re­sults that re­versed un­der stronger val­i­da­tion. The de­tailed ex­per­i­ment record keeps all 103 au­dited en­tries as op­tional ev­i­dence.

Useful en­try points:

Local OpenAI-compatible server

System de­sign

Benchmarks

The ex­per­i­ments that shaped TurboFieldfare

The coolest use for the Vision Pro

christianselig.com

July 29, 2026

Over the last year, I ad­mit­tedly haven’t used my Vision Pro a ton, but fairly re­cently I dis­cov­ered a su­per handy use that it’s ab­solutely in­cred­i­ble at, pro­vided you have the right tools.

After years of apart­ment life my girl­friend and I have re­cently be­gun the process of build­ing our (first! ex­cit­ing!) home, which is an ab­solute whirl­wind of choices and de­ci­sions. Maybe it’s where we’re both pro­gram­mers, but par­tic­u­larly for me, my soft­ware-ori­ented brain feels al­most in­com­pat­i­ble with this world of de­sign­ing a home (she’s the main one keep­ing this pro­ject mov­ing for­ward).

In soft­ware if you don’t end up lik­ing a fea­ture af­ter play­ing around with it for awhile, you can tweak it or re­move it en­tirely (try do­ing that with a poorly placed wall). Here, de­ci­sions feel huge and hard to com­mit to.

On top of this, the hous­ing mar­ket is re­ally ex­pen­sive right now in many parts of the world (ours def­i­nitely in­cluded) so mak­ing good use of every square foot feels para­mount.

(Don’t worry, we have a great ar­chi­tect who guides us skill­fully, but ul­ti­mately it’s up to us to sign off on every­thing!)

Enter the Vision Pro

Looking at PDFs of po­ten­tial floor plans, you don’t (or at least I don’t) de­velop much of a sense of scale or con­nec­tions with a bunch of black rec­tan­gles. This room says it’s 13 feet by 15 feet, I guess I can get out a mea­sur­ing tape, but what does that feel like? Would this hall­way be cramped? What all can you see when you first walk into the house?

Then it hit me: vir­tual re­al­ity! There’s other VR de­vices that have bet­ter gam­ing chops, but I’m not sure there’s a con­sumer de­vice out there that’s bet­ter than the Vision Pro on pa­per when it comes to ren­der­ing a vir­tual world with its high res­o­lu­tion screens and then plac­ing you in it with its abun­dance of sen­sors.

Problem is, how do you go from floor plan to vir­tual re­al­ity?

Let’s get cre­at­in’

I have a bit of ex­pe­ri­ence us­ing Fusion 360 (free 3D mod­el­ing soft­ware for hob­by­ists), so I had the idea to try to quickly build the floor plan up in 3D. Nothing fancy, just floors, ceil­ings, and walls, with holes for doors.

If you’re not fa­mil­iar with 3D de­sign tools, this is a great first pro­ject, you’re ba­si­cally just draw­ing the floor plan out in 2D and then ex­trud­ing out the walls. YouTube and ask­ing AI ques­tions are awe­some re­sources for learn­ing this handy skill in 2026. Sites like Fiverr are also a great re­source, I’m sure there’s an abun­dance of peo­ple there who could take a floor plan you’re in­ter­ested in and turn it into a 3D file for a rea­son­able price.

Leveling it up

A bunch of walls, ceil­ings, and holes in the wall is pretty pow­er­ful, but with­out ac­tual ob­jects or tex­ture in the space it’s still a bit hard to get scale when you’re walking around”.

You know how you walk into an empty apart­ment for the first time and you’re like Holy crap there’s so much room” and then you add your be­long­ings and there sud­denly is­n’t? It’s kinda like that, we need to add some mod­els and ma­te­ri­als to the space to ac­tu­ally ground our per­spec­tive a bit and give things scale, oth­er­wise it feels like walk­ing around a ware­house.

Adding tex­tures

Quick easy one. In Fusion, tap the A” key to bring up the Appearance panel and you can drag and drop a bunch of com­mon tex­tures like wood, stone, or paint onto ob­jects. This can re­ally help to break up the mo­not­ony of every­thing be­ing the same dull tex­ture and ac­tu­ally gives the place depth.

There’s even glass tex­tures for win­dows which is pretty handy.

IKEA

I can model a counter or a kitchen is­land (just a bunch of rec­tan­gles) but any­thing be­yond that and I’m tap­ping out. Thankfully IKEA has a mas­sive amount of fur­ni­ture with cor­re­spond­ing 3D mod­els, and even if you’re not a big IKEA per­son (wow aren’t you fancy) hav­ing any ap­prox­i­ma­tion for your fi­nal fur­ni­ture in your place re­ally helps you get an idea of how things fit to­gether and can be po­si­tioned.

Problem is, no easy way to get ac­cess to those 3D mod­els from what I can tell. Thankfully, with a script for the Tampermonkey browser ex­ten­sion you can eas­ily down­load them. Not every item on IKEA has a 3D model un­for­tu­nately, but seem­ingly the ma­jor­ity do.

This is su­per handy for adding a couch, a bed, a desk, an area rug, etc. to the space so when you walk into a bed­room your eyes kinda see the bed and desk and get a good idea for the scale of the room, which lets you un­der­stand if it’s a good size or if the di­men­sions are proper, and you can kinda imag­ine what it would be like to ex­ist in that space.

Downside is this down­loads a glb file, some­thing I was­n’t fa­mil­iar with. Fusion al­lows us to eas­ily im­port obj files, so we need to con­vert it to that. After try­ing a bunch of scripts and web­sites, I hon­estly found this web­site to be the best. Little janky, but it works.

Also note that the re­sult­ing obj file (and the folder it’s con­tained in) of­ten won’t ren­der tex­tures in­side Fusion 360, but when you ex­port them from Fusion they show up as ex­pected.

3D Warehouse

That cov­ers a lot of fur­ni­ture, but not every­thing. Maybe you want to see how a stand mixer looks on the counter or if you can fit a rice cooker com­fort­ably. Heck, you can find your car and put it in the garage for ref­er­ence.

This is where 3D Warehouse comes in, it’s a com­mu­nity-dri­ven web­site where peo­ple can up­load free 3D mod­els they cre­ate and eas­ily im­port them into Sketchup, an­other pop­u­lar piece of 3D mod­el­ing soft­ware (that might even work bet­ter than Fusion for house de­sign, I dunno, haven’t used it much).

Problem is, the files aren’t eas­ily us­able in Fusion, but I found a nice workaround.

If you open the URL on your iOS de­vice in­stead, it gives you the op­tion to view it in 3D in­side your cur­rent en­vi­ron­ment. Within that view, if you tap the share but­ton you get ac­cess to the ac­tual USDZ file that pow­ers this ex­pe­ri­ence, at which point you can just AirDrop it to your Mac. Nice!

From there, even though we love a USDZ file, Fusion seem­ingly does­n’t for im­port­ing ob­jects, so back to that pre­vi­ous web­site to con­vert to OBJ. From there it’s as sim­ple as bring­ing it into Fusion and plac­ing it.

This web­site is se­ri­ously awe­some, there’s a model for just about any­thing you’d put in your house.

Programming!

Now that I have some­thing I’m happy with, I re­mem­bered Apple plat­forms love the USDZ file for­mat for 3D mod­els, and sure enough Fusion had a USDZ ex­port op­tion. Nice!

With a file in hand, we’re get­ting some­where, so I AirDropped to it to my Vision Pro and opened it. This worked pretty okay for view­ing it, and you could even walk around a lit­tle bit, but un­less you live in a ware­house walk­ing the full length of a vir­tual house with­out bump­ing into some­thing and end­ing up in those VR fail com­pi­la­tions is pretty tricky. I want some­thing bet­ter.

This is where vibe cod­ing is per­fect: putting some­thing tech­ni­cally im­pres­sive to­gether when you would have never both­ered to take weeks to put it to­gether tra­di­tion­ally. Is the code per­fect? Nah, but the al­ter­na­tive is it never ex­ist­ing, and it’s just for fun” any­way.

After some in­tense typ­ing with a com­bi­na­tion of Claude and Codex over the course of a morn­ing I had some­thing I was pretty happy with. I named it Prospector (can’t re­mem­ber why) and I’ve been re­ally happy with it.

Here are a bunch of the fea­tures that level it up be­yond just view­ing a USDZ file in the Files app:

Controller sup­port! Walk around like you’re in a 3D video game, with mo­tion con­trols as well as the abil­ity to ro­tate the cam­era (you can still walk and look around nor­mally of course, this just aug­ments that).

Skybox! Pretty rudi­men­tary, but added a for­est style sky­box as the ex­te­rior of the world so if you look out a win­dow you don’t just see your cur­rent en­vi­ron­ment

Terrain fol­low­ing! If you tap the right D-pad but­ton the con­troller will map you to the ter­rain, so you’ll go up and down with any un­du­la­tions in the ter­rain. This is re­ally handy if you have a sur­vey of your prop­erty done so can ef­fec­tively have a full view of your prop­er­ty’s ter­rain and walk” through it

Toggle real life! If you hold your thumb and mid­dle fin­ger to­gether for a mo­ment it will tog­gle the vir­tual world on and off, this makes it feel safer if you want to take a step for­ward in vir­tual re­al­ity but are un­sure if you’re go­ing to bump into a table in real life.

Take flight! With the trig­gers you can fly up or down, great for go­ing be­tween floors for in­stance, and then if you press up on the D-pad it’ll re­set you to the ground level.

Speed mode! Pressing left on the D-pad makes you move faster, which is re­ally, well, fast when you’re cross­ing a big prop­erty rather than pok­ing around a sin­gle room.

With all those to­gether, this is a re­ally help­ful way to view USDZ files on the Vision Pro. You can re­ally eas­ily ex­plore a space and po­si­tion your­self in new ar­eas to look around, and even take off the head­set to show a spouse or friend so you can gather their in­put on any changes. Then mak­ing a change is as sim­ple as tweak­ing the file in Fusion, re-ex­port­ing, then re-run­ning.

Download link

If you want to play around with Prospector, my su­per janky, vibe coded USDZ viewer, I put the con­tents up on GitHub (it would take a fair bit more work to pol­ish this up into some­thing I’d be com­fort­able sub­mit­ting to the App Store).

Using it is as ad­ver­tised: janky. Take your USDZ file, im­port it into Xcode, and in­side ImmersiveView.swift change the USDZ file name to what­ever the name of your file is. You can also tweak the sky­box (Poly Haven has a bunch of awe­some op­tions). Then, just pair a con­troller to your Vision Pro and run the pro­ject on your de­vice.

Just know I coded like 0% of this, so if it’s ter­ri­ble you can’t judge me. Don’t you dare. I’m sen­si­tive. Judge the AI.

This is awe­some

No se­ri­ously, this feels so pow­er­ful. When we fi­nally de­cided on a plan, through the fancy Revit soft­ware our ar­chi­tect uses he was able to bring us on a 3D walk­through of our fu­ture house. It was cool, kinda like a Google Streetview-esque ex­pe­ri­ence where you could click to tele­port around the house and drag to look around.

But hon­estly, af­ter al­ready see­ing it in full, im­mer­sive 3D where you can look around and feel like you’re there, we al­ready felt like we knew the place. It re­ally feels like one day in the fu­ture vis­it­ing po­ten­tial de­signs in 3D will be a core part of the process, and if you have the hard­ware that fu­ture is pos­si­ble to­day.

Statement on behalf of UEFA and its 55 national associations

www.uefa.com

UEFA and its 55 mem­ber as­so­ci­a­tions stand as one. We unan­i­mously and un­equiv­o­cally re­ject FIFAs pro­posal to trans­fer own­er­ship in­ter­ests in the World Cup and other FIFA com­pe­ti­tions to pri­vate in­vestors.

The World Cup can­not be treated as an in­vest­ment prod­uct. It is one of foot­bal­l’s great­est sport­ing lega­cies. It has been built over gen­er­a­tions by play­ers, na­tional teams and sup­port­ers across every con­ti­nent. No part of it should ever be sur­ren­dered to pri­vate in­vestors. The World Cup is not for sale.

It is both ir­re­spon­si­ble and in­de­fen­si­ble that a pro­posal of such sig­nif­i­cance for foot­ball was con­ceived in se­cret and brought to the brink of ap­proval with­out any mean­ing­ful con­sul­ta­tion with those en­trusted with stew­ard­ing the game. This is not merely a pro­found fail­ure of lead­er­ship, but an ab­di­ca­tion of FIFAs duty as the cus­to­dian of world foot­ball.

National as­so­ci­a­tions around the world are now pre­sented with an ul­ti­ma­tum: ac­cept the ir­re­versible cap­ture of foot­bal­l’s great­est com­pe­ti­tions or bear the con­se­quences. This is not a democratic de­ci­sion”, but gov­er­nance by in­tim­i­da­tion — an act of co­er­cion un­wor­thy of an in­sti­tu­tion en­trusted with the stew­ard­ship of the global game.

But our op­po­si­tion goes far be­yond process.

The mo­ment ex­ter­nal in­vestors ac­quire own­er­ship in­ter­ests in FIFA com­pe­ti­tions, foot­ball changes for­ever. Commercial re­turn be­comes a per­ma­nent oblig­a­tion. Investor ex­pec­ta­tions be­come a daily pres­sure. From that mo­ment on­wards, every de­ci­sion on the in­ter­na­tional cal­en­dar, every de­ci­sion on com­pe­ti­tion for­mats and every de­ci­sion shap­ing the fu­ture of foot­ball is no longer dri­ven by what best serves the game, but by what best serves share­hold­ers.

This model has no place in world foot­ball. Football’s fu­ture can­not be dic­tated by the ex­pec­ta­tions of those whose first duty is to max­imise fi­nan­cial re­turn. Nor can the in­ter­ests of na­tional as­so­ci­a­tions, leagues, clubs, play­ers and sup­port­ers be­come sub­or­di­nate to in­vestor re­turns. Football can­not mort­gage its fu­ture for fi­nan­cial gain.

Europe’s po­si­tion is clear. We will never lend this model our le­git­i­macy. No one has the moral au­thor­ity to sell what they merely hold in trust for the next gen­er­a­tion.

As a re­sult of to­day’s dis­cus­sion, no UEFA na­tional teams will par­tic­i­pate in any FIFA com­pe­ti­tion for so long as these pro­pos­als re­main alive, un­less this pro­posal has been aban­doned in its en­tirety and bind­ing as­sur­ances have been given that FIFA will never again open its gov­er­nance or com­pe­ti­tions to pri­vate own­er­ship.

Nobody should be in any doubt: UEFA and its na­tional as­so­ci­a­tions will op­pose these plans with ab­solute de­ter­mi­na­tion.

There are mo­ments when in­sti­tu­tions are judged not by what they are pre­pared to ac­cept, but by what they refuse to com­pro­mise. This is one of those mo­ments.

Some things are sim­ply too im­por­tant to sell. The FIFA World Cup be­longs to foot­ball. It al­ways will. And so long as Europe has a voice, it will never be for sale.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.