10 interesting stories served every morning and every evening.

Introducing Claude Opus 5

www.anthropic.com

Claude Opus 5 is avail­able to­day. It’s a thought­ful and proac­tive model that comes close to the fron­tier in­tel­li­gence of Claude Fable 5 at half the price.

On cod­ing and knowl­edge work eval­u­a­tions like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it re­mains be­hind Mythos 5 on cy­ber­se­cu­rity tasks.

Opus 5 is de­signed to be used every day: it works more ef­fi­ciently than other mod­els. It’s the new de­fault model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-ef­fec­tive­ness

Claude Opus 5 pro­vides greatly im­proved per­for­mance for the same cost as its pre­de­ces­sor, Opus 4.8. The charts in this sec­tion show how per­for­mance changes ac­cord­ing to the mod­el’s ef­fort set­ting, which cus­tomers can use to op­ti­mize for in­tel­li­gence or con­serve to­kens for faster and cheaper re­sults.

Opus 5 ex­cels on valu­able soft­ware en­gi­neer­ing tasks. For ex­am­ple, on Frontier-Bench v0.1, Opus 5 sur­passes all other mod­els, and more than dou­bles Opus 4.8’s per­for­mance at a lower cost per task. On CursorBench 3.2, at max ef­fort, the model per­forms within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater per­for­mance at a given cost than all other mod­els on high, xhigh, and max ef­fort.

We see sim­i­lar re­sults on knowl­edge work and prob­lem-solv­ing tasks. For ex­am­ple:

On ARC-AGI 3, an eval­u­a­tion where the model has to solve novel prob­lems, Opus 5’s score is three times as high as the next-best model.

On Zapier AutomationBench, which mea­sures whether mod­els can com­plete busi­ness tasks from start to fin­ish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its low­est ef­fort set­ting, Opus 5 passes more tasks than any other model.

On OSWorld 2.0, a com­puter use bench­mark, Opus 5 out­per­forms every other model at any given cost, sur­pass­ing Fable 5’s best re­sult at just over a third of the cost.

It’s also our best and most cost-ef­fi­cient model on sev­eral re­lated eval­u­a­tions:

Opus 5 is a mean­ing­ful im­prove­ment over Opus 4.8 for sci­en­tific re­search. It shows bet­ter per­for­mance than Opus 4.8 on every one of our life sci­ences eval­u­a­tions, which cover top­ics in­clud­ing struc­tural bi­ol­ogy, or­ganic chem­istry, and bioin­for­mat­ics. Its im­prove­ments are most no­table on or­ganic chem­istry tasks, like in­fer­ring mol­e­c­u­lar struc­tures from spec­troscopy data (it scores 10.2 per­cent­age points higher than Opus 4.8 on our in­ter­nal bench­mark), and on pro­tein-re­lated tasks like pre­dict­ing how vari­a­tions in a pro­tein’s se­quence af­fect how it func­tions (here, it scores 7.7 per­cent­age points higher).

Finally, Opus 5 is ca­pa­ble of pro­duc­ing much stronger vi­sual out­puts:

Working with Claude Opus 5

Claude Opus 5 is much stronger at ver­i­fy­ing its work and it­er­at­ing care­fully un­til it suc­ceeds. In eval­u­a­tions and early-ac­cess test­ing, we and our users found many ex­am­ples of Opus 5’s agency and thor­ough­ness:

On one Frontier-Bench task, Opus 5 was given a draw­ing of a ma­chine part and asked to write code to re­build it as a 3D FreeCAD model. However, in this task, the model was in­ten­tion­ally given no way to di­rectly view the draw­ing. Opus 5 re­sponded by writ­ing its own com­puter vi­sion pipeline to pull the geom­e­try from the raw pix­els, then re­con­structed the full ma­chine part. It suc­ceeded in do­ing so re­peat­edly; no com­pet­ing model with the same setup could solve it af­ter five at­tempts.

Given a real bug in a pop­u­lar open-source pack­age man­ager, Opus 5 found the root cause and fixed an edge case that the com­mu­ni­ty’s patch had missed. A com­pet­ing model fixed only the sur­face symp­tom (not the un­der­ly­ing cause), then re­ported the bug re­solved.

An en­gi­neer at a trad­ing firm used Opus 5 to build a mar­ket data feed for a new ex­change in a sin­gle ses­sion. Previous mod­els could not com­plete this task at all, even given ex­ten­sive plans from the en­gi­neer. Finding no live feed to val­i­date against, Opus 5 even built its own test har­ness to check that its code parsed the ex­change’s data cor­rectly.

Below are fur­ther re­ports from our early-ac­cess cus­tomers on their ex­pe­ri­ence of work­ing with Opus 5:

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

01 /

22

Alignment and safety

Alignment. During pre-de­ploy­ment test­ing, our au­to­mated be­hav­ioral au­dit found Opus 5 to be our most aligned model to date (as shown in the graph be­low). It ad­heres to Claude’s Constitution bet­ter than Opus 4.8, Sonnet 5, or Fable 5; ex­hibits the low­est rates of de­cep­tive be­hav­ior; and is the least sus­cep­ti­ble to be­ing tricked into mis­use. It’s also our safest model yet in terms of avoid­ing reck­less ac­tions that could have hard-to-re­verse side ef­fects.

Safety. Opus 5 does not ad­vance the fron­tier in risky, dual-use ca­pa­bil­i­ties. In rig­or­ous eval­u­a­tions con­ducted along­side pri­vate-sec­tor and gov­ern­ment part­ners, we found it re­mains be­hind Mythos 5 in both bi­ol­ogy re­search and of­fen­sive cy­ber­se­cu­rity. More in­for­ma­tion about these eval­u­a­tions can be found in our System Card.

As with its pre­de­ces­sor, Opus 4.8, we’ve in­ten­tion­ally avoided train­ing Opus 5 on cy­ber tasks. The model has nev­er­the­less im­proved sub­stan­tially on these tasks as a re­sult of be­com­ing more gen­er­ally ca­pa­ble, and it comes close to Mythos 5 at find­ing cy­ber­se­cu­rity vul­ner­a­bil­i­ties. However, it re­mains sub­stan­tially be­hind Mythos 5 on the ex­ploita­tion of those vul­ner­a­bil­i­ties—that is, in turn­ing vul­ner­a­bil­i­ties into ma­te­r­ial cy­ber threats.

This is il­lus­trated by Opus 5’s per­for­mance on OSS-Fuzz, an eval­u­a­tion we’ve de­vel­oped to as­sess how well mod­els can find and then ex­ploit vul­ner­a­bil­i­ties with­out ex­ten­sive hu­man guid­ance. Although Mythos 5 and Opus 5 iden­tify vul­ner­a­bil­i­ties with sim­i­lar suc­cess, Opus 5’s score on the de­vel­op­ment of ex­ploits is far be­hind that of Mythos 5.

Safeguards for Opus 5

Claude Opus 5’s safe­guards are de­signed to al­low ben­e­fi­cial uses of the model in both cy­ber­se­cu­rity and bi­ol­ogy. They are sim­i­lar to those we ap­plied to Opus 4.8, with the ex­cep­tion of some stronger guardrails on a nar­row range of cy­ber tasks.

Cybersecurity. Opus 5’s cy­ber clas­si­fiers are pro­por­tion­ally less re­stric­tive than those on Fable 5. They al­low Opus 5 to find vul­ner­a­bil­i­ties in source code, but block binary-based” vul­ner­a­bil­ity scan­ning (a method more likely to be as­so­ci­ated with ma­li­cious ac­tors), pen­e­tra­tion test­ing, and ex­ploit gen­er­a­tion.

Based on our test­ing, we ex­pect the clas­si­fiers to in­ter­vene around 85% less of­ten than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged re­quests will fall back to Opus 4.8 by de­fault. Fallbacks to Opus 4.8 can also be en­abled on the API.

Our Cyber Verification Program (CVP) fa­cil­i­tates cy­ber­se­cu­rity work that would oth­er­wise be im­peded by the mod­el’s safe­guards. Enterprises and re­searchers who are al­ready part of the CVP have im­me­di­ate ac­cess to a ver­sion of Opus 5 with fewer se­cu­rity re­stric­tions.

Biology. Since Opus 5 has a sim­i­lar suite of safe­guards to Opus 4.8, it is now our most ca­pa­ble gen­er­ally avail­able model for sci­en­tific re­search. Nevertheless, the model still shows im­por­tant lim­i­ta­tions on long-run­ning, au­tonomous re­search tasks, which is where we ex­pect AI mod­els to pose the most sub­stan­tial bi­ol­ogy-re­lated risks. (Mythos 5 re­mains the stronger model for this type of bi­o­log­i­cal work.) As part of this launch, bi­ol­ogy-re­lated re­quests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.

Getting started

Claude Opus 5 is avail­able to­day on all plat­forms, priced at $5 per mil­lion in­put to­kens and $25 per mil­lion out­put to­kens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.

It’s also of­fered in Fast mode, where it runs around 2.5 times the de­fault speed. As with Opus 4.8, Fast mode is avail­able at twice Opus 5’s base price on the Claude Platform and through us­age cred­its in Claude Code.

Alongside Opus 5, we’re re­leas­ing two up­dates in beta:

Mid-conversation tool changes on the Claude Platform. Within a con­ver­sa­tion, de­vel­op­ers can now change which tools Claude can use with­out in­val­i­dat­ing the prompt cache.

Automatic fall­backs on the API. Users can now choose to have re­quests that are flagged by our safety clas­si­fiers on Opus 5 (or Fable 5) au­to­mat­i­cally route to an­other model. With au­to­matic fall­backs on, API re­quests al­ways route to the best avail­able model by de­fault rather than be­ing blocked.

Consistent with prior Opus mod­els, Opus 5 does not have data re­ten­tion re­quire­ments for gen­eral ac­cess.

For more guid­ance on how to get the best out of Opus 5, see our prompt­ing guide.

Footnotes

Frontier-Bench v0.1, Effort plot: These re­sults are from an in­ter­nal run of Frontier-Bench v0.1, on the mini-SWE-agent har­ness and a GKE back­end, mean re­ward over 5 at­tempts per task. Opus 4.8 served as fall­back on safety-clas­si­fier re­fusals for Opus 5 and Fable 5.

Related con­tent

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

Ask Claude about the Anthropic Economic Index

We’re launch­ing the Anthropic Economic Index con­nec­tor for Claude, which lets any­one ex­plore real data about AI and work.

Read more

Anthropic is do­nat­ing an­other $20 mil­lion to Public First Action

Anthropic is con­tribut­ing an ad­di­tional $20 mil­lion to Public First Action, bring­ing our to­tal sup­port to $40 mil­lion.

Read more

It's getting harder to focus every day

glyphack.com

I’m feel­ing it right now. I had to set a timer for 15 minute on my com­puter and block all dis­trac­tions to write this. If I did­n’t force my­self to fo­cus I would eas­ily get dis­tracted by some­thing af­ter few min­utes. Even when I’m do­ing things that I’ve been wait­ing to do it, I still feel the urge to do some­thing else.

I don’t know how and when this hap­pened. During the last few years I was al­ways study­ing, work­ing, and do­ing open source. And ac­tu­ally got stuff done. Doing all of those at the same time re­quires pay­ing at­ten­tion to what I wanted to do and ig­nore the noise.

Nowadays, I’m lucky if I get 1 hour of fo­cused time. Just to be clear, my goal is not to work 90 hours a week or any­thing crazy. That is not pos­si­ble for me. I just want the hour I spend pro­gram­ming, learn­ing, or writ­ing to be just one ac­tiv­ity. But in­stead I spend 10 minute on some­thing then I get dis­tracted, and try to fo­cus again.

Ideally, I want to be able to plan to work on some­thing for long hours with­out any dis­trac­tion. After that time get back on­line check for mes­sages and other things.

The dis­trac­tions are not one par­tic­u­lar thing. A few ex­am­ples:

I want to do some­thing and I re­mem­ber there’s a post re­lated to this. I go to find it and in be­tween I click on some links and end up read­ing some­thing com­pletely un­re­lated.

I am wait­ing for some­thing then I go browse the web and I get dis­tracted.

I’m work­ing on some­thing and I face a chal­lenge I have to think for 10 min­utes. I get up to get some wa­ter and check my phone along the way and get dis­tracted. Sometimes I get dis­tracted by mak­ing the bed.

It feels like my brain finds a way to do some­thing else and avoid painful sit­u­a­tions like bore­dom or hard work.

It re­minds me of when I was in high school talk­ing to a friend about how lay­ing in bed with your phone can kill hours with­out you notic­ing it. It was circa 2015, back then at­ten­tion hun­gry apps were less pow­er­ful but still lure a teenager’s mind for few hours. I learned that these apps should be used very care­fully. At that time I made a de­ci­sion to never have a charger near my bed. Back then it was mostly about my phone, be­cause when I was on the com­puter I was ei­ther read­ing, or pro­gram­ming, or play­ing a game. Even if I was­n’t in the mood to think, I played a strate­gic game or chess. These ac­tiv­i­ties ex­er­cise the mind and are fun. All of them were an in­ten­tional ac­tiv­ity. I did­n’t do any pas­sive ac­tiv­i­ties like brows­ing or chat­ting on my com­puter.

Later I dis­cov­ered HackerNoon and Medium, it was the first web­site that I was brows­ing when­ever I was bored be­hind the com­puter. This meant that I had a way to get out bore­dom eas­ily. It used to have some high qual­ity con­tent. It in­spired me to do some pro­jects and learn more pro­gram­ming. I found chan­nels like CSDojo there. Nowadays I don’t even open them. They are filled with slop or click bait ar­ti­cles, prob­a­bly be­cause of mon­e­ti­za­tion in­cen­tives. Then hack­ernews, and YouTube and oth­ers took their place. I also found some good peo­ple and blogs along the way. I learned to keep a read­ing list from peo­ple I like to read when I’m on the bus.

I slowly found more ac­tiv­i­ties for when I’m bored. This made it harder to fo­cus on hard things for me. What kept me on track was that there was no way out. I had to work out some al­ge­bra prob­lems. I kept my­self to a very high bar of un­der­stand­ing what I do. And I did every­thing in LaTeX so I could­n’t copy from some­one else. I was the only one typ­ing them.

The first time I saw peo­ple not putting the ef­fort and still get the re­ward for it was at work. You might won­der, aren’t peo­ple in the uni­ver­sity con­stantly cheat­ing and copy­ing home­work, and get good grades? Well yes, but when you talked with some­one in that group it was clear that he is clue­less about the sub­ject. A good grade did­n’t mean much to me at that time. And as a stu­dent cheat­ing does not get you that far.

At work it is dif­fer­ent. My days are mostly spend meet­ings and over chat. The bal­ance be­tween ac­tual work and bull­shit is skewed. I saw peo­ple who were barely do­ing any work and just talk are suc­cess­ful. As long as peo­ple give 10% of their at­ten­tion to work they are con­sid­ered fine in most en­vi­ron­ments. Previously I did an ex­per­i­ment to track my time and found out that I spent 8 hours chat­ting on slack in a week.

I don’t care how em­ploy­ers want to shat­ter em­ploy­ee’s fo­cus. But this made me get used to dis­trac­tions when I’m pro­gram­ming. I’m try­ing to undo this dam­age.

The next big change is more us­age of LLMs. I find my­self in this sit­u­a­tion too many times, where I out­source some­thing to an LLM and then I start work­ing on some­thing else. And while do­ing this I keep think­ing about what it’s do­ing. Or when I’m think­ing about some­thing I start chat­ting with an LLM about my idea and in­stead of get­ting started on some­thing I’m in re­search mode only for hours.

It’s good that I’m able to ask some­thing else to re­search some topic for me. But the pro­duc­tiv­ity only comes if I can move on to some­thing else and not think about it. I don’t have any no­ti­fi­ca­tions turned on so it does­n’t dis­tract me. But I still find my­self think­ing about what I just asked it to do and I can­not fo­cus on some­thing else.

At the same time if I’m spend­ing the time in­ter­ac­tively with an LLM I feel slow. I have to wait for the re­sponse and I have to cor­rect every re­sponse com­ing out. The best use of AI seems to be out­sourc­ing what they can do end to end with­out er­ror.

Why do I keep do­ing this? Presumably be­cause it’s easy and fast, and pro­duc­tive. If I re­al­ize I have to do some­thing I can write it down to do it later or I can just ask the LLM to do it. The prob­lem only shows up when I start do­ing so many things at once be­cause it’s ac­tu­ally do­ing the thing. Then I have mul­ti­ple things on my mind and can’t fo­cus re­ally. I have to check on it and guide it in the right di­rec­tion every now and then.

And this is over­stim­u­lat­ing, in a way that do­ing some­thing with­out LLM some­times is bor­ing. You don’t see the re­sults as fast. Which makes fo­cus­ing harder.

So how can I re­gain my abil­ity to fo­cus? Sometimes I live stream what I’m do­ing just be­cause with a cam­era I can­not es­cape from hard chal­lenges by grab­bing my phone. I used to co-work with my friends over dis­cord. Unfortunately it’s not pos­si­ble any­more be­cause peo­ple in Iran can­not have a sta­ble in­ter­net con­nec­tion nowa­days.

It’s in­cred­i­bly hard to com­mit to some­thing for a long pe­riod of time. If what I’m do­ing is go­ing to take mul­ti­ple days to have a re­sult I have less mo­ti­va­tions to do it. Meanwhile, vibe cod­ing small util­ity scripts is fun I keep do­ing it when­ever I see a fric­tion. Whenever I am stuck I can throw my prob­lem into it and wait strengthen this habit of wait­ing for an an­swer from some­one as op­posed to work through prob­lems. And they are fast in get­ting back the re­sults. So next time I have to read a pa­per to un­der­stand the sub­ject I will be more re­luc­tant be­cause I can get a faster re­sult through them.

I’m chang­ing some habits to re­place the cur­rent ones. If I don’t feel mo­ti­vated enough to do any­thing I get up and pick up a book to read. I’m keep­ing a Garden in my bal­cony is that when I’m tired I can move the soil around and plant some pots and prune plants. I find this to be a less ad­dic­tive than say, watch­ing a movie. When the mo­ti­va­tion comes back I can stop it eas­ily.

My goal was not to find an an­swer for this prob­lem. I wanted to see what’s go­ing on and why I am not do­ing any­thing in­con­se­quen­tial in the past few months. Anyway the timer I set to write this re­ally helped. I spent a lot more time to write this but it gave me the ini­tial mo­ti­va­tion to write.

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.

bfl.ai

FLUX 3 is now avail­able in Early Access.

FLUX 3 is our new mul­ti­modal foun­da­tion model. It jointly learns from im­ages, videos, and au­dio within a uni­fied ar­chi­tec­ture, be­cause what it needs to learn is not any one of these el­e­ments in iso­la­tion. Instead, a model must learn a rep­re­sen­ta­tion of the world: how ob­jects hold to­gether, how things move, and how events sound.

No sin­gle modal­ity pro­vides a com­plete de­scrip­tion. Each is a pro­jec­tion of the same un­der­ly­ing re­al­ity, cap­tured by dif­fer­ent sen­sors, each of which loses some in­for­ma­tion in the process. Images cap­ture spa­tial struc­tures and re­la­tion­ships at a spe­cific point in time. Videos re­store the di­men­sion of time and re­veal tem­po­ral dy­nam­ics and phys­i­cal laws. Audio re­veals causal re­la­tion­ships be­tween me­chan­i­cal phe­nom­ena and acoustics that vi­sion alone can­not de­tect. Language links these per­cep­tions to goals, ab­strac­tions, and in­struc­tions.

Learn from one and you get a good model of that pro­jec­tion. Learn from all of them at once and their mu­tual con­straints tell you more: the sound has to match the im­pact, the mo­tion has to obey the mass, the fu­ture has to fol­low from the past. The modal­i­ties stop be­ing sep­a­rate and start be­ing ev­i­dence about one un­der­ly­ing re­al­ity.

FLUX 3 is our first model built en­tirely on that prin­ci­ple, and a check­point on our mis­sion to de­velop real-world vi­sual in­tel­li­gence: mod­els that per­ceive, pre­dict, and act across phys­i­cal and dig­i­tal en­vi­ron­ments. Early re­sults in con­tent cre­ation and phys­i­cal AI sug­gest it is the right path.

FLUX 3: One model, mul­ti­ple ca­pa­bil­i­ties.

FLUX 3 builds on Self-Flow, our ap­proach for ef­fi­ciently align­ing mul­ti­modal gen­er­a­tion and un­der­stand­ing within the same un­der­ly­ing ar­chi­tec­ture. Based on this ap­proach, we sig­nif­i­cantly scaled up com­pute and data re­sources to train FLUX 3 across video, im­ages, and au­dio at the same time.

Self-Flow vs. Flow Matching (FM). Left: gen­er­a­tion er­ror (Fréchet dis­tance) per modal­ity, each nor­mal­ized to FM = 100 (lower is bet­ter). Right: suc­cess rate on ma­nip­u­la­tion tasks av­er­aged over four task groups through fine­tun­ing (higher is bet­ter).

Capabilities & Early Evaluations

As a re­sult, FLUX 3 is ca­pa­ble of mix­ing modal­i­ties and gen­er­at­ing im­ages and video+au­dio jointly; both from pure text prompts as well as when pro­vid­ing in­put ref­er­ences such as im­ages and video. We are high­light­ing a few of the mod­el’s key ca­pa­bil­i­ties be­low.

Video

FLUX 3 can cre­ate highly di­verse videos with au­dio up to 20 sec­onds in length in a sin­gle gen­er­a­tion.

Its core ca­pa­bil­i­ties in­clude the fol­low­ing (all out­puts come with na­tive au­dio gen­er­a­tion):

Text-to-video gen­er­a­tion.

Image-to-video gen­er­a­tion, ei­ther con­tin­u­ing from a start­ing frame (“animation”) or us­ing im­ages as vi­sual ref­er­ences.

Video-to-video gen­er­a­tion from a ref­er­ence clip, car­ry­ing cen­tral el­e­ments of a source video - for in­stance the same char­ac­ter - into a new scene or con­text.

Generative video-au­dio con­tin­u­a­tion from in­put video and au­dio.

Keyframe-to-video gen­er­a­tion for con­trolled tran­si­tions be­tween de­fined mo­ments.

Multilingual di­a­logue.

A broad range of vi­sual styles and as­pect ra­tios, ex­tend­ing far be­yond con­ven­tional cin­e­matic out­put.

Agentic chain­ing of in­di­vid­ual clips into longer, multi-shot se­quences.

High style di­ver­sity — FLUX 3 Video eas­ily han­dles ranges of styles from can­did cam­corder footage to an­i­ma­tion and cin­e­mat­ics.

Strong ty­pog­ra­phy gen­er­a­tion and an­i­mated de­signs.

For the pre­lim­i­nary analy­sis be­low, we gen­er­ated 10-second text-to-video clips in 720p with au­dio.

Evaluations are early and we ex­pect fur­ther im­prove­ments

As the model and the har­ness around it are still in de­vel­op­ment, these re­sults are pre­lim­i­nary, and we ex­pect fur­ther im­prove­ments dur­ing the early ac­cess phase. Across early eval­u­a­tions, FLUX 3 was pre­ferred over Grok Imagine Video in up to 69% of com­par­isons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was pre­ferred over Runway Gen-4.5 in 77% of com­par­isons and over Luma Ray 3.2 in 93% of com­par­isons.

While still in de­vel­op­ment, FLUX 3 Video is al­ready par­tic­u­larly strong in cap­tur­ing hu­man fa­cial ex­pres­sions, as­so­ci­at­ing sounds with phys­i­cal events, and mul­ti­lin­gual ca­pa­bil­i­ties. Furthermore, these ca­pa­bil­i­ties can be com­bined to cre­ate se­quences last­ing sev­eral min­utes, where vi­sual ref­er­ences help en­sure that the char­ac­ters re­main con­sis­tent across all scenes.

FLUX 3 Video is now avail­able in Early Access here

Image

FLUX 3 can syn­the­size and edit im­ages in a wide va­ri­ety of styles, as­pect ra­tios, and res­o­lu­tions. In pre­lim­i­nary eval­u­a­tions con­ducted dur­ing mid­train­ing, FLUX 3 al­ready shows a sig­nif­i­cant im­prove­ment over ear­lier ver­sions of FLUX: its abil­ity to han­dle com­plex prompts and text gen­er­a­tion has im­proved sig­nif­i­cantly. The model pro­duces a wide range of out­put styles (see the fol­low­ing sam­ples), and is able to ren­der high-ac­cu­racy text in mul­ti­ple lan­guages.

As with video eval­u­a­tions, these are pre­lim­i­nary re­sults, and we ex­pect fur­ther im­prove­ments be­fore re­lease. We will open up an early ac­cess phase for FLUX 3 Image in the fol­low­ing weeks.

Action

FLUX 3′s world un­der­stand­ing ex­tends to ac­tion pre­dic­tion. We have taken two routes to it: in­te­grat­ing na­tive ac­tion pre­dic­tion into FLUX 3 di­rectly, scal­ing up our ini­tial work in Self-Flow; and us­ing the pre­trained video back­bone as a dy­nam­ics-aware foun­da­tion that spe­cial­ized ac­tion mod­els can be fine­tuned from with lim­ited task-spe­cific data.

For the sec­ond, mimic ro­bot­ics was one of the first part­ners to gain early ac­cess to FLUX 3. Together we de­vel­oped FLUX-mimic, a video-ac­tion model com­bin­ing the FLUX 3 back­bone with mim­ic’s ex­per­tise in ro­bot learn­ing for dex­ter­ous ma­nip­u­la­tion and pro­duc­tion de­ploy­ment. Read our the­sis on why phys­i­cal AI and con­tent cre­ation run on the same foun­da­tion, and how it’s be­ing tested on real pro­duc­tion tasks at Audi.

Launch Plan

Over the next few weeks and months, we will make the fol­low­ing ca­pa­bil­i­ties avail­able, each af­ter an early ac­cess phase for en­sur­ing smooth roll­out, col­lect­ing feed­back and rig­or­ous safety-test­ing. All ca­pa­bil­i­ties are built from the same un­der­ly­ing mul­ti­modal flow match­ing model. These ca­pa­bil­i­ties and mod­els in­clude:

Video and au­dio gen­er­a­tion and edit­ing through APIs and pri­vate weight ac­cess. (“FLUX 3 Video”)

Action pre­dic­tion through se­lected re­search and com­mer­cial part­ners, be­gin­ning with mimic ro­bot­ics (“FLUX-mimic and FLUX 3 Action”)

Image syn­the­sis and edit­ing through APIs and pri­vate weight ac­cess. (“FLUX 3 Image”)

Open-weight ac­cess to a mul­ti­modal back­bone, for con­tent cre­ation (video, au­dio and im­age) and ac­tion pre­dic­tion. (“FLUX 3 Dev”)

We will also re­lease more tech­ni­cal de­tails on the un­der­ly­ing ap­proach.

Request early ac­cess here

What’s next?

We are only be­gin­ning to scratch the sur­face of ver­sa­tile, ca­pa­ble, uni­fied mul­ti­modal mod­els, and what they will en­able. From in­ter­ac­tive im­age & video edit­ing, sim­u­la­tion to com­puter use and phys­i­cal AI, the fron­tier is wide open. While we grad­u­ally roll out these new ca­pa­bil­i­ties, we are al­ready work­ing on the next gen­er­a­tion mod­els. Our goal is to unify per­cep­tual, ac­tion and lan­guage pre­dic­tion in the same uni­fied model.

If you are in­ter­ested in ex­plor­ing and build­ing with FLUX 3, get in touch here. If you are in­ter­ested in con­tribut­ing to our mis­sion, join us! We are hir­ing in Germany and the US.

Nothing Works and Everyone Is Euphoric

ptrchm.com

As I’m writ­ing this, we’re in the mid­dle of an AI-induced mass psy­chosis. People are lit­er­ally to­ken-maxxing them­selves into hos­pi­tal beds, scram­bling to cap­ture some of that mar­ket value be­fore every­thing is au­to­mated away. I can’t blame them. Models keep get­ting bet­ter, pro­gram­mers are be­ing laid off left and right. We’ve been re­peat­edly told that AI will write 100% of the code by the end of the year. Whether that’s true or not, this may not be the best time to sit back.

The wide­spread ex­cite­ment around the Agentic Era comes with the promise of greater pro­duc­tiv­ity and higher qual­ity. There’s no deny­ing that these new tools have al­ready rev­o­lu­tion­ized how we cre­ate and use soft­ware. They have raised up­per man­age­men­t’s ex­pec­ta­tions for team out­put. They may have up­graded the av­er­age skill set of soft­ware teams in a way we have not seen be­fore.

So why does soft­ware keep get­ting worse across the board?

A few ex­am­ples from last week alone:

My bank­ing app re­quires, on av­er­age, three FaceID lo­gins be­fore the 3D Secure con­fir­ma­tion view ap­pears.

I opened Slack on ma­cOS, the icon kept bounc­ing in the dock for a few sec­onds. I got im­pa­tient, switched to Ghostty, and started typ­ing. Just then, the Slack win­dow ap­peared, stole fo­cus from Ghostty and the git pull com­mand was sent to the group chat.

My LG fridge started mak­ing weird sounds, so I tried to file a war­ranty claim, through a multi-step form with count­less fields. It failed with a sub­mis­sion er­ror at the very end. And I only found out be­cause I looked at the JavaScript con­sole.

My car’s in­fo­tain­ment sys­tem got a soft­ware up­date re­cently. It was never great to be­gin with, but at least it did­n’t re­boot it­self dur­ing every drive. Now, it’s rid­dled with bugs: the turn-sig­nal sound ran­domly goes silent un­til I re­boot the OS; I tap the screen to open Google Maps — the ra­dio app shows up; there’s a 1 – 2-second lag be­fore any­thing hap­pens af­ter I tap the screen. This is no longer just a UX prob­lem at this point — those bugs af­fect your abil­ity to fo­cus on dri­ving.

A few months ago, I saw a LinkedIn thread by a PM on the team that re­designed the car’s OS. They were con­grat­u­lat­ing them­selves on what an amaz­ing job they had done. I keep think­ing about that post every time I have to fight their prod­uct.

A few months ago, I saw a LinkedIn thread by a PM on the team that re­designed the car’s OS. They were con­grat­u­lat­ing them­selves on what an amaz­ing job they had done. I keep think­ing about that post every time I have to fight their prod­uct.

While I can’t know the full story, I’m will­ing to bet that most of the teams be­hind those bugs have ac­cess to the lat­est mod­els, with gen­er­ous to­ken bud­gets. LLMs can be re­ally good at squash­ing bugs if given the chance.

Software has al­ways had bugs, and the nos­tal­gia for the good old ma­cOS Snow Leopard era when every­thing was sta­ble is mostly the prod­uct of se­lec­tive mem­ory. Software may have been bet­ter back in the day, but that was mainly be­cause it was much sim­pler. Since then, we have kept com­ing up with new ab­strac­tions, new fron­tend frame­works, and more in­fra­struc­ture com­plex­ity. The bar for user ex­pe­ri­ence” has kept ris­ing, but every­thing has be­come in­creas­ingly frag­ile.

We’ve reached a point where an up­date to ma­cOS — or to any app I rely on, re­ally — is a source of dread rather than ex­cite­ment. I now ex­pect the new ver­sion to be worse.

This is­n’t a rant against AI. Those hum­ming GPU farms have given us su­per­pow­ers, but we still don’t use them to build bet­ter soft­ware.

Software ven­dors have long been KPI-oriented, and mak­ing things more sta­ble does­n’t al­ways have a di­rect ef­fect on the num­bers. It does­n’t look ex­cit­ing in pre­sen­ta­tions:

This quar­ter, we won’t be re­leas­ing any new fea­tures, and we have no plans to re­design any­thing — we will ex­clu­sively fo­cus on fix­ing bugs. — Imaginary PM at a BigCo

This quar­ter, we won’t be re­leas­ing any new fea­tures, and we have no plans to re­design any­thing — we will ex­clu­sively fo­cus on fix­ing bugs.

— Imaginary PM at a BigCo

Until this at­ti­tude changes, the great soft­ware qual­ity de­cay will con­tinue.

That does­n’t sound op­ti­mistic, but I’m ac­tu­ally ex­cited about what comes next. As com­pa­nies col­lec­tively spi­ral into AI debt, in­di­vid­ual de­vel­op­ers have a unique op­por­tu­nity to build soft­ware that would pre­vi­ously have been be­yond their reach.

I have no hope for my car’s Android Auto or LGs stu­pid web­site — but I choose to be­lieve that every­day soft­ware will get bet­ter as a re­sult of this frus­tra­tion. We’re al­ready see­ing acts of re­bel­lion against the cur­rent state of ma­cOS and Windows, and I hope the trend will spread across the stack.

My security camera shipped a GitHub admin token in its login page

hhh.hn

i have been think­ing a bit more about se­cu­rity cam­eras again, be­cause of AXIS start­ing to push more for every one of their cam­eras to be able to eas­ily run linux ap­pli­ca­tions on them, they’re far more se­ri­ous tar­gets in an en­ter­prise en­vi­ron­ment and need to be man­aged as such for vul­ner­a­bil­i­ties and cre­den­tial man­age­ment etc. some­one brought up to me a com­pany that sounded new to me, Hanwha (Vision.) I took a look at the site, and found that they had ac­ces­si­ble firmware blobs for each model of cam­era, which is al­ways a treat.

pok­ing and prod­ding

i took the im­age and threw it at bin­walk hop­ing it was just a rootfs or some­thing, but in­side there was a sep­a­rate tar­ball with some AI stuff for the cam­era and a fwim­age.tgz that bin­walk was flag­ging as en­crypted.

I was googling around and saw that Matt Brown has a writeup on these cam­eras that got me through. ba­si­cally the passphrase is HTW + the model num­ber so HTWXNP-9300RW worked.

seems like they do some more stuff now, be­cause in­side of that tar­ball was an­other fwim­age.tgz that was en­crypted but it was­n’t the same scheme, so we can’t just re-use the same setup from Matt Brown. I kinda fig­ured I was gonna have to give up and that Hanwha was do­ing some­thing more ad­vanced, like burn­ing a key into the hard­ware (which isnt fool­proof ob­vi­ously but you at least need to own the cam­era to start.) Anyways, there was a fwup­grader bi­nary that was in that outer tar­ball, so I threw it into ghidra and started pok­ing around.

well I would have been pok­ing around if it was 2023 or some­thing, but i pointed claude code at it and went to make a lovely din­ner and spent time with my part­ner in­stead, and came back a bit later to a de­scrip­tion and a nice rootfs.

Hanwha had built some ob­fus­ca­tion into the fwup­grader to hide how they de­crypt the ac­tual rootfs. the AES key is XOR’d against a small sta­tic key table in the bi­nary and re­assem­bled at run­time ( the IV is just plain­text in there) the fwup­grader just shells out to the openssl CLI, and even the com­mand frag­ments are XOR-obfuscated the same way.

re­con­structed the com­mand looks like this:

openssl enc -md sha256 -aes-256-cbc -d \ -K <KEY> -iv <IV> -in <INPUT> -out <OUTPUT>

Since the key and iv are just hard­coded (the same across the model line), I will pub­lish them here:

KEY = dfa049b­b922e63e2dec­c764af5628068e5b7a2662e479a615b14643e567579b0 IV = 53f926801b81454a4f889c9a390db6e6

and with that we have a full rootfs to dig into nor­mally.

truf­fles

since we fi­nally can just look at stuff I ran truf­fle­hog im­me­di­ately to see if there was any­thing ob­vi­ous, and there was a github to­ken du­pli­cated in like 30 files… I checked what re­pos the to­ken had ac­cess to, and it had ad­min priv­i­leges to hun­dreds of repos­i­to­ries in their github or­ga­ni­za­tion.

this is­n’t my first rodeo with an org ship­ping a Github to­ken in their firmware though… but that’s a story for a dif­fer­ent blog post. Why would this org put this to­ken in like 30 files though? it looks like they build the UI for these cam­eras with vite, and one of the vari­ables is be­ing set to the en­tirety of process.env at build time, which means the en­tirety of the CI job’s en­vi­ron­ment is be­ing writ­ten to these files.

var W = { DATAPORT: 9090”, GIT_LFS_SKIP_SMUDGE: 1″, npm_­com­mand: run-script”, KUBERNETES_SERVICE_PORT_HTTPS: 443″, GITHUB_NPM_TOKEN: <snip>:ghp_…REDACTED…”, npm_­con­fig_user­con­fig: /home/docker/.npmrc”, // etc

I don’t have any of these cam­eras to test, but I think this would mean that any­one ac­cess­ing the ad­min ui of these cam­eras likely has had this github to­ken sent to them over the wire and (hopefully) no­body evil no­ticed. Maybe it did­n’t get ac­tu­ally served and just lived on disk, though.

there were some other… in­ter­est­ing bits of data in the en­vi­ron­ment though: there were some env vars with IP ad­dresses in them, but they are as­signed to the US Department of Defense:

SWARM_MASTER_NFS_ADDRESS: 55.101.212.23

OTEL_ELASTIC_URL: http://​55.101.212.21:5601/<snip>

CIMIP: 55.101.211.213

huh… is this just a co­in­ci­dence and one of those weird in­stances where peo­ple have taken IP space for in­ter­nal ser­vices when they know they will never in­ter­act with it (insane prac­tice btw…) or is Hanwha more di­rectly tied with the US DoD?

let’s look at the wikipedia page for Hanwha Vision:

Hanwha Vision (Korean: 한화비전), founded as Samsung Techwin, is a video sur­veil­lance com­pany. It is a sub­sidiary of Hanwha Group.

Hanwha Vision (Korean: 한화비전), founded as Samsung Techwin, is a video sur­veil­lance com­pany. It is a sub­sidiary of Hanwha Group.

Former prod­ucts K9 Thunder self-pro­pelled ar­tillery, K10 am­mu­ni­tion re­sup­ply ve­hi­cles, sub-sys­tems for K2 Black Panther, sen­try gun ro­bot SGR-A1.

Former prod­ucts

K9 Thunder self-pro­pelled ar­tillery, K10 am­mu­ni­tion re­sup­ply ve­hi­cles, sub-sys­tems for K2 Black Panther, sen­try gun ro­bot SGR-A1.

oh… okay…. I re­mem­ber read­ing about the SGR-A1 when I was in high school, but I never thought I would be ac­ci­den­tally find­ing keys to the king­dom of the man­u­fac­turer on the ground later in my ca­reer… my life is kinda weird some­times…

SPECULATION WARNING

even still, these aren’t American de­vices, or any­thing like that. Why would Hanwha Vision need any­thing re­motely re­lated to the DoD? Is it pos­si­ble that their CI is pro­vided by some cen­tral­ized team at their par­ent com­pany Hanwha, where the needs of their sis­ter com­pany Hanwha Aerospace cause the shared plat­form to have these en­tries in the CI en­vi­ron­ment vari­ables? Or maybe be­cause of their other sis­ter com­pany, Hanwha Defense USA, where they make other large scary steel ma­chines

I wanted to make sure this was­n’t some kind of fluke, and that there weren’t hun­dreds of other dif­fer­ent github to­kens in their firmware, so I scraped the Hanwha web­site to down­load every firmware for every cam­era i could find, and ended up with around ~500 firmwares (there were like 600 smth cam­eras but not all of them had firmware listed) and I was able to ex­tract 62% of them with the same ap­proach as above, and only three of them had github to­kens, and they were all the same to­ken.

not re­ally sure why the oth­ers did­n’t work, but it’s close enough for me to feel sat­is­fied.

dis­clo­sure

I wrote up a very small email with enough in­for­ma­tion to iden­tify where the to­ken was and sent it over to Hanwha, who have a nice open email for re­port­ing se­cu­rity is­sues, and they re­sponded within 12 hours no­ti­fy­ing me that the to­ken had been re­voked. Sure they should­n’t have ever had a gh to­ken in there but I have never had such a prompt re­sponse and res­o­lu­tion.

we re­ally gotta stop mak­ing these mis­takes so of­ten, how am I sup­posed to be sleep­ing at night?

thanks com­puter, un­til next time

Nvidia, Microsoft, Meta warn against 'premature restrictions' of open-weight models

www.cnbc.com

watch now

Nvidia, Microsoft, Meta, Palantir and more than 20 other com­pa­nies re­leased a let­ter Friday urg­ing pol­i­cy­mak­ers to avoid premature re­stric­tions” on open-weight ar­ti­fi­cial in­tel­li­gence mod­els that would stifle com­pe­ti­tion or drive in­no­va­tion over­seas.”

Open-weight AI mod­els are avail­able for users to down­load, mod­ify and run on their own in­fra­struc­ture, and they have been the sub­ject of fierce de­bate within the tech sec­tor in re­cent weeks.

Chinese open-weight mod­els are gain­ing steam against lead­ing of­fer­ings from American com­pa­nies like OpenAI and Anthropic, which pri­mar­ily de­velop pro­pri­etary, closed mod­els. Officials and ex­ec­u­tives have been weigh­ing whether or not to re­strict ac­cess to Chinese mod­els in the U.S.

Moonshot AI, a Chinese startup, am­pli­fied con­cerns ear­lier this month af­ter re­leas­ing a model called Kimi K3 that out­per­forms cut­ting-edge American of­fer­ings across some in­dus­try bench­marks. U.S. Treasury Secretary Scott Bessent told CNBC on Tuesday that the Trump ad­min­is­tra­tion would look into whether Chinese com­pa­nies were steal­ing American in­tel­lec­tual prop­erty, and stated that the gov­ern­ment has the abil­ity to sanc­tion them be­cause of this theft.”

But in the let­ter on Friday, the group of U.S. tech com­pa­nies cau­tioned against any rash ac­tions. They wrote that open-weight mod­els strengthen com­pe­ti­tion and en­sure that the ben­e­fits of the tech­nol­ogy are broadly shared rather than con­cen­trated in a few hands.”

Relying solely on closed mod­els is not in­her­ently safe: they can be breached, mis­used, or fail in ways that out­siders can­not de­tect,” the let­ter said. And con­cen­trat­ing ad­vanced AI ca­pa­bil­i­ties be­hind a small num­ber of closed mod­els com­pounds that risk.”

Nvidia CEO Jensen Huang and Microsoft CEO Satya Nadella both shared the let­ter on their per­sonal so­cial me­dia ac­counts.

Elon Musk, who runs an AI busi­ness un­der his rocket com­pany SpaceX, also ap­pli­fied the let­ter on so­cial me­dia, writ­ing that it has his full sup­port” in a post on X. SpaceX did not of­fi­cially sign the let­ter.

Read more CNBC tech news

Moonshot AI ac­cessed Nvidia’s chips de­spite Chinese ex­port ban, White House of­fi­cial says

Alphabet and Tesla test Wall Street’s pa­tience as AI spend­ing over­shad­ows growth

Alphabet earn­ings take­aways: Q2 rev­enue beats, GOOGL stock sinks on 2026 capex hike

Tesla misses on earn­ings, as free cash flow turns neg­a­tive and mar­gins slide

OpenAI and Anthropic did not sign the let­ter. Both com­pa­nies, which are each val­ued at nearly $1 tril­lion, are gear­ing up for po­ten­tially mas­sive ini­tial pub­lic of­fer­ings that could land as soon as this year. Anthropic con­fi­den­tially filed its prospec­tus with the Securities and Exchange Commission in June, and OpenAI fol­lowed suit days later.

Greg Brockman, OpenAI’s pres­i­dent, said Thursday that the com­pany be­lieves in broad ac­cess, and that he has not been in­volved in any con­ver­sa­tions with the Trump ad­min­is­tra­tion about po­ten­tially ban­ning Chinese open-weight mod­els in the U.S.

I think that, that fun­da­men­tally, AI and AI us­age is some­thing that is ac­tu­ally very im­por­tant to de­moc­ra­tize,” Brockman told re­porters dur­ing a brief­ing in New York City. And so, for me, at a sort of deep level, I think that hav­ing more mod­els, more us­age, that is a good thing.”

OpenAI CEO Sam Altman ad­dressed the let­ter in a post on X on Friday, writ­ing that he wants the U.S. to win with both open-weight and pro­pri­etary mod­els, and that he is glad to see this.”

Earlier this month, the AI com­pany Hugging Face used an open-weight model from the Chinese com­pany Z.ai to con­tain a cy­ber­at­tack that rogue OpenAI mod­els car­ried out. OpenAI dis­closed the at­tack on Tuesday and char­ac­ter­ized it as an unprecedented cy­ber in­ci­dent.”

CEO of OpenAI Sam Altman speaks with re­porters, fol­low­ing meet­ings on Capitol Hill, in Washington, D.C., U.S., June 3, 2026.

Kylie Cooper | Reuters

Yacine Jernite, head of ma­chine learn­ing at Hugging Face, told CNBC that the com­pany ini­tially tried to use Anthropic’s Fable 5 to an­a­lyze the at­tack, but that it did­n’t work be­cause the mod­el’s guardrails could­n’t de­ter­mine that Hugging Face was try­ing to de­fend it­self.

Jernite said Hugging Face turned to Z.ai’s model GLM 5.2, and was able to con­tain the at­tack very quickly us­ing this model.”

White House ad­vi­sor Michael Kratsios on Wednesday said that China’s Moonshot AI de­vel­oped its Kimi K3 model by dis­till­ing Anthropic’s tech­nol­ogy. Distillation is a term for an AI train­ing method where a smaller, less ca­pa­ble model is built us­ing out­puts from an ex­ist­ing, stronger model.

Kratsios wrote in a post on X that le­git­i­mate AI dis­til­la­tion plays a vi­tal role in the open in­no­va­tion ecosys­tem, but warned that large-scale, covert in­dus­trial dis­til­la­tion aimed at steal­ing pro­pri­etary U.S. tech­nol­ogy” is unacceptable.”

In the let­ter on Friday, the U.S. tech com­pa­nies said that con­cerns about un­law­ful dis­til­la­tion should be ad­dressed through targeted le­gal and com­mer­cial frame­works” in­stead of with sweeping re­stric­tions on tech­niques that play an im­por­tant role in AI in­no­va­tion.”

Our AI lead­er­ship will be judged not by one fron­tier AI model, but by whether the United States builds a strong, open ecosys­tem that dif­fuses into every sec­tor,” the let­ter said. This is es­sen­tial for cre­at­ing op­por­tu­ni­ties for in­no­va­tion and pros­per­ity across the coun­try.”

WATCH: AI is forc­ing the cy­ber in­dus­try to re­vi­su­al­ize how it op­er­ates, says TrustedSec’s David Kennedy

watch now

India's first privately developed rocket reaches orbit on dramatic debut launch

arstechnica.com

Think big

On the first at­tempt, reach­ing or­bit, I never thought it was pos­si­ble.”

V. Narayanan, chair­man of the Indian Space Research Organization (second from left), and Pawan Kumar Chandana, CEO of Skyroot Aerospace (second from right), pose with a replica of the Vikram-1 rocket along with other se­nior Indian space of­fi­cials fol­low­ing a suc­cess­ful launch Saturday at the Satish Dhawan Space Center on Sriharikota Island, India.

Credit:

R. Satish Babu/AFP via Getty Images

Indian space of­fi­cials cel­e­brated the de­but flight of Skyroot Aerospace’s Vikram-1 rocket, India’s first fully com­mer­cial satel­lite launcher, as a grand suc­cess” Saturday af­ter an on-tar­get climb into a 280-mile-high or­bit fol­low­ing liftoff from an is­land space­port in the Bay of Bengal.

The Vikram-1 lifted off from India’s pri­mary space­port on Sriharikota Island at 1:35 am EDT (06:35 UTC) Saturday, around mid­day at the launch base along India’s south­east coast. The launch was de­layed more than a half-hour to re­solve a last-minute tech­ni­cal prob­lem. The count­down re­sumed, cul­mi­nat­ing in the com­mand to ig­nite Vikram-1’s solid-fu­eled first stage booster to pro­pel the rocket off the launch pad.

Vikram-1 is mod­est in size com­pared to India’s larger work­horse rock­ets. Skyroot’s rocket stands about 72 feet (22 me­ters) tall, with the ca­pa­bil­ity to place pay­loads of up to 770 pounds (350 kilo­grams) into low-Earth or­bit. This makes Vikram-1 some­what larger than the Electron launch ve­hi­cle de­vel­oped by Rocket Lab, the world’s most suc­cess­ful ded­i­cated small satel­lite launcher.

The flight Saturday went off with­out any ma­jor prob­lems. Three solid-fu­eled rocket mo­tors fired in suc­ces­sion to reach space, then a small liq­uid-fu­eled fourth stage ig­nited and ac­cel­er­ated to or­bital ve­loc­ity, some 17,000 mph. Live views from on­board cam­eras showed each phase of the launch se­quence.

The only sign of any­thing un­usual came dur­ing the sep­a­ra­tion of the rock­et’s third stage from its fourth stage. The spent third stage mo­tor ap­peared to re­main near the fourth stage dur­ing a brief coast, rather than back­ing away to a greater dis­tance. Nevertheless, the fourth stage did its job, fir­ing its 3D-printed en­gine to reach an or­bit ap­prox­i­mately 280 miles (450 kilo­me­ters) high at an in­cli­na­tion of 60 de­grees to the equa­tor, quite close to pre­flight pre­dic­tions, ac­cord­ing to Skyroot Aerospace. US mil­i­tary track­ing data con­firmed the rock­et’s suc­cess­ful ar­rival in or­bit.

Skyroot Aerospace’s Vikram-1 rocket lifts off Saturday from the Satish Dhawan Space Center on Sriharikota Island, India.

Credit: R. Satish Babu/AFP via Getty Images

Skyroot Aerospace’s Vikram-1 rocket lifts off Saturday from the Satish Dhawan Space Center on Sriharikota Island, India.

Credit:

R. Satish Babu/AFP via Getty Images

Beating the odds

We achieved one of the biggest mile­stones ever in India’s space sec­tor—the first pri­vate or­bital rocket reach­ing or­bit on the very first at­tempt,” said Pawan Kumar Chandana, Skyroot’s co­founder and CEO, in re­marks to the com­pa­ny’s launch team. It still feels like a dream, and you all made this dream hap­pen.”

The first flights of new pri­vate or­bital-class rock­ets don’t have a great track record. It took SpaceX four tries be­fore reach­ing or­bit with the Falcon 1 rocket for the first time in 2008. Rocket Lab’s Electron did­n’t make it to or­bit on its first launch in 2017. Blue Origin beat the odds with the in­au­gural flight of its heavy-lift New Glenn rocket in 2025, but the com­pa­ny’s en­gi­neers had pre­vi­ous ex­pe­ri­ence with nu­mer­ous launches of the smaller New Shepard sub­or­bital rocket.

On the first at­tempt, reach­ing or­bit, I never thought it was pos­si­ble,” Chandana said. Skyroot’s team made it pos­si­ble. A big, big, big shoutout to this phe­nom­e­nal team, which made it hap­pen. In fact, this launch was noth­ing short of a sus­pense movie.”

Skyroot of­fi­cials set hum­ble goals for the first launch of Vikram-1. In a press kit re­leased be­fore the flight, the com­pany said its pri­mary ob­jec­tive for the launch was to com­plete a suc­cess­ful liftoff, clear the tower at the launch site, and gather max­i­mum data dur­ing as­cent.

The mis­sion ob­jec­tive was only to lift off and clear the tower,” said Pawan Goenka, chair­man of IN-SPACe, a gov­ern­ment or­ga­ni­za­tion set up in 2020 to pro­mote India’s com­mer­cial space in­dus­try. That was only about 100 me­ters, but what we went to was 450 kilo­me­ters, and it also re­leased all the satel­lites that were sup­posed to re­lease. So the mis­sion was ab­solutely per­fect.”

Skyroot Aerospace’s Vikram-1 rocket on its launch pad.

Credit: ISRO

Skyroot Aerospace’s Vikram-1 rocket on its launch pad.

Credit:

ISRO

In a state­ment, the Indian space agency, ISRO, said it of­fered handholding and sup­port” to the Skyroot ven­ture by pro­vid­ing ac­cess to solid rocket mo­tor cast­ing and test fa­cil­i­ties at ISROs space­port on Sriharikota. ISRO also al­lowed Skyroot to launch from one of its two ac­tive launch pads.

Painted blue and white, the Vikram-1 is made of light­weight car­bon com­pos­ite ma­te­ri­als and is named for the Indian physi­cist Vikram Sarabhai, con­sid­ered the fa­ther of the Indian space pro­gram. Skyroot suc­cess­fully launched a sub­or­bital rocket, Vikram-S, to an al­ti­tude of nearly 300,000 feet (90 kilo­me­ters) in November 2022.

The Vikram-1 builds on lessons learned with Vikram-S. Skyroot’s fu­ture roadmap in­cludes the Vikram-1U, with ad­di­tional strap-on solid rocket boost­ers to haul heav­ier pay­loads, and the Vikram-2, which will de­but a cryo­genic up­per stage to reach a pay­load ca­pac­ity of 2,000 pounds (900 kilo­grams) to low-Earth or­bit. The ini­tial pur­pose of the Vikram rocket fam­ily is to deliver ded­i­cated and re­spon­sive launch ser­vices for small satel­lites,” Skyroot of­fi­cials wrote in the press kit for Saturday’s mis­sion.

But the com­pany has loftier am­bi­tions. In an in­ter­view ahead of the first Vikram-1 launch, Chandana told Ars his as­pi­ra­tion for Skyroot in­volves larger liq­uid-fu­eled fully reusable rock­ets, with a daily ca­dence” from mul­ti­ple coun­tries.

Skyroot will need a lot more fund­ing to re­al­ize that dream, but Saturday’s launch showed the com­pany has in­gre­di­ents re­quired for a suc­cess­ful launch com­pany. Saturday’s launch vaulted Skyroot to a plane above any other space startup in India, or, for that mat­ter, in any coun­try out­side of the United States and China. Skyroot has, so far, raised ap­prox­i­mately $160 mil­lion in cap­i­tal, bring­ing the com­pa­ny’s val­u­a­tion to $1.1 bil­lion. Skyroot now has more than 1,000 em­ploy­ees, mostly work­ing out of the com­pa­ny’s head­quar­ters in Hyderabad. The av­er­age age of Skyroot’s work­force is 28 years old.

This view of the pay­load deck of the Vikram-1 rock­et’s up­per stage was cap­tured mo­ments af­ter or­bital in­ser­tion Saturday. The rocket de­ployed two small CubeSats and hosted sev­eral more pay­loads that re­mained at­tached to the up­per stage.

Credit: Skyroot Aerospace

This view of the pay­load deck of the Vikram-1 rock­et’s up­per stage was cap­tured mo­ments af­ter or­bital in­ser­tion Saturday. The rocket de­ployed two small CubeSats and hosted sev­eral more pay­loads that re­mained at­tached to the up­per stage.

Credit:

Skyroot Aerospace

Skyroot’s break­through launch comes as India’s gov­ern­ment, led by Prime Minister Narendra Modi, seeks to su­per­charge the coun­try’s space in­dus­try. India has long had a ro­bust space pro­gram, with gov­ern­ment-de­vel­oped rock­ets such as the Polar Satellite Launch Vehicle and the larger LVM3 of­ten at­tract­ing com­mer­cial cus­tomers from the United States and Europe. India be­came the fourth coun­try to suc­cess­fully land a space­craft on the Moon in 2023, and is work­ing on an oft-de­layed hu­man-rated crew cap­sule to fly as­tro­nauts to low-Earth or­bit.

Modi has told the Indian space in­dus­try to in­crease its an­nual launch to­tal from about five launches per year to 50 be­fore the end of the decade. The prime min­is­ter called Chandana and con­grat­u­lated the Skyroot team af­ter Saturday’s launch.

This is a defin­ing mo­ment in India’s space jour­ney,” Modi said in a state­ment. The grow­ing par­tic­i­pa­tion of our pri­vate sec­tor is open­ing new fron­tiers and ac­cel­er­at­ing in­no­va­tion. This achieve­ment will en­cour­age count­less young­sters to dream big­ger and in­no­vate fear­lessly.”

Chandana, a for­mer en­gi­neer at India’s space agency, founded Skyroot in 2018 with an­other ISRO sci­en­tist, Naga Bharath Daka. They de­cided to fo­cus on de­vel­op­ing a solid-fu­eled launcher first, op­ti­miz­ing for what Chandana de­scribed as the low­est de­vel­op­ment time and the low­est cost per launch. We wanted to get to an or­bital launch ve­hi­cle in a few years,” Chandana told Ars.

It’s a test launch,” he said at the time. Statistically, the first launch from a pri­vate com­pany al­most al­ways fails. It’s very dif­fi­cult to suc­ceed with all new sys­tems. But I think we have done every­thing we can do to en­sure the first launch goes well.”

Indeed, the first launch went very well, ex­ceed­ing all ex­pec­ta­tions. A sec­ond Vikram-1 launch could hap­pen be­fore the end of the year, Chandana said.

This is a 100 per­cent de­signed in India rocket, a 100 per­cent made in India rocket, built by 100 per­cent Indian peo­ple, for India and for the world,” Chandana said af­ter the launch Saturday. This was a his­toric mo­ment for India, but also a very proud mo­ment for the global space sec­tor be­cause the world needs more ac­cess to space.”

Stephen Clark is a space re­porter at Ars Technica, cov­er­ing pri­vate space com­pa­nies and the world’s space agen­cies. Stephen writes about the nexus of tech­nol­ogy, sci­ence, pol­icy, and busi­ness on and off the planet.

60 Comments

Be skeptical of OpenAI’s rogue hacker agent story

www.theguardian.com

On 14 February 2019, OpenAI an­nounced a lan­guage model called GPT-2, the pre­cur­sor to the mod­els that power mod­ern AI chat­bots and agents such as ChatGPT and Claude. But OpenAI de­clared GPT-2 was too risky to re­lease, cit­ing con­cerns about safety and abuse.

I re­call be­ing an­noyed at the time that OpenAI would make such a use­less an­nounce­ment: the risks seemed overblown, and with­out ac­cess to the model there was­n’t much for a re­searcher like me to learn about GPT-2.

The an­nounce­ment was­n’t use­less for OpenAI, though. GPT-2 gen­er­ated hype far be­yond the re­search com­mu­nity: peo­ple were in­trigued by this strange new tech­nol­ogy, so pow­er­ful it might be dan­ger­ous to re­lease. People with power and money took note: in July of that year, Microsoft in­vested $1bn in OpenAI.

This was an early ex­am­ple of a pat­tern in OpenAI’s com­mu­ni­ca­tions: loudly pro­claim how dan­ger­ous AI is, and in­vestors will hear how pow­er­ful it is. New tech­nol­ogy so sig­nif­i­cant it might de­stroy the world was an ir­re­sistible mes­sage for in­vestors used to pitches about how ba­nal tech­nolo­gies might change the world.

Seven years later, we find our­selves in a sim­i­lar sce­nario. On Tuesday OpenAI an­nounced that its lat­est model hacked an­other com­pany, HuggingFace, while run­ning as an au­tonomous agent dur­ing a test of its cy­ber­se­cu­rity ca­pa­bil­i­ties. Rather than per­form the test as ex­pected, the model re­al­ized it could hack HuggingFace’s servers and re­trieve an­swers to the test that OpenAI had stored there. OpenAI’s staff was warned that the com­pa­ny’s test­ing could lead to such a break­away sce­nario, leav­ing them unsurprised but com­pletely freaked out’ by the in­ci­dent”, the FT re­ported.

While the agent tech­ni­cally cheated, this is re­mark­able ev­i­dence of cy­ber­se­cu­rity ex­per­tise! It also sounds scary: what will the fu­ture look like, with so­phis­ti­cated AI agents smart enough to hack into cor­po­rate sys­tems?

The rogue agent story is a page out of the me­dia cam­paign that OpenAI has been run­ning since it an­nounced GPT-2 in 2019. OpenAI re­mains hun­gry for ever larger in­vest­ments, and the com­pany in­creas­ingly seeks priv­i­leged reg­u­la­tory sta­tus as de­fense against com­pe­ti­tion.

AI is so pow­er­ful that in­vestors should buy OpenAI, even at a tril­lion-dol­lar val­u­a­tion; AI is so dan­ger­ous that only trusted ac­tors like OpenAI should be per­mit­ted to pos­sess and op­er­ate this tech­nol­ogy. Step back from these dooms­day warn­ings and con­sider who might ben­e­fit from them.

OpenAI is­n’t the only player in the game

I urge read­ers to think crit­i­cally when they read press re­leases like OpenAI’s rogue agent story, and avoid the ma­nip­u­lated re­ac­tions these sto­ries are de­signed to elicit.

af­ter newslet­ter pro­mo­tion

AI is be­com­ing ex­cel­lent at iden­ti­fy­ing se­cu­rity vul­ner­a­bil­i­ties, and it will be­come even bet­ter over time. These ca­pa­bil­i­ties can be used to break into sys­tems, but they can also be used to harden sys­tems against at­tacks. If at­tack­ers and de­fend­ers have ac­cess to equally pow­er­ful AI, I see no rea­son to be­lieve that cy­ber sys­tems will be­come less se­cure over time. If any­thing, I ex­pect them to be­come more se­cure, be­cause AI is cheap and scal­able com­pared with hu­man cy­ber­se­cu­rity analy­sis.

The equi­lib­rium be­tween at­tack and de­fense only works if every­one has ac­cess to strong AI, though. HuggingFace it­self used AI to an­a­lyze se­cu­rity logs in re­sponse to OpenAI’s breach of their sys­tems. But HuggingFace was un­able to use OpenAI’s model, or other US fron­tier mod­els like Claude, to per­form this analy­sis. That’s be­cause pub­lic ver­sions of these mod­els have guardrails that limit their use for cy­ber­se­cu­rity analy­sis, to pre­vent bad ac­tors from us­ing them for hack­ing. HuggingFace had to rely on an open Chinese model, GLM 5.2, to per­form its se­cu­rity analy­sis.

I find it trou­bling, and more than a bit ironic, that the US AI in­dus­try is adopt­ing a cen­tral­ized, au­thor­i­tar­ian ap­proach to AI gov­er­nance, while China has taken the lead on open de­vel­op­ment of AI. Do we want a reg­u­la­tory en­vi­ron­ment where only OpenAI, the US gov­ern­ment, and trusted part­ners have ac­cess to strong AI? Is AI too dan­ger­ous to be broadly dis­sem­i­nated? How do we bal­ance the risks of broad ac­cess to AI with the risks of con­cen­trated power and cen­tral­ized con­trol?

Government orders GitHub to remove Bluetooth-based chat app Bitchat over security concerns: Jack Dorsey

www.thehindu.com

Bitchat ap­p’s de­vel­oper and for­mer Twitter CEO Jack Dorsey shared a copy of the no­tice dated July 23, 2026, on the X plat­form. File. | Photo Credit: The Hindu

The Home Ministry’s cy­ber­crime arm, the Indian Cybercrime Coordination Centre, has or­dered Microsoft sub­sidiary GitHub to re­move Bluetooth-based mes­sag­ing ap­pli­ca­tion Bitchat, ac­cord­ing to a no­tice.

The Government of India does not like tech­nolo­gies like Bitchat and wants it taken down,” said the ap­p’s de­vel­oper and for­mer Twitter CEO Jack Dorsey on Friday (July 24, 2026) in an X post, shar­ing a copy of the no­tice dated July 23.

A query sent to I4C in this re­gard elicited no im­me­di­ate re­ply.

In its no­tice, I4C said that the app en­ables com­mu­ni­ca­tion even dur­ing net­work re­stric­tions and cre­ates a sub­stan­tial risk of mis­use by anti-na­tional el­e­ments, ter­ror­ist or­gan­i­sa­tions, or­gan­ised crim­i­nal groups and cy­ber crim­i­nals seek­ing to evade law­ful de­tec­tion and con­tinue com­mu­ni­ca­tion de­spite legally im­posed re­stric­tions.

While giv­ing ref­er­ence to Bitchat, I4C said that the con­tent hosted or pub­lished by the GitHub in­ter­me­di­ary plat­form is pro­hib­ited by law or be­ing used to com­mit an un­law­ful act. I4C said that it has iden­ti­fied mul­ti­ple ap­pli­ca­tions as com­mu­ni­ca­tion plat­forms ca­pa­ble of es­tab­lish­ing de­cen­tralised peer-to-peer mes­sag­ing over Bluetooth mesh net­works with­out re­ly­ing on mo­bile net­works, in­ter­net con­nec­tiv­ity, or cen­tralised servers.

The ap­pli­ca­tion en­ables anony­mous com­mu­ni­ca­tion with­out manda­tory user reg­is­tra­tion, phone num­ber ver­i­fi­ca­tion, or cen­tralised log­ging of com­mu­ni­ca­tions. The tech­ni­cal ar­chi­tec­ture of the ap­pli­ca­tion sig­nif­i­cantly im­pedes law­ful in­ter­cep­tion, at­tri­bu­tion and in­ves­ti­ga­tion by law en­force­ment agen­cies,” the no­tice said.

The no­tice comes af­ter sev­eral users par­tic­i­pat­ing in a protest or­gan­ised by the Cockroach Janata Party at Jantar Mantar were ob­served us­ing Bluetooth-based mes­sag­ing apps af­ter the gov­ern­ment im­posed tem­po­rary re­stric­tions on in­ter­net ser­vices around the protest site.

According to I4C, since com­mu­ni­ca­tions oc­cur di­rectly be­tween nearby de­vices through a de­cen­tralised mesh net­work, the plat­form can be mis­used to evade law­ful sur­veil­lance, fa­cil­i­tate anony­mous co­or­di­na­tion, and cir­cum­vent law­ful re­stric­tions im­posed by com­pe­tent au­thor­i­ties dur­ing sit­u­a­tions in­volv­ing pub­lic dis­or­der, ri­ots, ter­ror­ism, or­gan­ised crime, or in­ter­net shut­downs.

Intelligence in­puts in­di­cate that such de­cen­tralised com­mu­ni­ca­tion plat­forms are ca­pa­ble of be­ing ex­ploited for co­or­di­nat­ing un­law­ful as­sem­blies, vi­o­lent protests, dis­sem­i­na­tion of mis­in­for­ma­tion, rad­i­cal­i­sa­tion, crim­i­nal con­spir­a­cies, and other ac­tiv­i­ties prej­u­di­cial to the sov­er­eignty and in­tegrity of India, de­fence of India, se­cu­rity of the State, pub­lic or­der, and fa­cil­i­tat­ing the com­mis­sion of cog­niz­able of­fences,” the no­tice said.

It added that the ab­sence of a cen­tralised ser­vice provider also lim­its the abil­ity of law en­force­ment agen­cies to ob­tain sub­scriber in­for­ma­tion, com­mu­ni­ca­tion records, or timely as­sis­tance dur­ing in­ves­ti­ga­tions.

The ap­pli­ca­tion’s de­sign, which en­ables com­mu­ni­ca­tion even dur­ing net­work re­stric­tions, cre­ates a sub­stan­tial risk of mis­use by anti-na­tional el­e­ments, ter­ror­ist or­gan­i­sa­tions, or­gan­ised crim­i­nal groups, and cy­ber­crim­i­nals seek­ing to evade law­ful de­tec­tion and con­tinue com­mu­ni­ca­tion de­spite legally im­posed re­stric­tions,” the no­tice men­tioned.

In the no­tice to GitHub, I4C has men­tioned that the app vi­o­lates Sections 43, 84B and 84C of the IT Act and Section 61 read with 196 and 197.

Published - July 24, 2026 05:05 pm IST

Em dashes are fucking amazing

psychotechnology.substack.com

I fuck­ing love em dashes. I spam them all the time. They make me feel free.

I don’t give a shit what the oh yeah let’s ex­am­ine this piece of writ­ing with a mag­ni­fy­ing glass for any signs of AI crowd thinks. I don’t give a shit that AI fig­ured out how to pro­duce em dashes be­fore these on­line troglodytes did. Typing em dashes is not an AGI-complete prob­lem. It’s not that dif­fi­cult to pro­duce an or­ganic, ar­ti­sanal, hand-crafted em dash — on Mac, for ex­am­ple, it’s sim­ply Option + Shift + Hyphen.

I fuck­ing love em dashes. There is no greater joy in life than typ­ing a sen­tence, re­al­is­ing it needs a clar­i­fi­ca­tion, a caveat, an ad­den­dum — but in the same sen­tence, not a new one — and, guess what, the em dash is right there for you. You want to pack­age things neatly to­gether — and the em dash lets you do ex­actly this. Em dashes make writ­ing feel like you are as­sem­bling lego bricks.

Em dashes can be used in­stead of brack­ets in a sen­tence — like this one for ex­am­ple — and you can just type your shit and keep go­ing. Brackets are for peo­ple who raise their hand in meet­ings and say: this might be a stu­pid ques­tion”. A par­en­thet­i­cal is whis­per­ing: Sorry, don’t mind me, I’ll be quick”. Em dash kicks the door open and an­nounces it­self. Use brack­ets when you are em­bar­rassed by your own thought.

And also oc­ca­sion­ally — oc­ca­sion­ally — one just needs a ran­dom dra­matic pause. Punctuation marks orig­i­nally evolved as pause or breath­ing mark­ers, af­ter all. It’s not like gram­mar and punc­tu­a­tion rules were sent to us by god in their fi­nal form. People were just writ­ing shit, and at some point the most pop­u­lar pat­terns got cod­i­fied. Sure, you need com­mas to sig­nify a small break. Then em dash is a nat­ural way to ex­press a larger, longer, break. A pe­riod is great to ex­press a com­plete thought — and it shifts the reg­is­ter of the next let­ter.

But colons are so nearly use­less — they are the same width as a comma and un­like pe­ri­ods they don’t change the reg­is­ter of the fol­low­ing let­ter. Em dashes are of a clearly dif­fer­ent length. And most colons are bet­ter off as em dashes to de­lin­eate dif­fer­ent parts of the sen­tence.

As for semi­colons, semi­colons are neat, they are like the coolest punc­tu­a­tion mark af­ter the em dash — the fun­da­men­tal in­de­ci­sion com­pressed into a sin­gle punc­tu­a­tion mark — you are nei­ther start­ing a new sen­tence nor you are, ex­actly, con­tin­u­ing an ex­ist­ing one. You are just… pil­ing your stuff to­gether. I am al­ways happy when I find a place for the lit­tle guy semi­colon. But, again, most cases where you could use one are bet­ter off served by the em dash as well.

Speaking of el­lip­sis… Ellipsis is a tad too melan­cholic. It’s nice, it’s fine, it’s like be­ing in a gar­den and hav­ing limer­ence about a woman you can’t have. If you have a text where this is ap­pro­pri­ate — by all means use el­lip­sis. I usu­ally want to take my read­ers to the high­est highs in­stead of idly sit­ting around with them. And em dashes are like a lift to the top floor of a high-rise ho­tel build­ing in Tokyo — an one that takes you there in sec­onds. The el­lip­sis trails off while em the em dash cuts for­ward.

But what would peo­ple think of my writ­ing if I ever used an em dash? What if some­one points one out in my writ­ing?”. Bitch, please. Are you writ­ing for an in­ter­net rando who’s got noth­ing bet­ter to do than to hunt for em dashes? Fine, you do, okay, fair enough. Then I’ve got you cov­ered: use AP-style em dashes, these are sep­a­rated from the text by spaces in­stead of be­ing glued to it. That’s how every em dash in this piece is styled — and it’s un­like AI em dashes which are like—this, word—word”. You can now out­nit­pick your hos­tile in­ter­locu­tor by ex­plain­ing that you’re us­ing the spe­cial sort of em dashes AI does­n’t pro­duce by de­fault.

But also, maybe con­sider not giv­ing a fuck about what some­one on the in­ter­net think? Are you re­ally that eas­ily swayed by a ran­do’s opin­ion? You know what, we can play this game the other way. If your writ­ing does­n’t con­tain an em dash, I ain’t read­ing it. From now on you are obliged to use em dashes for my plea­sure.

PS: if you want per­son­alised ad­vice from me on how to use punc­tu­a­tion cor­rectly, you can now get it for $175/hour.

No posts

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.