10 interesting stories served every morning and every evening.

Introducing Claude Opus 5

www.anthropic.com

Claude Opus 5 is avail­able to­day. It’s a thought­ful and proac­tive model that comes close to the fron­tier in­tel­li­gence of Claude Fable 5 at half the price.

On cod­ing and knowl­edge work eval­u­a­tions like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it re­mains be­hind Mythos 5 on cy­ber­se­cu­rity tasks.

Opus 5 is de­signed to be used every day: it works more ef­fi­ciently than other mod­els. It’s the new de­fault model on Claude Max, and the strongest model on Claude Pro.

Performance and cost-ef­fec­tive­ness

Claude Opus 5 pro­vides greatly im­proved per­for­mance for the same cost as its pre­de­ces­sor, Opus 4.8. The charts in this sec­tion show how per­for­mance changes ac­cord­ing to the mod­el’s ef­fort set­ting, which cus­tomers can use to op­ti­mize for in­tel­li­gence or con­serve to­kens for faster and cheaper re­sults.

Opus 5 ex­cels on valu­able soft­ware en­gi­neer­ing tasks. For ex­am­ple, on Frontier-Bench v0.1, Opus 5 sur­passes all other mod­els, and more than dou­bles Opus 4.8’s per­for­mance at a lower cost per task. On CursorBench 3.2, at max ef­fort, the model per­forms within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater per­for­mance at a given cost than all other mod­els on high, xhigh, and max ef­fort.

We see sim­i­lar re­sults on knowl­edge work and prob­lem-solv­ing tasks. For ex­am­ple:

On ARC-AGI 3, an eval­u­a­tion where the model has to solve novel prob­lems, Opus 5’s score is three times as high as the next-best model.

On Zapier AutomationBench, which mea­sures whether mod­els can com­plete busi­ness tasks from start to fin­ish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its low­est ef­fort set­ting, Opus 5 passes more tasks than any other model.

On OSWorld 2.0, a com­puter use bench­mark, Opus 5 out­per­forms every other model at any given cost, sur­pass­ing Fable 5’s best re­sult at just over a third of the cost.

It’s also our best and most cost-ef­fi­cient model on sev­eral re­lated eval­u­a­tions:

Opus 5 is a mean­ing­ful im­prove­ment over Opus 4.8 for sci­en­tific re­search. It shows bet­ter per­for­mance than Opus 4.8 on every one of our life sci­ences eval­u­a­tions, which cover top­ics in­clud­ing struc­tural bi­ol­ogy, or­ganic chem­istry, and bioin­for­mat­ics. Its im­prove­ments are most no­table on or­ganic chem­istry tasks, like in­fer­ring mol­e­c­u­lar struc­tures from spec­troscopy data (it scores 10.2 per­cent­age points higher than Opus 4.8 on our in­ter­nal bench­mark), and on pro­tein-re­lated tasks like pre­dict­ing how vari­a­tions in a pro­tein’s se­quence af­fect how it func­tions (here, it scores 7.7 per­cent­age points higher).

Finally, Opus 5 is ca­pa­ble of pro­duc­ing much stronger vi­sual out­puts:

Working with Claude Opus 5

Claude Opus 5 is much stronger at ver­i­fy­ing its work and it­er­at­ing care­fully un­til it suc­ceeds. In eval­u­a­tions and early-ac­cess test­ing, we and our users found many ex­am­ples of Opus 5’s agency and thor­ough­ness:

On one Frontier-Bench task, Opus 5 was given a draw­ing of a ma­chine part and asked to write code to re­build it as a 3D FreeCAD model. However, in this task, the model was in­ten­tion­ally given no way to di­rectly view the draw­ing. Opus 5 re­sponded by writ­ing its own com­puter vi­sion pipeline to pull the geom­e­try from the raw pix­els, then re­con­structed the full ma­chine part. It suc­ceeded in do­ing so re­peat­edly; no com­pet­ing model with the same setup could solve it af­ter five at­tempts.

Given a real bug in a pop­u­lar open-source pack­age man­ager, Opus 5 found the root cause and fixed an edge case that the com­mu­ni­ty’s patch had missed. A com­pet­ing model fixed only the sur­face symp­tom (not the un­der­ly­ing cause), then re­ported the bug re­solved.

An en­gi­neer at a trad­ing firm used Opus 5 to build a mar­ket data feed for a new ex­change in a sin­gle ses­sion. Previous mod­els could not com­plete this task at all, even given ex­ten­sive plans from the en­gi­neer. Finding no live feed to val­i­date against, Opus 5 even built its own test har­ness to check that its code parsed the ex­change’s data cor­rectly.

Below are fur­ther re­ports from our early-ac­cess cus­tomers on their ex­pe­ri­ence of work­ing with Opus 5:

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

On FrontierCode 1.1, Claude Opus 5 ap­proaches Fable-level per­for­mance at half the cost. Within Devin, it also shows par­tic­u­lar strength on dif­fi­cult de­bug­ging and root-cause analy­sis tasks.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 de­liv­ers near Fable 5 in­tel­li­gence at Opus speed and cost. On CursorBench it’s just un­der Fable 5 and has many of the same be­hav­iors. We are ex­cited to see how de­vel­op­ers use it in Cursor.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

Claude Opus 5 topped Zapier’s AutomationBench leader­board with­out spend­ing more to­kens than prior Claude mod­els. It took a raw ac­count-health work­book and ran a full churn-pre­ven­tion se­quence end to end: flag­ging at-risk ac­counts, alert­ing the right owner, and sum­ma­riz­ing for re­ten­tion ops. Previous mod­els did­n’t pass; Opus 5 hit 100%.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

On our ge­nomics analy­sis work, Claude Opus 5 be­haves more like a care­ful sci­en­tist than any model we’ve run. It reaches for the right sta­tis­ti­cal tests to rule out con­founders, cross-checks its own re­sults by in­de­pen­dent meth­ods, and stays on track through long multi-step analy­ses.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 came out ahead of every model in its fam­ily on our in­ter­nal evals. It is­n’t just bet­ter on our hard­est agen­tic cod­ing tasks, up 22% over Opus 4.7, it’s stead­ier, with far less vari­ance run to run. For the mil­lions of builders on Lovable, that con­sis­tency is the whole game. Reliable re­sults, build af­ter build.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

Claude Opus 5 is the biggest leap in the Opus fam­ily since 4.5. On the same full-stack app builds, the front end shows it first: the best an­i­ma­tions, games, and 3D work we have seen from an Opus model.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

We’re lov­ing Claude Opus 5. For the kind of open-ended an­a­lyt­i­cal work our agent han­dles, it’s a strict up­grade over Opus 4.8, and the gains are biggest ex­actly where it mat­ters: the harder, vaguer tasks. Responses are clearer and more con­cise, and we see im­proved ef­fi­ciency at higher ef­fort lev­els too.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 is a strik­ing im­prove­ment over Opus 4.8 for the fi­nan­cial re­search work­flows our an­a­lysts run every day. It stands out on nu­mer­i­cal rea­son­ing, table work, and sharper crit­i­cal think­ing where pre­ci­sion mat­ters.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 de­liv­ers the in­dus­try in­tel­li­gence and ac­cu­racy that is es­sen­tial for the analy­sis of spe­cial­ized en­ter­prise con­tent. Box found that Opus 5 out­per­forms Opus 4.8 by 8% and de­liv­ers no­table per­for­mance gains in the data analy­sis (11% im­prove­ment) and due dili­gence (17% im­prove­ment) work­flows that tech­nol­ogy, health­care, and pub­lic sec­tor or­ga­ni­za­tions rely on daily.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 is a clear gen­er­a­tional step up from Opus 4.8. Over one week­end I gave it a chief-of-staff role over my dev en­vi­ron­ments: it built its own mon­i­tor, drove each box, and pulled me in only for the judg­ment calls.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

Claude Opus 5 made large scale changes across our Fundamental Research Assistant code­base, adapt­ing to feed­back through­out an agen­tic work­flow and ex­plain­ing its rea­son­ing more clearly than any model we’ve used. It han­dled work we would nor­mally have bro­ken into much smaller pieces.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

On some of our hard­est fi­nan­cial-mod­el­ing tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both ac­cu­racy and ef­fi­ciency. Its per­for­mance floor is ma­te­ri­ally higher, es­pe­cially on deep fi­nance do­main logic. Across ef­fort lev­els it av­er­aged 9 per­cent­age points higher ac­cu­racy with a third fewer turns and tool calls and 60% less time.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 checks its own work the way a real fron­tend de­vel­oper would. On our bench­mark it opened its pages in a browser at desk­top and phone widths, caught a prod­uct hid­den be­low the mo­bile fold and an off-screen check­out but­ton, and fixed both be­fore hand­ing the work back.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5 is a clear step up in per­for­mance on le­gal agent work com­pared to prior Opus mod­els, and we saw the biggest gains in prac­tice ar­eas like cor­po­rate gov­er­nance and ar­bi­tra­tion. We were also im­pressed with Opus 5’s abil­ity to main­tain qual­ity at lower rea­son­ing lev­els, achiev­ing sim­i­lar per­for­mance while gen­er­at­ing 26% fewer to­kens on av­er­age com­pared to Opus 4.8 at max rea­son­ing.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s biggest gains for us are on longer-hori­zon work: build­ing a full deck, then re­vis­ing it. Artifact qual­ity is what de­cides which model we ship, and this is the clear­est step up we’ve seen — bet­ter vi­sual un­der­stand­ing, cleaner for­mat­ting, fewer slide is­sues.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

Claude Opus 5’s judg­ment is what stands out. Handing off a PR, it does­n’t rush to pub­lish: it ver­i­fies the branches, checks the tem­plate, and thinks through test im­pli­ca­tions so the hand­off is clean. The older mod­els tended to jump ahead and get caught on our checks.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

During a rearchi­tect­ing ses­sion, Claude Opus 5 pushed back on a de­sign I pro­posed, and it did­n’t fold when I in­sisted. Instead, it ex­plained ex­actly what was valu­able in my idea, nar­rowed its ob­jec­tion to a sin­gle de­sign ques­tion, and pro­posed a com­pro­mise that kept the good part while fix­ing the flaw. That’s the kind of judg­ment that lets us trust it with less over­sight.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

On first-turn red­lines, Claude Opus 5 scored the high­est of any model we tested, nearly dou­ble Opus 4.8. Commenting is bet­ter too: on NDAs it gets to the red­line in less time and with fewer passes, with ac­cu­racy main­tained or bet­ter.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger haz­ard spot­ter on sub­tle, code­base-spe­cific is­sues. We’re adopt­ing it for pro­duc­tion work­loads.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

We will def­i­nitely mi­grate a num­ber of use cases in Cosmos, our uni­fied agent plat­form. We’re look­ing for­ward to in­creas­ingly us­ing Claude Opus 5 for code re­view, and I am con­fi­dent in say­ing we would rather peo­ple be us­ing Opus 5 than Opus 4.8.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

What stands out about Claude Opus 5 is judg­ment. It thinks harder be­fore it writes a sin­gle line, catches its own log­i­cal faults dur­ing plan­ning rather than af­ter the fact, and rea­sons about why an an­swer is right, not just whether it works. It’s the clear­est jump in prob­lem-solv­ing we’ve seen from one Claude model to the next, and we’re look­ing for­ward to see­ing it adopted in JetBrains IDEs.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

Claude Opus 5 is the strongest Opus model we’ve tested on our trad­ing bench­mark, and it gets there us­ing roughly a sev­enth of the rea­son­ing to­kens and un­der half the la­tency of Opus 4.8. Better an­swers at a frac­tion of the com­pute.

01 /

22

Alignment and safety

Alignment. During pre-de­ploy­ment test­ing, our au­to­mated be­hav­ioral au­dit found Opus 5 to be our most aligned model to date (as shown in the graph be­low). It ad­heres to Claude’s Constitution bet­ter than Opus 4.8, Sonnet 5, or Fable 5; ex­hibits the low­est rates of de­cep­tive be­hav­ior; and is the least sus­cep­ti­ble to be­ing tricked into mis­use. It’s also our safest model yet in terms of avoid­ing reck­less ac­tions that could have hard-to-re­verse side ef­fects.

Safety. Opus 5 does not ad­vance the fron­tier in risky, dual-use ca­pa­bil­i­ties. In rig­or­ous eval­u­a­tions con­ducted along­side pri­vate-sec­tor and gov­ern­ment part­ners, we found it re­mains be­hind Mythos 5 in both bi­ol­ogy re­search and of­fen­sive cy­ber­se­cu­rity. More in­for­ma­tion about these eval­u­a­tions can be found in our System Card.

As with its pre­de­ces­sor, Opus 4.8, we’ve in­ten­tion­ally avoided train­ing Opus 5 on cy­ber tasks. The model has nev­er­the­less im­proved sub­stan­tially on these tasks as a re­sult of be­com­ing more gen­er­ally ca­pa­ble, and it comes close to Mythos 5 at find­ing cy­ber­se­cu­rity vul­ner­a­bil­i­ties. However, it re­mains sub­stan­tially be­hind Mythos 5 on the ex­ploita­tion of those vul­ner­a­bil­i­ties—that is, in turn­ing vul­ner­a­bil­i­ties into ma­te­r­ial cy­ber threats.

This is il­lus­trated by Opus 5’s per­for­mance on OSS-Fuzz, an eval­u­a­tion we’ve de­vel­oped to as­sess how well mod­els can find and then ex­ploit vul­ner­a­bil­i­ties with­out ex­ten­sive hu­man guid­ance. Although Mythos 5 and Opus 5 iden­tify vul­ner­a­bil­i­ties with sim­i­lar suc­cess, Opus 5’s score on the de­vel­op­ment of ex­ploits is far be­hind that of Mythos 5.

Safeguards for Opus 5

Claude Opus 5’s safe­guards are de­signed to al­low ben­e­fi­cial uses of the model in both cy­ber­se­cu­rity and bi­ol­ogy. They are sim­i­lar to those we ap­plied to Opus 4.8, with the ex­cep­tion of some stronger guardrails on a nar­row range of cy­ber tasks.

Cybersecurity. Opus 5’s cy­ber clas­si­fiers are pro­por­tion­ally less re­stric­tive than those on Fable 5. They al­low Opus 5 to find vul­ner­a­bil­i­ties in source code, but block binary-based” vul­ner­a­bil­ity scan­ning (a method more likely to be as­so­ci­ated with ma­li­cious ac­tors), pen­e­tra­tion test­ing, and ex­ploit gen­er­a­tion.

Based on our test­ing, we ex­pect the clas­si­fiers to in­ter­vene around 85% less of­ten than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged re­quests will fall back to Opus 4.8 by de­fault. Fallbacks to Opus 4.8 can also be en­abled on the API.

Our Cyber Verification Program (CVP) fa­cil­i­tates cy­ber­se­cu­rity work that would oth­er­wise be im­peded by the mod­el’s safe­guards. Enterprises and re­searchers who are al­ready part of the CVP have im­me­di­ate ac­cess to a ver­sion of Opus 5 with fewer se­cu­rity re­stric­tions.

Biology. Since Opus 5 has a sim­i­lar suite of safe­guards to Opus 4.8, it is now our most ca­pa­ble gen­er­ally avail­able model for sci­en­tific re­search. Nevertheless, the model still shows im­por­tant lim­i­ta­tions on long-run­ning, au­tonomous re­search tasks, which is where we ex­pect AI mod­els to pose the most sub­stan­tial bi­ol­ogy-re­lated risks. (Mythos 5 re­mains the stronger model for this type of bi­o­log­i­cal work.) As part of this launch, bi­ol­ogy-re­lated re­quests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.

Getting started

Claude Opus 5 is avail­able to­day on all plat­forms, priced at $5 per mil­lion in­put to­kens and $25 per mil­lion out­put to­kens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.

It’s also of­fered in Fast mode, where it runs around 2.5 times the de­fault speed. As with Opus 4.8, Fast mode is avail­able at twice Opus 5’s base price on the Claude Platform and through us­age cred­its in Claude Code.

Alongside Opus 5, we’re re­leas­ing two up­dates in beta:

Mid-conversation tool changes on the Claude Platform. Within a con­ver­sa­tion, de­vel­op­ers can now change which tools Claude can use with­out in­val­i­dat­ing the prompt cache.

Automatic fall­backs on the API. Users can now choose to have re­quests that are flagged by our safety clas­si­fiers on Opus 5 (or Fable 5) au­to­mat­i­cally route to an­other model. With au­to­matic fall­backs on, API re­quests al­ways route to the best avail­able model by de­fault rather than be­ing blocked.

Consistent with prior Opus mod­els, Opus 5 does not have data re­ten­tion re­quire­ments for gen­eral ac­cess.

For more guid­ance on how to get the best out of Opus 5, see our prompt­ing guide.

Footnotes

Frontier-Bench v0.1, Effort plot: These re­sults are from an in­ter­nal run of Frontier-Bench v0.1, on the mini-SWE-agent har­ness and a GKE back­end, mean re­ward over 5 at­tempts per task. Opus 4.8 served as fall­back on safety-clas­si­fier re­fusals for Opus 5 and Fable 5.

Related con­tent

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

Ask Claude about the Anthropic Economic Index

We’re launch­ing the Anthropic Economic Index con­nec­tor for Claude, which lets any­one ex­plore real data about AI and work.

Read more

Anthropic is do­nat­ing an­other $20 mil­lion to Public First Action

Anthropic is con­tribut­ing an ad­di­tional $20 mil­lion to Public First Action, bring­ing our to­tal sup­port to $40 mil­lion.

Read more

It's getting harder to focus every day

glyphack.com

I’m feel­ing it right now. I had to set a timer for 15 minute on my com­puter and block all dis­trac­tions to write this. If I did­n’t force my­self to fo­cus I would eas­ily get dis­tracted by some­thing af­ter few min­utes. Even when I’m do­ing things that I’ve been wait­ing to do it, I still feel the urge to do some­thing else.

I don’t know how and when this hap­pened. During the last few years I was al­ways study­ing, work­ing, and do­ing open source. And ac­tu­ally got stuff done. Doing all of those at the same time re­quires pay­ing at­ten­tion to what I wanted to do and ig­nore the noise.

Nowadays, I’m lucky if I get 1 hour of fo­cused time. Just to be clear, my goal is not to work 90 hours a week or any­thing crazy. That is not pos­si­ble for me. I just want the hour I spend pro­gram­ming, learn­ing, or writ­ing to be just one ac­tiv­ity. But in­stead I spend 10 minute on some­thing then I get dis­tracted, and try to fo­cus again.

Ideally, I want to be able to plan to work on some­thing for long hours with­out any dis­trac­tion. After that time get back on­line check for mes­sages and other things.

The dis­trac­tions are not one par­tic­u­lar thing. A few ex­am­ples:

I want to do some­thing and I re­mem­ber there’s a post re­lated to this. I go to find it and in be­tween I click on some links and end up read­ing some­thing com­pletely un­re­lated.

I am wait­ing for some­thing then I go browse the web and I get dis­tracted.

I’m work­ing on some­thing and I face a chal­lenge I have to think for 10 min­utes. I get up to get some wa­ter and check my phone along the way and get dis­tracted. Sometimes I get dis­tracted by mak­ing the bed.

It feels like my brain finds a way to do some­thing else and avoid painful sit­u­a­tions like bore­dom or hard work.

It re­minds me of when I was in high school talk­ing to a friend about how lay­ing in bed with your phone can kill hours with­out you notic­ing it. It was circa 2015, back then at­ten­tion hun­gry apps were less pow­er­ful but still lure a teenager’s mind for few hours. I learned that these apps should be used very care­fully. At that time I made a de­ci­sion to never have a charger near my bed. Back then it was mostly about my phone, be­cause when I was on the com­puter I was ei­ther read­ing, or pro­gram­ming, or play­ing a game. Even if I was­n’t in the mood to think, I played a strate­gic game or chess. These ac­tiv­i­ties ex­er­cise the mind and are fun. All of them were an in­ten­tional ac­tiv­ity. I did­n’t do any pas­sive ac­tiv­i­ties like brows­ing or chat­ting on my com­puter.

Later I dis­cov­ered HackerNoon and Medium, it was the first web­site that I was brows­ing when­ever I was bored be­hind the com­puter. This meant that I had a way to get out bore­dom eas­ily. It used to have some high qual­ity con­tent. It in­spired me to do some pro­jects and learn more pro­gram­ming. I found chan­nels like CSDojo there. Nowadays I don’t even open them. They are filled with slop or click bait ar­ti­cles, prob­a­bly be­cause of mon­e­ti­za­tion in­cen­tives. Then hack­ernews, and YouTube and oth­ers took their place. I also found some good peo­ple and blogs along the way. I learned to keep a read­ing list from peo­ple I like to read when I’m on the bus.

I slowly found more ac­tiv­i­ties for when I’m bored. This made it harder to fo­cus on hard things for me. What kept me on track was that there was no way out. I had to work out some al­ge­bra prob­lems. I kept my­self to a very high bar of un­der­stand­ing what I do. And I did every­thing in LaTeX so I could­n’t copy from some­one else. I was the only one typ­ing them.

The first time I saw peo­ple not putting the ef­fort and still get the re­ward for it was at work. You might won­der, aren’t peo­ple in the uni­ver­sity con­stantly cheat­ing and copy­ing home­work, and get good grades? Well yes, but when you talked with some­one in that group it was clear that he is clue­less about the sub­ject. A good grade did­n’t mean much to me at that time. And as a stu­dent cheat­ing does not get you that far.

At work it is dif­fer­ent. My days are mostly spend meet­ings and over chat. The bal­ance be­tween ac­tual work and bull­shit is skewed. I saw peo­ple who were barely do­ing any work and just talk are suc­cess­ful. As long as peo­ple give 10% of their at­ten­tion to work they are con­sid­ered fine in most en­vi­ron­ments. Previously I did an ex­per­i­ment to track my time and found out that I spent 8 hours chat­ting on slack in a week.

I don’t care how em­ploy­ers want to shat­ter em­ploy­ee’s fo­cus. But this made me get used to dis­trac­tions when I’m pro­gram­ming. I’m try­ing to undo this dam­age.

The next big change is more us­age of LLMs. I find my­self in this sit­u­a­tion too many times, where I out­source some­thing to an LLM and then I start work­ing on some­thing else. And while do­ing this I keep think­ing about what it’s do­ing. Or when I’m think­ing about some­thing I start chat­ting with an LLM about my idea and in­stead of get­ting started on some­thing I’m in re­search mode only for hours.

It’s good that I’m able to ask some­thing else to re­search some topic for me. But the pro­duc­tiv­ity only comes if I can move on to some­thing else and not think about it. I don’t have any no­ti­fi­ca­tions turned on so it does­n’t dis­tract me. But I still find my­self think­ing about what I just asked it to do and I can­not fo­cus on some­thing else.

At the same time if I’m spend­ing the time in­ter­ac­tively with an LLM I feel slow. I have to wait for the re­sponse and I have to cor­rect every re­sponse com­ing out. The best use of AI seems to be out­sourc­ing what they can do end to end with­out er­ror.

Why do I keep do­ing this? Presumably be­cause it’s easy and fast, and pro­duc­tive. If I re­al­ize I have to do some­thing I can write it down to do it later or I can just ask the LLM to do it. The prob­lem only shows up when I start do­ing so many things at once be­cause it’s ac­tu­ally do­ing the thing. Then I have mul­ti­ple things on my mind and can’t fo­cus re­ally. I have to check on it and guide it in the right di­rec­tion every now and then.

And this is over­stim­u­lat­ing, in a way that do­ing some­thing with­out LLM some­times is bor­ing. You don’t see the re­sults as fast. Which makes fo­cus­ing harder.

So how can I re­gain my abil­ity to fo­cus? Sometimes I live stream what I’m do­ing just be­cause with a cam­era I can­not es­cape from hard chal­lenges by grab­bing my phone. I used to co-work with my friends over dis­cord. Unfortunately it’s not pos­si­ble any­more be­cause peo­ple in Iran can­not have a sta­ble in­ter­net con­nec­tion nowa­days.

It’s in­cred­i­bly hard to com­mit to some­thing for a long pe­riod of time. If what I’m do­ing is go­ing to take mul­ti­ple days to have a re­sult I have less mo­ti­va­tions to do it. Meanwhile, vibe cod­ing small util­ity scripts is fun I keep do­ing it when­ever I see a fric­tion. Whenever I am stuck I can throw my prob­lem into it and wait strengthen this habit of wait­ing for an an­swer from some­one as op­posed to work through prob­lems. And they are fast in get­ting back the re­sults. So next time I have to read a pa­per to un­der­stand the sub­ject I will be more re­luc­tant be­cause I can get a faster re­sult through them.

I’m chang­ing some habits to re­place the cur­rent ones. If I don’t feel mo­ti­vated enough to do any­thing I get up and pick up a book to read. I’m keep­ing a Garden in my bal­cony is that when I’m tired I can move the soil around and plant some pots and prune plants. I find this to be a less ad­dic­tive than say, watch­ing a movie. When the mo­ti­va­tion comes back I can stop it eas­ily.

My goal was not to find an an­swer for this prob­lem. I wanted to see what’s go­ing on and why I am not do­ing any­thing in­con­se­quen­tial in the past few months. Anyway the timer I set to write this re­ally helped. I spent a lot more time to write this but it gave me the ini­tial mo­ti­va­tion to write.

98.css

jdan.github.io

A de­sign sys­tem for build­ing faith­ful recre­ations of old UIs.

Intro

98.css is a CSS li­brary for build­ing in­ter­faces that look like Windows 98. See more on GitHub.

My First VB4 Program

Hello, world!

This li­brary re­lies on the us­age of se­man­tic HTML. To make a but­ton, you’ll need to use a <button>. Input el­e­ments re­quire la­bels. Icon but­tons rely on aria-la­bel. This page will guide you through that process, but ac­ces­si­bil­ity is a pri­mary goal of this pro­ject.

You can over­ride many of the styles of your el­e­ments while main­tain­ing the ap­pear­ance pro­vided by this li­brary. Need more padding on your but­tons? Go for it. Need to add some color to your in­put la­bels? Be our guest.

This li­brary does not con­tain any JavaScript, it merely styles your HTML with some CSS. This means 98.css is com­pat­i­ble with your fron­tend frame­work of choice.

Here is an ex­am­ple of 98.css used with React, and an ex­am­ple with vanilla JavaScript. The fastest way to use 98.css is to im­port it from unpkg.

<link rel=“stylesheet” href=“https://​unpkg.com/​98.css >

You can in­stall 98.css from the GitHub re­leases page, or from npm.

npm in­stall 98.css

Components

Button

A com­mand but­ton, also re­ferred to as a push but­ton, is a con­trol that causes the ap­pli­ca­tion to per­form some ac­tion when the user clicks it.

A stan­dard but­ton mea­sures 75px wide and 23px tall, with a raised outer and in­ner bor­der. They are given 12px of hor­i­zon­tal padding by de­fault.

<button>Click me</​but­ton> <input type=“sub­mit” /> <input type=“re­set” />

You can add the class de­fault to any but­ton to ap­ply ad­di­tional styling, use­ful when com­mu­ni­cat­ing to the user what de­fault ac­tion would hap­pen in the ac­tive win­dow if the Enter key was pressed on Windows 98.

<button class=“de­fault”>OK</​but­ton>

When but­tons are clicked, the raised bor­ders be­come sunken. The fol­low­ing but­ton is sim­u­lated to be in the pressed (active) state.

<button>I am be­ing pressed</​but­ton>

Disabled but­tons main­tain the same raised bor­der, but have a washed out” ap­pear­ance in their la­bel.

<button dis­abled>I can­not be clicked</​but­ton>

Button fo­cus is com­mu­ni­cated with a dot­ted bor­der, set 4px within the con­tents of the but­ton. The fol­low­ing ex­am­ple is sim­u­lated to be fo­cused.

<button>I am fo­cused</​but­ton>

Checkbox

A check box rep­re­sents an in­de­pen­dent or non-ex­clu­sive choice.

Checkboxes are rep­re­sented with a sunken panel, pop­u­lated with a check” icon when se­lected, next to a la­bel in­di­cat­ing the choice.

Note: You must in­clude a cor­re­spond­ing la­bel af­ter your check­box, us­ing the <label> el­e­ment with a for at­tribute pointed at the id of your in­put. This en­sures the check­box is easy to use with as­sis­tive tech­nolo­gies, on top of en­sur­ing a good user ex­pe­ri­ence for all (navigating with the tab key, be­ing able to click the en­tire la­bel to se­lect the box).

This is a check­box

<input type=“check­box” id=“ex­am­ple1″> <label for=“ex­am­ple1”>This is a check­box</​la­bel>

Checkboxes can be se­lected and dis­abled with the stan­dard checked and dis­abled at­trib­utes.

When group­ing in­puts, wrap each in­put in a con­tainer with the field-row class. This en­sures a con­sis­tent spac­ing be­tween in­puts.

I am checked

I am in­ac­tive

I am in­ac­tive but still checked

<div class=“field-row”> <input checked type=“check­box” id=“ex­am­ple2”> <label for=“ex­am­ple2″>I am checked</​la­bel> </div> <div class=“field-row”> <input dis­abled type=“check­box” id=“ex­am­ple3”> <label for=“ex­am­ple3″>I am in­ac­tive</​la­bel> </div> <div class=“field-row”> <input checked dis­abled type=“check­box” id=“ex­am­ple4”> <label for=“ex­am­ple4″>I am in­ac­tive but still checked</​la­bel> </div>

OptionButton

An op­tion but­ton, also re­ferred to as a ra­dio but­ton, rep­re­sents a sin­gle choice within a lim­ited set of mu­tu­ally ex­clu­sive choices. That is, the user can choose only one set of op­tions.

Option but­tons can be used via the ra­dio type on an in­put el­e­ment.

Option but­tons can be grouped by spec­i­fy­ing a shared name at­tribute on each in­put. Just as be­fore: when group­ing in­puts, wrap each in­put in a con­tainer with the field-row class to en­sure a con­sis­tent spac­ing be­tween in­puts.

Yes

No

<div class=“field-row”> <input id=“ra­dio5″ type=“ra­dio” name=“first-ex­am­ple”> <label for=“ra­dio5”>Yes</​la­bel> </div> <div class=“field-row”> <input id=“ra­dio6” type=“ra­dio” name=“first-ex­am­ple”> <label for=“ra­dio6″>No</​la­bel> </div>

Option but­tons can also be checked and dis­abled with their cor­re­spond­ing HTML at­trib­utes.

Peanut but­ter should be smooth

I un­der­stand why peo­ple like crunchy peanut but­ter

Crunchy peanut but­ter is good

<div class=“field-row”> <input id=“ra­dio7″ type=“ra­dio” name=“sec­ond-ex­am­ple”> <label for=“ra­dio7”>Peanut but­ter should be smooth</​la­bel> </div> <div class=“field-row”> <input checked dis­abled id=“ra­dio8” type=“ra­dio” name=“sec­ond-ex­am­ple”> <label for=“ra­dio8″>I un­der­stand why peo­ple like crunchy peanut but­ter</​la­bel> </div> <div class=“field-row”> <input dis­abled id=“ra­dio9″ type=“ra­dio” name=“sec­ond-ex­am­ple”> <label for=“ra­dio9”>Crunchy peanut but­ter is good</​la­bel> </div>

GroupBox

A group box is a spe­cial con­trol you can use to or­ga­nize a set of con­trols. A group box is a rec­tan­gu­lar frame with an op­tional la­bel that sur­rounds a set of con­trols.

A group box can be used by wrap­ping your el­e­ments with the field­set tag. It con­tains a sunken outer bor­der and a raised in­ner bor­der, re­sem­bling an en­graved box around your con­trols.

<fieldset> <div class=“field-row”>Se­lect one:</​div> <div class=“field-row”> <input id=“ra­dio10” type=“ra­dio” name=“field­set-ex­am­ple”> <label for=“ra­dio10″>Din­ers</​la­bel> </div> <div class=“field-row”> <input id=“ra­dio11″ type=“ra­dio” name=“field­set-ex­am­ple”> <label for=“ra­dio11”>Drive-Ins</​la­bel> </div> <div class=“field-row”> <input id=“ra­dio12” type=“ra­dio” name=“field­set-ex­am­ple”> <label for=“ra­dio12″>Dives</​la­bel> </div> </fieldset>

You can pro­vide your group with a la­bel by plac­ing a leg­end el­e­ment within the field­set.

<fieldset> <legend>Today’s mood</​leg­end> <div class=“field-row”> <input id=“ra­dio13″ type=“ra­dio” name=“field­set-ex­am­ple2″> <label for=“ra­dio13”>Claire Saffitz</label> </div> <div class=“field-row”> <input id=“ra­dio14” type=“ra­dio” name=“field­set-ex­am­ple2”> <label for=“ra­dio14″>Brad Leone</label> </div> <div class=“field-row”> <input id=“ra­dio15″ type=“ra­dio” name=“field­set-ex­am­ple2″> <label for=“ra­dio15”>Chris Morocco</label> </div> <div class=“field-row”> <input id=“ra­dio16” type=“ra­dio” name=“field­set-ex­am­ple2”> <label for=“ra­dio16″>Carla Lalli Music</label> </div> </fieldset>

TextBox

A text box (also re­ferred to as an edit con­trol) is a rec­tan­gu­lar con­trol where the user en­ters or ed­its text. It can be de­fined to sup­port a sin­gle line or mul­ti­ple lines of text.

Text boxes can ren­dered by spec­i­fy­ing a text type on an in­put el­e­ment. As with check­boxes and ra­dio but­tons, you should pro­vide a cor­re­spond­ing la­bel with a prop­erly set for at­tribute, and wrap both in a con­tainer with the field-row class.

Occupation

<div class=“field-row”> <label for=“tex­t17″>Oc­cu­pa­tion</​la­bel> <input id=“tex­t17” type=“text” /> </div>

Additionally, you can make use of the field-row-stacked class to po­si­tion your la­bel above the in­put in­stead of be­side it.

Address (Line 1)

Address (Line 2)

<div class=“field-row-stacked” style=“width: 200px”> <label for=“tex­t18”>Ad­dress (Line 1)</label> <input id=“tex­t18″ type=“text” /> </div> <div class=“field-row-stacked” style=“width: 200px”> <label for=“tex­t19″>Ad­dress (Line 2)</label> <input id=“tex­t19” type=“text” /> </div>

To sup­port mul­ti­ple lines in the user’s in­put, use the textarea el­e­ment in­stead.

Additional notes

<div class=“field-row-stacked” style=“width: 200px”> <label for=“tex­t20”>Ad­di­tional notes</​la­bel> <textarea id=“tex­t20″ rows=“8”></​textarea> </div>

Text boxes can also be dis­abled and have value with their cor­re­spond­ing HTML at­trib­utes.

Favorite color

<div class=“field-row”> <label for=“tex­t21″>Fa­vorite color</​la­bel> <input id=“tex­t21” dis­abled type=“text” value=“Win­dows Green”/> </div>

Slider

A slider, some­times called a track­bar con­trol, con­sists of a bar that de­fines the ex­tent or range of the ad­just­ment and an in­di­ca­tor that shows the cur­rent value for the con­trol…

Sliders can ren­dered by spec­i­fy­ing a range type on an in­put el­e­ment.

Volume: Low

High

<div class=“field-row” style=“width: 300px”> <label for=“range22”>Vol­ume:</​la­bel> <label for=“range23″>Low</​la­bel> <input id=“range23” type=“range” min=“1” max=“11″ value=“5” /> <label for=“range24″>High</​la­bel> </div>

You can make use of the has-box-in­di­ca­tor class re­place the de­fault in­di­ca­tor with a box in­di­ca­tor, fur­ther­more the slider can be wrapped with a div us­ing is-ver­ti­cal to dis­play the in­put ver­ti­cally.

Note: To change the length of a ver­ti­cal slider, the in­put width and div height.

Cowbell

<div class=“field-row”> <label for=“range25″>Cow­bell</​la­bel> <div class=“is-ver­ti­cal”> <input id=“range25″ class=“has-box-in­di­ca­tor” type=“range” min=“1” max=“3″ step=“1” value=“2″ /> </div> </div>

Dropdown

A drop-down list box al­lows the se­lec­tion of only a sin­gle item from a list. In its closed state, the con­trol dis­plays the cur­rent value for the con­trol. The user opens the list to change the value.

Dropdowns can be ren­dered by us­ing the se­lect and op­tion el­e­ments.

<select> <option>5 - Incredible!</option> <option>4 - Great!</option> <option>3 - Pretty good</​op­tion> <option>2 - Not so great</​op­tion> <option>1 - Unfortunate</option> </select>

By de­fault, the first op­tion will be se­lected. You can change this by giv­ing one of your op­tion el­e­ments the se­lected at­tribute.

<select> <option>5 - Incredible!</option> <option>4 - Great!</option> <option se­lected>3 - Pretty good</​op­tion> <option>2 - Not so great</​op­tion> <option>1 - Unfortunate</option> </select>

Window

The fol­low­ing com­po­nents il­lus­trate how to build com­plete win­dows us­ing 98.css.

Title Bar

At the top edge of the win­dow, in­side its bor­der, is the ti­tle bar (also ref­fered to as the cap­tion or cap­tion bar), which ex­tends across the width of the win­dow. The ti­tle bar iden­ti­fies the con­tents of the win­dow.

Include com­mand but­tons as­so­ci­ated with the com­mon com­mands of the pri­mary win­dow in the ti­tle bar. These but­tons act as short­cuts to spe­cific win­dow com­mands.

You can build a com­plete ti­tle bar by mak­ing use of three classes, ti­tle-bar, ti­tle-bar-text, and ti­tle-bar-con­trols.

A Title Bar

<div class=“ti­tle-bar”> <div class=“ti­tle-bar-text”>A Title Bar</div> <div class=“ti­tle-bar-con­trols”> <button aria-la­bel=“Close”></​but­ton> </div> </div>

We make use of aria-la­bel to ren­der the Close but­ton, to let as­sis­tive tech­nolo­gies know the in­tent of this but­ton. You may also use Minimize”, Maximize”, Restore” and Help” like so:

A Title Bar

A Maximized Title Bar

A Helpful Bar

<div class=“ti­tle-bar”> <div class=“ti­tle-bar-text”>A Title Bar</div> <div class=“ti­tle-bar-con­trols”> <button aria-la­bel=“Min­i­mize”></​but­ton> <button aria-la­bel=“Max­i­mize”></​but­ton> <button aria-la­bel=“Close”></​but­ton> </div> </div>

<br />

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.

bfl.ai

FLUX 3 is now avail­able in Early Access.

FLUX 3 is our new mul­ti­modal foun­da­tion model. It jointly learns from im­ages, videos, and au­dio within a uni­fied ar­chi­tec­ture, be­cause what it needs to learn is not any one of these el­e­ments in iso­la­tion. Instead, a model must learn a rep­re­sen­ta­tion of the world: how ob­jects hold to­gether, how things move, and how events sound.

No sin­gle modal­ity pro­vides a com­plete de­scrip­tion. Each is a pro­jec­tion of the same un­der­ly­ing re­al­ity, cap­tured by dif­fer­ent sen­sors, each of which loses some in­for­ma­tion in the process. Images cap­ture spa­tial struc­tures and re­la­tion­ships at a spe­cific point in time. Videos re­store the di­men­sion of time and re­veal tem­po­ral dy­nam­ics and phys­i­cal laws. Audio re­veals causal re­la­tion­ships be­tween me­chan­i­cal phe­nom­ena and acoustics that vi­sion alone can­not de­tect. Language links these per­cep­tions to goals, ab­strac­tions, and in­struc­tions.

Learn from one and you get a good model of that pro­jec­tion. Learn from all of them at once and their mu­tual con­straints tell you more: the sound has to match the im­pact, the mo­tion has to obey the mass, the fu­ture has to fol­low from the past. The modal­i­ties stop be­ing sep­a­rate and start be­ing ev­i­dence about one un­der­ly­ing re­al­ity.

FLUX 3 is our first model built en­tirely on that prin­ci­ple, and a check­point on our mis­sion to de­velop real-world vi­sual in­tel­li­gence: mod­els that per­ceive, pre­dict, and act across phys­i­cal and dig­i­tal en­vi­ron­ments. Early re­sults in con­tent cre­ation and phys­i­cal AI sug­gest it is the right path.

FLUX 3: One model, mul­ti­ple ca­pa­bil­i­ties.

FLUX 3 builds on Self-Flow, our ap­proach for ef­fi­ciently align­ing mul­ti­modal gen­er­a­tion and un­der­stand­ing within the same un­der­ly­ing ar­chi­tec­ture. Based on this ap­proach, we sig­nif­i­cantly scaled up com­pute and data re­sources to train FLUX 3 across video, im­ages, and au­dio at the same time.

Self-Flow vs. Flow Matching (FM). Left: gen­er­a­tion er­ror (Fréchet dis­tance) per modal­ity, each nor­mal­ized to FM = 100 (lower is bet­ter). Right: suc­cess rate on ma­nip­u­la­tion tasks av­er­aged over four task groups through fine­tun­ing (higher is bet­ter).

Capabilities & Early Evaluations

As a re­sult, FLUX 3 is ca­pa­ble of mix­ing modal­i­ties and gen­er­at­ing im­ages and video+au­dio jointly; both from pure text prompts as well as when pro­vid­ing in­put ref­er­ences such as im­ages and video. We are high­light­ing a few of the mod­el’s key ca­pa­bil­i­ties be­low.

Video

FLUX 3 can cre­ate highly di­verse videos with au­dio up to 20 sec­onds in length in a sin­gle gen­er­a­tion.

Its core ca­pa­bil­i­ties in­clude the fol­low­ing (all out­puts come with na­tive au­dio gen­er­a­tion):

Text-to-video gen­er­a­tion.

Image-to-video gen­er­a­tion, ei­ther con­tin­u­ing from a start­ing frame (“animation”) or us­ing im­ages as vi­sual ref­er­ences.

Video-to-video gen­er­a­tion from a ref­er­ence clip, car­ry­ing cen­tral el­e­ments of a source video - for in­stance the same char­ac­ter - into a new scene or con­text.

Generative video-au­dio con­tin­u­a­tion from in­put video and au­dio.

Keyframe-to-video gen­er­a­tion for con­trolled tran­si­tions be­tween de­fined mo­ments.

Multilingual di­a­logue.

A broad range of vi­sual styles and as­pect ra­tios, ex­tend­ing far be­yond con­ven­tional cin­e­matic out­put.

Agentic chain­ing of in­di­vid­ual clips into longer, multi-shot se­quences.

High style di­ver­sity — FLUX 3 Video eas­ily han­dles ranges of styles from can­did cam­corder footage to an­i­ma­tion and cin­e­mat­ics.

Strong ty­pog­ra­phy gen­er­a­tion and an­i­mated de­signs.

For the pre­lim­i­nary analy­sis be­low, we gen­er­ated 10-second text-to-video clips in 720p with au­dio.

Evaluations are early and we ex­pect fur­ther im­prove­ments

As the model and the har­ness around it are still in de­vel­op­ment, these re­sults are pre­lim­i­nary, and we ex­pect fur­ther im­prove­ments dur­ing the early ac­cess phase. Across early eval­u­a­tions, FLUX 3 was pre­ferred over Grok Imagine Video in up to 69% of com­par­isons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was pre­ferred over Runway Gen-4.5 in 77% of com­par­isons and over Luma Ray 3.2 in 93% of com­par­isons.

While still in de­vel­op­ment, FLUX 3 Video is al­ready par­tic­u­larly strong in cap­tur­ing hu­man fa­cial ex­pres­sions, as­so­ci­at­ing sounds with phys­i­cal events, and mul­ti­lin­gual ca­pa­bil­i­ties. Furthermore, these ca­pa­bil­i­ties can be com­bined to cre­ate se­quences last­ing sev­eral min­utes, where vi­sual ref­er­ences help en­sure that the char­ac­ters re­main con­sis­tent across all scenes.

FLUX 3 Video is now avail­able in Early Access here

Image

FLUX 3 can syn­the­size and edit im­ages in a wide va­ri­ety of styles, as­pect ra­tios, and res­o­lu­tions. In pre­lim­i­nary eval­u­a­tions con­ducted dur­ing mid­train­ing, FLUX 3 al­ready shows a sig­nif­i­cant im­prove­ment over ear­lier ver­sions of FLUX: its abil­ity to han­dle com­plex prompts and text gen­er­a­tion has im­proved sig­nif­i­cantly. The model pro­duces a wide range of out­put styles (see the fol­low­ing sam­ples), and is able to ren­der high-ac­cu­racy text in mul­ti­ple lan­guages.

As with video eval­u­a­tions, these are pre­lim­i­nary re­sults, and we ex­pect fur­ther im­prove­ments be­fore re­lease. We will open up an early ac­cess phase for FLUX 3 Image in the fol­low­ing weeks.

Action

FLUX 3′s world un­der­stand­ing ex­tends to ac­tion pre­dic­tion. We have taken two routes to it: in­te­grat­ing na­tive ac­tion pre­dic­tion into FLUX 3 di­rectly, scal­ing up our ini­tial work in Self-Flow; and us­ing the pre­trained video back­bone as a dy­nam­ics-aware foun­da­tion that spe­cial­ized ac­tion mod­els can be fine­tuned from with lim­ited task-spe­cific data.

For the sec­ond, mimic ro­bot­ics was one of the first part­ners to gain early ac­cess to FLUX 3. Together we de­vel­oped FLUX-mimic, a video-ac­tion model com­bin­ing the FLUX 3 back­bone with mim­ic’s ex­per­tise in ro­bot learn­ing for dex­ter­ous ma­nip­u­la­tion and pro­duc­tion de­ploy­ment. Read our the­sis on why phys­i­cal AI and con­tent cre­ation run on the same foun­da­tion, and how it’s be­ing tested on real pro­duc­tion tasks at Audi.

Launch Plan

Over the next few weeks and months, we will make the fol­low­ing ca­pa­bil­i­ties avail­able, each af­ter an early ac­cess phase for en­sur­ing smooth roll­out, col­lect­ing feed­back and rig­or­ous safety-test­ing. All ca­pa­bil­i­ties are built from the same un­der­ly­ing mul­ti­modal flow match­ing model. These ca­pa­bil­i­ties and mod­els in­clude:

Video and au­dio gen­er­a­tion and edit­ing through APIs and pri­vate weight ac­cess. (“FLUX 3 Video”)

Action pre­dic­tion through se­lected re­search and com­mer­cial part­ners, be­gin­ning with mimic ro­bot­ics (“FLUX-mimic and FLUX 3 Action”)

Image syn­the­sis and edit­ing through APIs and pri­vate weight ac­cess. (“FLUX 3 Image”)

Open-weight ac­cess to a mul­ti­modal back­bone, for con­tent cre­ation (video, au­dio and im­age) and ac­tion pre­dic­tion. (“FLUX 3 Dev”)

We will also re­lease more tech­ni­cal de­tails on the un­der­ly­ing ap­proach.

Request early ac­cess here

What’s next?

We are only be­gin­ning to scratch the sur­face of ver­sa­tile, ca­pa­ble, uni­fied mul­ti­modal mod­els, and what they will en­able. From in­ter­ac­tive im­age & video edit­ing, sim­u­la­tion to com­puter use and phys­i­cal AI, the fron­tier is wide open. While we grad­u­ally roll out these new ca­pa­bil­i­ties, we are al­ready work­ing on the next gen­er­a­tion mod­els. Our goal is to unify per­cep­tual, ac­tion and lan­guage pre­dic­tion in the same uni­fied model.

If you are in­ter­ested in ex­plor­ing and build­ing with FLUX 3, get in touch here. If you are in­ter­ested in con­tribut­ing to our mis­sion, join us! We are hir­ing in Germany and the US.

The Beam Engine

glinscott.github.io

Power from Steam in the Industrial Revolution

By Gary Linscott

This is a beam en­gine. It pro­duced about fif­teen horse­power con­tin­u­ously, roughly as much power as 150 peo­ple. Engines like this turned steam into the power that drove the Industrial Revolution. This ar­ti­cle builds the en­gine up from first prin­ci­ples, us­ing in­ter­ac­tive fig­ures to ex­plore each idea (try ro­tat­ing the en­gine above with two fin­gers, or pinch­ing to zoom in­drag­ging the en­gine above, or zoom­ing with ⌘/Ctrl + scroll). Let’s start our jour­ney through the en­gine with steam.

Steam

Below, we have a pot filled with wa­ter and a fire un­der­neath. As the fire heats the wa­ter, some of it be­gins to boil and turns into steam.

Steam un­der­goes an amaz­ing trans­for­ma­tion: it ex­pands to 1,700 times the vol­ume of the orig­i­nal wa­ter. One cup of wa­ter be­comes roughly 400 litres of steam, enough to fill two bath­tubs. If the steam does­n’t have enough room to ex­pand it will push on all the walls of the con­tainer. This push on every wall is pres­sure, and we will mea­sure it in at­mos­pheres, mul­ti­ples of the or­di­nary pres­sure of the air around us. The steam also presses on the sur­face of the wa­ter, which trans­mits the pres­sure evenly to every­where the wa­ter touches.

Now we need a way to har­ness the prop­er­ties of steam.

Pistons and cylin­ders

A pis­ton is a round disc that fits snugly in­side a cylin­der. Steam pushes on one face of the pis­ton and a rod trans­mits the force else­where. The force de­pends on two things: the pres­sure of the steam and the area of the pis­ton. At a pres­sure dif­fer­ence of one at­mos­phere, each square cen­time­tre of pis­ton pro­vides about one kilo­gram of force.

Early boiler builders did­n’t know how to safely har­ness high-pres­sure steam.1 Instead, to get more force they made the pis­ton wider. Because area grows with the square of the di­am­e­ter, dou­bling the width of a pis­ton gives it four times the area and four times the force at the same pres­sure. This is why early steam en­gines had enor­mous cylin­ders, some­times wide enough for a per­son to stand in­side. In the fig­ure be­low, the boiler pres­sure never changes; try in­creas­ing only the bore un­til the pis­ton can lift the car.

With steam push­ing on our pis­ton, we can do real work. But low-pres­sure steam is not very strong. To move heavy ma­chin­ery, en­gi­neers turned to a sur­pris­ing source: the at­mos­phere.

The weight of air

Air feels weight­less, but only be­cause we are sur­rounded by it. Imagine a col­umn of air one cen­time­tre square, ex­tend­ing from your hand all the way to the top of the at­mos­phere. That col­umn weighs about one kilo­gram, so the at­mos­phere presses on every square cen­time­tre with roughly one kilo­gram of force.

We do not feel this enor­mous pres­sure be­cause the air and fluid in­side us push back at the same pres­sure. But if the pres­sure falls on one side of a sur­face, the pres­sure on the other side re­mains. This is what hap­pens when you drink through a straw. Your mouth low­ers the pres­sure in­side the straw, and the at­mos­phere push­ing on the drink in the cup forces it up­ward.

Otto von Guericke gave a spec­tac­u­lar demon­stra­tion of this ef­fect in 1654. He joined two cop­per hemi­spheres into a sphere about half a me­tre across and pumped out the air. To the amaze­ment of the ob­servers, teams of horses could not pull the halves apart. The at­mos­phere was clamp­ing them to­gether with about two tonnes of force! As soon as he opened a valve and let the air back in, they came apart by hand.

Creating a vac­uum was ex­tremely dif­fi­cult at first. Guericke had to la­bo­ri­ously pump the air out of his sphere, but steam gives us a much faster way to make one. If we fill a ves­sel with steam and then cool it with a spray of wa­ter, the steam con­denses back into roughly 1/1,700 of its vol­ume.

Fill a cylin­der with steam, con­dense it un­der­neath a pis­ton, and the at­mos­phere will drive the pis­ton down into the vac­uum. A near-per­fect vac­uum gives us the same pres­sure dif­fer­ence we used ear­lier: about one kilo­gram of force for every square cen­time­tre of pis­ton. A pis­ton half a me­tre across could col­lect al­most two tonnes of force from the at­mos­phere.

Newcomen’s en­gine

In the early 1700s, mines were get­ting deeper, and flood­ing was be­com­ing a huge prob­lem. Once a shaft reached be­low the wa­ter table, wa­ter seeped in con­tin­u­ously and had to be pumped out day and night. The pumps were dri­ven by teams of horses walk­ing in cir­cles. As one team tired, an­other took over, but the deep­est mines still flooded dur­ing wet weather and valu­able coal had to be aban­doned. A new so­lu­tion was needed, and steam would pro­vide the an­swer.

Thomas Newcomen sup­plied tools to the mines and knew that flood­ing was both a huge prob­lem and an op­por­tu­nity. He spent years turn­ing the vac­uum pis­ton stroke into an en­gine that could run all day. He con­nected the pis­ton to one end of a huge rock­ing beam and hung heavy pump rods from the other. The at­mos­phere drove the pis­ton down and lifted the pump rods; their weight then pulled the pis­ton back up while the cylin­der filled with steam again.

Newcomen’s first suc­cess­ful en­gine was in­stalled at a coal mine near Dudley in 1712. It ran at about twelve strokes per minute, lift­ing roughly forty-five litres of wa­ter fifty me­tres on every stroke. Unlike the horses, it could con­tinue around the clock with­out food or rest. Similar en­gines soon ap­peared in mines from Cornwall to Newcastle.3

Newcomen’s en­gine worked! But it used an ex­tra­or­di­nary amount of coal. The cold wa­ter sprayed di­rectly into the cylin­der, chill­ing a huge mass of iron along with the steam. Roughly three quar­ters of the steam was wasted heat­ing the cylin­der back up on every stroke.

The mines were happy with this trade­off be­cause they burned slack, small pieces of coal that were con­sid­ered waste. Anywhere else, the fuel cost was sim­ply too much. This kept the steam en­gine stuck in coal mines for the next fifty years.

The boiler

Why did Newcomen use the at­mos­phere to push the pis­ton in­stead of the steam it­self? His boiler was sim­ply not strong enough. The haystack boiler pro­duced only about a twen­ti­eth of an at­mos­phere above the sur­round­ing air. It was built from thin cop­per or iron plates joined with riv­ets, and the wide walls and weak seams could not safely hold much pres­sure.

James Watt, who we will meet in the next sec­tion, used the wag­gon boiler shown be­low. Water sat in the broad cham­ber above the fur­nace, the hot gases passed un­der­neath, and steam col­lected be­neath the rounded roof.

The broad bot­tom was good at catch­ing heat, but the wag­gon shape was ter­ri­ble at hold­ing pres­sure. Raise the steam pres­sure in the fig­ure be­low and com­pare what hap­pens to the rounded roof, the flat sides and the in­ward-curved bot­tom.

The fig­ure also shows why later builders curved the whole boiler out­ward like the roof. They rolled iron plate into long cylin­ders, re­mov­ing the flat sides and in­ward-curved bot­tom. They kept the boil­ers nar­row be­cause mak­ing a cylin­der wider in­creases the force try­ing to split it open, even when the pres­sure stays the same.4 Better iron and riv­et­ing then made much higher pres­sures pos­si­ble, and around 1800 Richard Trevithick was run­ning en­gines at sev­eral at­mos­pheres.

Now Newcomen’s use of a vac­uum makes sense. His boiler could push with per­haps fifty grams per square cen­time­tre above at­mos­pheric pres­sure. By con­dens­ing the steam and let­ting the at­mos­phere push the pis­ton in­stead, he got close to one kilo­gram per square cen­time­tre, around twenty times as much force from the same boiler.

Watt’s sep­a­rate con­denser

In 1765, Watt was re­pair­ing a model Newcomen en­gine at the University of Glasgow. He was amazed by how much steam it con­sumed and be­gan try­ing to un­der­stand where it all went. He dis­cussed the prob­lem with his col­league Joseph Black, who was study­ing the heat ab­sorbed while wa­ter boils. Black called it la­tent heat. For a kilo­gram of wa­ter, boil­ing it away takes more than five times as much en­ergy as heat­ing it from freez­ing to boil­ing.

With this knowl­edge, Watt cal­cu­lated the ex­act amount of wa­ter needed to con­dense the vol­ume of steam in the cylin­der. He was sur­prised to find that this ex­act amount barely made a vac­uum at all: the con­dens­ing steam dumped its la­tent heat into the spray, warm­ing the wa­ter un­til it stopped con­dens­ing any­thing. Adding in more cold wa­ter just cooled the cylin­der down more, wast­ing steam to heat the cylin­der back up on the next stroke. Watt’s bril­liant in­sight was to add a sec­ond ves­sel that could stay cold while the cylin­der stayed hot.5

At the end of the stroke, a valve opened and the steam rushed into the cold ves­sel, called the con­denser. As the steam turned back into wa­ter, the pres­sure fell in the con­denser and, through the con­nect­ing pipe, in the cylin­der as well. A small air pump dri­ven by the en­gine drew out the con­densed wa­ter, along with any air that had leaked in, on every stroke. Keeping the cylin­der hot and the con­denser cold cut coal con­sump­tion by about two thirds! Watt and his busi­ness part­ner Matthew Boulton turned the sav­ing into a busi­ness model, charg­ing cus­tomers one third of the money they saved on coal.

Better tools for mak­ing pre­cise cylin­ders al­lowed Watt to make an­other im­por­tant change: he closed the top of the cylin­der and used steam on both sides of the pis­ton. Steam pushed down while the con­denser low­ered the pres­sure be­low; on the re­turn stroke, the same thing hap­pened in the op­po­site di­rec­tion. This was the dou­ble-act­ing en­gine. Below, we can com­pare it with the sin­gle-act­ing cylin­der it re­placed.

The same cylin­der now pro­duced power on both strokes, and the steady push-pull made the en­gine much bet­ter suited to dri­ving ma­chin­ery. But get­ting steam in and out of the cylin­der was now more com­pli­cated. One end had to con­nect to the boiler while the other con­nected to the ex­haust, and then the two con­nec­tions had to switch be­fore the pis­ton re­turned.

The slide valve

Early steam en­gines used sev­eral sep­a­rate valves and link­ages to route the steam. Our en­gine does all of this with one slide valve. It moves only a few cen­time­tres, con­nect­ing one end of the cylin­der to fresh steam and the other to the ex­haust. As the pis­ton reaches the end of its stroke, the valve slides across and swaps the two con­nec­tions.

The valve sits in­side the steam chest, an iron box bolted to the side of the cylin­der and kept full of fresh steam. Three ports open into the chest. The two outer ports con­nect to the ends of the cylin­der, while the mid­dle one car­ries away the ex­haust. The valve is shaped like a wide, hol­low D. One edge un­cov­ers a cylin­der port and lets fresh steam en­ter, while the hol­low back joins the other cylin­der port to the ex­haust.

The valve needs to move in per­fect syn­chro­niza­tion with the pis­ton, or the en­gine will not work. This mo­tion comes from an ec­cen­tric on the en­gine’s ro­tat­ing shaft. The ec­cen­tric is a cir­cu­lar disc mounted slightly off-cen­tre, so its cen­tre trav­els in a small cir­cle as the shaft turns. A strap around the disc fol­lows this mo­tion and dri­ves the valve rod back and forth. Its po­si­tion on the shaft is cho­sen so the next steam port be­gins open­ing be­fore the pis­ton reaches the end of its stroke.

Now, we can see how the pis­ton, valve gear and ec­cen­tric work on our beam en­gine.

Using less steam

We can save a sur­pris­ing amount of coal by clos­ing the steam port be­fore the pis­ton reaches the end of its stroke. The trapped steam con­tin­ues to ex­pand and push the pis­ton, al­though its pres­sure falls as the vol­ume grows. Closing the valve at halfway, called cut­off, uses half as much steam while still pro­duc­ing about 85 per­cent of the ideal work.6 Watt patented this idea in 1782. Later com­pound en­gines sent the ex­haust from one cylin­der into a larger cylin­der, then some­times into a third, ex­tract­ing more work as the steam ex­panded.

Measuring the work

Everything we have just dis­cussed hap­pens in­side an opaque cylin­der. In 1796, Watt’s as­sis­tant John Southern built an in­stru­ment that let them see in­side. A small spring-loaded pis­ton moved a pen­cil up and down with the pres­sure, while a card moved side­ways with the main pis­ton. The re­sult­ing in­di­ca­tor di­a­gram showed the pres­sure through the en­tire stroke, and the area in­side the loop mea­sured the work pro­duced.

A leak­ing pis­ton, late cut­off and re­stricted ex­haust each pro­duce a dif­fer­ent shape, al­low­ing an en­gi­neer to di­ag­nose the en­gine from a sin­gle card. Boulton & Watt found the in­stru­ment so valu­able that they kept it se­cret for years.7

We can now con­trol the steam and pro­duce power in both di­rec­tions, but the pis­ton still moves back and forth. This is called rec­i­p­ro­cat­ing mo­tion. Pumps can use it di­rectly, but the mills dri­ving the Industrial Revolution needed ro­ta­tion.

Making ro­ta­tion

To turn the pis­ton’s back-and-forth mo­tion into ro­ta­tion, our beam en­gine uses a crank, al­though Watt’s first ro­tat­ing en­gines could not use one.8 A pin off­set from the cen­tre of the shaft is joined to the pis­ton by a con­nect­ing rod. The push on the pin turns the shaft, but not equally through the rev­o­lu­tion. Twice per turn the crank and con­nect­ing rod line up, at po­si­tions called dead cen­tres, where the pis­ton pushes straight through the shaft and pro­duces no ro­ta­tion at all. With noth­ing to carry it past these points, the en­gine would stop the first time the crank reached one.

The large fly­wheel fixes this prob­lem. It stores en­ergy while the crank has good lever­age, then re­turns that en­ergy to keep the en­gine spin­ning past the dead cen­tres. In the fig­ure be­low, the shaded band in the in­set shows the fly­wheel col­lect­ing and re­pay­ing en­ergy through each rev­o­lu­tion. Try the fly­wheel mass slider: a heav­ier wheel changes speed less, giv­ing the en­gine a smooth and steady ro­ta­tion.

Our en­gine can now turn a shaft with­out stop­ping. But join­ing the pis­ton rod to the crank turns out to be harder than it looks.

The beam and the par­al­lel mo­tion

Now, look closely at the con­nect­ing rod in the fig­ure be­low. As the crank turns, its pin moves side­ways as well as up and down. The pis­ton rod can­not fol­low it be­cause it must travel straight through the seal at the top of the cylin­der. If we con­nect them di­rectly, the rod pushes the pis­ton side­ways and quickly de­stroys the seal.

The beam car­ried the side­ways load into a large round bear­ing, which work­shops could make ac­cu­rately. But its end moved in an arc, and Watt still needed the pis­ton rod to travel in a straight line.

His in­ge­nious so­lu­tion was the par­al­lel mo­tion, patented in 1784. A set of hinged links joins the beam to a fixed point on the en­gine. As the beam pulls the pis­ton rod side­ways in one di­rec­tion, an­other link pulls it al­most ex­actly the same amount in the other. The two curves can­cel, leav­ing a path that is re­mark­ably close to a straight line. Watt was so pleased with the mech­a­nism that he wrote he was more proud of the par­al­lel mo­tion than of any other me­chan­i­cal in­ven­tion I have ever made.”

Our pis­ton can now turn the crank with­out be­ing pulled side­ways. At the far end of the beam, we also get a con­ve­nient source of back-and-forth mo­tion, which the en­gine uses to keep its boiler filled with wa­ter.

The pump

As the en­gine runs, the boiler turns wa­ter into steam. To keep it go­ing, we need to re­place that wa­ter with­out stop­ping. We can’t sim­ply con­nect a wa­ter tank, be­cause the pres­sure in­side the boiler would push the wa­ter back out. Instead, the far end of the beam dri­ves the small pump be­side the base of the en­gine, forc­ing fresh wa­ter into the boiler.

Inside the pump is a nar­row plunger and two one-way check valves. As the plunger rises, the pres­sure falls, the in­let valve opens and wa­ter en­ters from the tank. On the way down, the pres­sure rises, clos­ing the in­let valve and open­ing the out­let to­wards the boiler. The chang­ing wa­ter pres­sure op­er­ates both valves au­to­mat­i­cally.

The pump must pro­duce slightly more pres­sure than the boiler, but it does not need to move much wa­ter on each stroke. Making the plunger nar­row keeps the re­quired force small, for the same pres­sure-times-area rea­son that made our en­gine pis­ton wide. A small part of the en­gine’s power can now keep the boiler full, while the rest turns the fly­wheel.

Powering the mill

We talked about why mills need ro­ta­tion, but not how they used it. Before steam en­gines, wa­ter-pow­ered mills had to sit be­side a river. The flow­ing wa­ter turned a large wa­ter­wheel, which drove a main shaft, and iron shafts, pul­leys and leather belts car­ried that ro­ta­tion through the build­ing to power the ma­chines. Our ex­am­ple mill here has a saw for cut­ting wood and a power loom which wove cloth. Click ei­ther ma­chine to shift its belt onto the loose pul­ley; that ma­chine will coast to a stop while the shaft and the other ma­chine con­tinue run­ning.

It is not in­tu­itive that a leather belt can trans­mit enough power to drive a ma­chine that ten strong peo­ple could not. With only fric­tion be­tween the iron pul­leys and the leather pro­vid­ing the con­nec­tion, it seems that the belt would slip. The physics un­der­ly­ing fric­tion is fas­ci­nat­ing. Imagine a huge ship tied to an iron bol­lard on the dock with a rope. Tension in the first small part of the rope presses it against the iron, and the re­sult­ing fric­tion re­duces the ten­sion that reaches the next part, and so on around the post.9

Now, let’s re­turn to leather belts and iron pul­leys. A belt is in­stalled un­der ten­sion, so at rest its two sides pull with roughly equal force. Once the ma­chine needs power, fric­tion trans­fers some of that pull from the re­turn­ing side to the dri­ving side.10

The power trans­ferred by a belt is the dif­fer­ence in ten­sion mul­ti­plied by the belt speed. At full mill scale, a six­teen-foot fly­wheel at sixty rev­o­lu­tions per minute has a belt speed of about fif­teen me­tres a sec­ond. If the load makes one side pull with 2,000 new­tons more than the other, a foot-wide leather belt can carry about forty horse­power.

The fight for wa­ter

Richard Arkwright’s wa­ter-pow­ered mill at Cromford opened in 1771, and the fac­tory sys­tem that fol­lowed cre­ated fierce de­mand for the best river sites. Water-powered mills were also de­pen­dent on the weather: a dry sea­son could shut down the fac­tory.

Steam pump­ing en­gines of­fered a so­lu­tion. An en­gine lifted the wa­ter that had passed be­neath the wheel back up the hill, al­low­ing the same wa­ter to fall through the wheel again. This kept the smooth turn of the wa­ter wheel, but wasted coal mov­ing the wa­ter. Watt sold six­teen to twenty horse­power pump­ing en­gines to de­liver ten horse­power to the ma­chines.

Watt’s dou­ble-act­ing en­gine, beam, crank and fly­wheel let the en­gine di­rectly turn the line shaft. This met a huge de­mand from mill own­ers who wanted to build near work­ers and ma­te­ri­als rather than around a par­tic­u­lar stretch of river.11 One prob­lem re­mained, though: every time a ma­chine was turned on or off, the load on the en­gine changed.

The gov­er­nor

The beam en­gine still needed a way to keep its speed con­stant. Imagine it at the be­gin­ning of the day, turn­ing at 30 rpm with no ma­chines con­nected. When the first ma­chine is con­nected, it draws power from the en­gine and slows it down. The en­gine dri­ver could open the throt­tle by hand un­til the shaft re­turned to 30 rpm, but this was tir­ing work, and mis­takes had se­vere con­se­quences. A cast-iron fly­wheel could burst if it spun too quickly.

Instead, Watt adapted a de­vice used on wind­mills to ad­just the steam au­to­mat­i­cally.12 Bevel gears turn a ver­ti­cal spin­dle, and two heavy balls hang from hinged arms at­tached to it. As the en­gine speeds up, the balls swing out­ward and lift a slid­ing col­lar. A fork and long rod carry this mo­tion across the en­gine and turn the steam cock to­wards closed. When the en­gine slows, the balls fall and open the cock again.13

The whole ma­chine

Let’s re­turn to the com­plete en­gine from the be­gin­ning of the ar­ti­cle. Every mech­a­nism we stud­ied on its own is here, run­ning in its place. The fig­ure fol­lows the power once along its whole path, from the boiler steam to the belt that leaves for the mill.14

Epilogue

The beam en­gine was a prod­uct of the tools and sci­ence of its time. Watt used a beam and par­al­lel mo­tion partly be­cause the work­shops of the 1780s could not make long, ac­cu­rate guides for a crosshead. As plan­ing ma­chines im­proved dur­ing the nine­teenth cen­tury, those straight guides be­came prac­ti­cal. The heavy beam was no longer re­quired, and by the 1860s most new mill en­gines drove the fly­wheel di­rectly.15

Line shafts and leather belts out­lived the beam en­gine, re­main­ing above fac­tory floors well into the twen­ti­eth cen­tury. Electric mo­tors fi­nally gave each ma­chine its own source of ro­ta­tion. Wires re­placed the long shafts and belts, and stop­ping one lathe no longer changed the load on a cen­tral en­gine dri­ving the en­tire mill.

The most dra­matic change was how much power newer en­gines ex­tracted from coal. Corliss valves con­trolled steam ex­pan­sion more pre­cisely, com­pound en­gines ex­panded it through sev­eral cylin­ders, and tur­bines even­tu­ally re­placed the pis­ton with a con­tin­u­ously ro­tat­ing wheel. Newcomen con­verted only about half a per­cent of the heat into use­ful work. Watt’s con­denser raised the use­ful share to roughly three per­cent, enough for steam power to move away from the coal mines. By the 1890s, high pres­sure and com­pound ex­pan­sion pushed large ma­rine en­gines such as the Titanic’s be­yond ten per­cent. A mod­ern steam tur­bine plant con­verts more than forty per­cent.

Footnotes

Thomas Savery tried to use higher-pres­sure steam in the 1690s with boil­ers made from sol­dered cop­per. The fire could soften the sol­der, and the leak­ing joints needed fre­quent re­pair. Newcomen took a dif­fer­ent route. Because the steam in his boiler was barely above at­mos­pheric pres­sure, he could use thin lead and wrought-iron plates joined with riv­ets. The seams still leaked and the metal cor­roded, but the boiler did not have to con­tain the pres­sure that Savery’s pump re­quired. ↩

Thomas Savery tried to use higher-pres­sure steam in the 1690s with boil­ers made from sol­dered cop­per. The fire could soften the sol­der, and the leak­ing joints needed fre­quent re­pair. Newcomen took a dif­fer­ent route. Because the steam in his boiler was barely above at­mos­pheric pres­sure, he could use thin lead and wrought-iron plates joined with riv­ets. The seams still leaked and the metal cor­roded, but the boiler did not have to con­tain the pres­sure that Savery’s pump re­quired. ↩

Casting a large iron cylin­der was much eas­ier than mak­ing the in­side straight and round. Newcomen’s cylin­ders were ground by hand, then sealed with a leather flap cov­ered by a layer of wa­ter, which could fol­low the un­even bore. Denis Papin had pro­posed the vac­uum-pis­ton prin­ci­ple in 1690: a small amount of wa­ter boiled be­neath a pis­ton and pushed it up­ward, then con­den­sa­tion al­lowed the at­mos­phere to force it down again. His ap­pa­ra­tus demon­strated a sin­gle stroke but did not be­come a con­tin­u­ously run­ning en­gine. ↩

Casting a large iron cylin­der was much eas­ier than mak­ing the in­side straight and round. Newcomen’s cylin­ders were ground by hand, then sealed with a leather flap cov­ered by a layer of wa­ter, which could fol­low the un­even bore. Denis Papin had pro­posed the vac­uum-pis­ton prin­ci­ple in 1690: a small amount of wa­ter boiled be­neath a pis­ton and pushed it up­ward, then con­den­sa­tion al­lowed the at­mos­phere to force it down again. His ap­pa­ra­tus demon­strated a sin­gle stroke but did not be­come a con­tin­u­ously run­ning en­gine. ↩

A pis­ton 50 centimetres across has about 2,000 square cen­time­tres of area, enough to col­lect two tonnes of force from a per­fect vac­uum. After al­low­ing for leaks and the weight of the pump rods, it might do about four kilo­watts of use­ful work. A horse can sus­tain much less than one horse­power over a work­ing day, so re­plac­ing the en­gine re­quired a re­lay of per­haps fif­teen or twenty an­i­mals. Watt later sold his en­gines by the num­ber of horses they re­placed, and fixed one horse­power at 33,000 foot-pounds per minute. ↩

A pis­ton 50 centimetres across has about 2,000 square cen­time­tres of area, enough to col­lect two tonnes of force from a per­fect vac­uum. After al­low­ing for leaks and the weight of the pump rods, it might do about four kilo­watts of use­ful work. A horse can sus­tain much less than one horse­power over a work­ing day, so re­plac­ing the en­gine re­quired a re­lay of per­haps fif­teen or twenty an­i­mals. Watt later sold his en­gines by the num­ber of horses they re­placed, and fixed one horse­power at 33,000 foot-pounds per minute. ↩

For a bar­rel with ra­dius r and length L, the cut has an area of 2rL, so pres­sure p pushes the halves apart with a force of 2prL. Two edges of length L re­sist that force, leav­ing pr in each me­tre of plate. At two at­mos­pheres and a half-me­tre ra­dius, this is about ten tonnes per me­tre. The stress run­ning length­wise is only half as large, which is why a cylin­dri­cal boiler tends to split along its length like a sausage. A sphere di­vides the load equally and is stronger still, but it was much harder to make from rolled and riv­eted plate. ↩

For a bar­rel with ra­dius r and length L, the cut has an area of 2rL, so pres­sure p pushes the halves apart with a force of 2prL. Two edges of length L re­sist that force, leav­ing pr in each me­tre of plate. At two at­mos­pheres and a half-me­tre ra­dius, this is about ten tonnes per me­tre. The stress run­ning length­wise is only half as large, which is why a cylin­dri­cal boiler tends to split along its length like a sausage. A sphere di­vides the load equally and is stronger still, but it was much harder to make from rolled and riv­eted plate. ↩

Watt’s en­gine needed a much more ac­cu­rate cylin­der than Newcomen’s loose, wa­ter-sealed pis­ton. Around 1775, the iron­mas­ter John Wilkinson built a bor­ing mill with a rigid cut­ting bar sup­ported at both ends, adapt­ing tech­niques he had de­vel­oped for bor­ing can­nons. In 1776, Matthew Boulton re­ported that a 50-inch cylin­der in­stalled at Tipton var­ied by less than the thick­ness of an old shilling. This ac­cu­racy kept the steam from leak­ing around Watt’s pis­ton and made the new en­gine prac­ti­cal. ↩

Watt’s en­gine needed a much more ac­cu­rate cylin­der than Newcomen’s loose, wa­ter-sealed pis­ton. Around 1775, the iron­mas­ter John Wilkinson built a bor­ing mill with a rigid cut­ting bar sup­ported at both ends, adapt­ing tech­niques he had de­vel­oped for bor­ing can­nons. In 1776, Matthew Boulton re­ported that a 50-inch cylin­der in­stalled at Tipton var­ied by less than the thick­ness of an old shilling. This ac­cu­racy kept the steam from leak­ing around Watt’s pis­ton and made the new en­gine prac­ti­cal. ↩

For a cylin­der with vol­ume V and pres­sure p, ad­mit­ting steam for the full stroke pro­duces work pV. With cut­off at half stroke, the ad­mit­ted steam pro­duces pV/​2 dur­ing the first half. As it ex­pands through the rest of the cylin­der, it adds about 0.35 pV more, as­sum­ing it fol­lows Boyle’s law and re­mains hot. This gives 85 per­cent of the full-stroke work from half the steam. Cutting off ear­lier saves still more steam, but even­tu­ally the falling pres­sure be­comes too weak to over­come fric­tion and the poor lever­age near dead cen­tre. ↩

For a cylin­der with vol­ume V and pres­sure p, ad­mit­ting steam for the full stroke pro­duces work pV. With cut­off at half stroke, the ad­mit­ted steam pro­duces pV/​2 dur­ing the first half. As it ex­pands through the rest of the cylin­der, it adds about 0.35 pV more, as­sum­ing it fol­lows Boyle’s law and re­mains hot. This gives 85 per­cent of the full-stroke work from half the steam. Cutting off ear­lier saves still more steam, but even­tu­ally the falling pres­sure be­comes too weak to over­come fric­tion and the poor lever­age near dead cen­tre. ↩

The pres­sure and vol­ume graph out­lived the me­chan­i­cal in­di­ca­tor. In 1834, Émile Clapeyron used the same type of di­a­gram to ex­plain Sadi Carnot’s the­ory of heat en­gines, and ther­mo­dy­nam­ics still plots pres­sure against vol­ume to­day. ↩

The pres­sure and vol­ume graph out­lived the me­chan­i­cal in­di­ca­tor. In 1834, Émile Clapeyron used the same type of di­a­gram to ex­plain Sadi Carnot’s the­ory of heat en­gines, and ther­mo­dy­nam­ics still plots pres­sure against vol­ume to­day. ↩

James Pickard patented the use of a crank on a steam en­gine in 1780, so William Murdoch de­signed the sun-and-planet gear as a way around the patent. A gear at­tached to the con­nect­ing rod trav­elled around a sec­ond gear on the fly­wheel shaft, turn­ing the shaft twice for every cy­cle of the beam. Once Pickard’s patent ex­pired in 1794, builders re­turned to the much sim­pler crank used on our en­gine.

Murdoch’s mech­a­nism be­longs to the epicyclic, or plan­e­tary, fam­ily of gears. Several planet gears can share the load while the in­put and out­put re­main on the same axis, mak­ing the arrange­ment com­pact and strong. Planetary gears ap­pear in cord­less drills, bi­cy­cle hubs, au­to­matic trans­mis­sions and wind tur­bines. Hybrid cars even use them to di­vide power be­tween the en­gine, elec­tric mo­tor and wheels. ↩

James Pickard patented the use of a crank on a steam en­gine in 1780, so William Murdoch de­signed the sun-and-planet gear as a way around the patent. A gear at­tached to the con­nect­ing rod trav­elled around a sec­ond gear on the fly­wheel shaft, turn­ing the shaft twice for every cy­cle of the beam. Once Pickard’s patent ex­pired in 1794, builders re­turned to the much sim­pler crank used on our en­gine.

Murdoch’s mech­a­nism be­longs to the epicyclic, or plan­e­tary, fam­ily of gears. Several planet gears can share the load while the in­put and out­put re­main on the same axis, mak­ing the arrange­ment com­pact and strong. Planetary gears ap­pear in cord­less drills, bi­cy­cle hubs, au­to­matic trans­mis­sions and wind tur­bines. Hybrid cars even use them to di­vide power be­tween the en­gine, elec­tric mo­tor and wheels. ↩

If we slice the wrapped rope into small pieces and add the force vec­tors, each piece has a small in­ward force equal to the lo­cal ten­sion mul­ti­plied by the an­gle it cov­ers. Friction can re­move up to μ times that in­ward force, about 0.3 for rope on cast iron.

Repeating that frac­tional re­duc­tion pro­duces the ex­po­nen­tial e−μθ. With a 2,000-newton pull, about the weight of an up­right pi­ano, one turn around the post leaves 300 new­tons, two leave 46, and three leave seven. ↩

If we slice the wrapped rope into small pieces and add the force vec­tors, each piece has a small in­ward force equal to the lo­cal ten­sion mul­ti­plied by the an­gle it cov­ers. Friction can re­move up to μ times that in­ward force, about 0.3 for rope on cast iron.

Repeating that frac­tional re­duc­tion pro­duces the ex­po­nen­tial e−μθ. With a 2,000-newton pull, about the weight of an up­right pi­ano, one turn around the post leaves 300 new­tons, two leave 46, and three leave seven. ↩

My security camera shipped a GitHub admin token in its login page

hhh.hn

i have been think­ing a bit more about se­cu­rity cam­eras again, be­cause of AXIS start­ing to push more for every one of their cam­eras to be able to eas­ily run linux ap­pli­ca­tions on them, they’re far more se­ri­ous tar­gets in an en­ter­prise en­vi­ron­ment and need to be man­aged as such for vul­ner­a­bil­i­ties and cre­den­tial man­age­ment etc. some­one brought up to me a com­pany that sounded new to me, Hanwha (Vision.) I took a look at the site, and found that they had ac­ces­si­ble firmware blobs for each model of cam­era, which is al­ways a treat.

pok­ing and prod­ding

i took the im­age and threw it at bin­walk hop­ing it was just a rootfs or some­thing, but in­side there was a sep­a­rate tar­ball with some AI stuff for the cam­era and a fwim­age.tgz that bin­walk was flag­ging as en­crypted.

I was googling around and saw that Matt Brown has a writeup on these cam­eras that got me through. ba­si­cally the passphrase is HTW + the model num­ber so HTWXNP-9300RW worked.

seems like they do some more stuff now, be­cause in­side of that tar­ball was an­other fwim­age.tgz that was en­crypted but it was­n’t the same scheme, so we can’t just re-use the same setup from Matt Brown. I kinda fig­ured I was gonna have to give up and that Hanwha was do­ing some­thing more ad­vanced, like burn­ing a key into the hard­ware (which isnt fool­proof ob­vi­ously but you at least need to own the cam­era to start.) Anyways, there was a fwup­grader bi­nary that was in that outer tar­ball, so I threw it into ghidra and started pok­ing around.

well I would have been pok­ing around if it was 2023 or some­thing, but i pointed claude code at it and went to make a lovely din­ner and spent time with my part­ner in­stead, and came back a bit later to a de­scrip­tion and a nice rootfs.

Hanwha had built some ob­fus­ca­tion into the fwup­grader to hide how they de­crypt the ac­tual rootfs. the AES key is XOR’d against a small sta­tic key table in the bi­nary and re­assem­bled at run­time ( the IV is just plain­text in there) the fwup­grader just shells out to the openssl CLI, and even the com­mand frag­ments are XOR-obfuscated the same way.

re­con­structed the com­mand looks like this:

openssl enc -md sha256 -aes-256-cbc -d \ -K <KEY> -iv <IV> -in <INPUT> -out <OUTPUT>

Since the key and iv are just hard­coded (the same across the model line), I will pub­lish them here:

KEY = dfa049b­b922e63e2dec­c764af5628068e5b7a2662e479a615b14643e567579b0 IV = 53f926801b81454a4f889c9a390db6e6

and with that we have a full rootfs to dig into nor­mally.

truf­fles

since we fi­nally can just look at stuff I ran truf­fle­hog im­me­di­ately to see if there was any­thing ob­vi­ous, and there was a github to­ken du­pli­cated in like 30 files… I checked what re­pos the to­ken had ac­cess to, and it had ad­min priv­i­leges to hun­dreds of repos­i­to­ries in their github or­ga­ni­za­tion.

this is­n’t my first rodeo with an org ship­ping a Github to­ken in their firmware though… but that’s a story for a dif­fer­ent blog post. Why would this org put this to­ken in like 30 files though? it looks like they build the UI for these cam­eras with vite, and one of the vari­ables is be­ing set to the en­tirety of process.env at build time, which means the en­tirety of the CI job’s en­vi­ron­ment is be­ing writ­ten to these files.

var W = { DATAPORT: 9090”, GIT_LFS_SKIP_SMUDGE: 1″, npm_­com­mand: run-script”, KUBERNETES_SERVICE_PORT_HTTPS: 443″, GITHUB_NPM_TOKEN: <snip>:ghp_…REDACTED…”, npm_­con­fig_user­con­fig: /home/docker/.npmrc”, // etc

I don’t have any of these cam­eras to test, but I think this would mean that any­one ac­cess­ing the ad­min ui of these cam­eras likely has had this github to­ken sent to them over the wire and (hopefully) no­body evil no­ticed. Maybe it did­n’t get ac­tu­ally served and just lived on disk, though.

there were some other… in­ter­est­ing bits of data in the en­vi­ron­ment though: there were some env vars with IP ad­dresses in them, but they are as­signed to the US Department of Defense:

SWARM_MASTER_NFS_ADDRESS: 55.101.212.23

OTEL_ELASTIC_URL: http://​55.101.212.21:5601/<snip>

CIMIP: 55.101.211.213

huh… is this just a co­in­ci­dence and one of those weird in­stances where peo­ple have taken IP space for in­ter­nal ser­vices when they know they will never in­ter­act with it (insane prac­tice btw…) or is Hanwha more di­rectly tied with the US DoD?

let’s look at the wikipedia page for Hanwha Vision:

Hanwha Vision (Korean: 한화비전), founded as Samsung Techwin, is a video sur­veil­lance com­pany. It is a sub­sidiary of Hanwha Group.

Hanwha Vision (Korean: 한화비전), founded as Samsung Techwin, is a video sur­veil­lance com­pany. It is a sub­sidiary of Hanwha Group.

Former prod­ucts K9 Thunder self-pro­pelled ar­tillery, K10 am­mu­ni­tion re­sup­ply ve­hi­cles, sub-sys­tems for K2 Black Panther, sen­try gun ro­bot SGR-A1.

Former prod­ucts

K9 Thunder self-pro­pelled ar­tillery, K10 am­mu­ni­tion re­sup­ply ve­hi­cles, sub-sys­tems for K2 Black Panther, sen­try gun ro­bot SGR-A1.

oh… okay…. I re­mem­ber read­ing about the SGR-A1 when I was in high school, but I never thought I would be ac­ci­den­tally find­ing keys to the king­dom of the man­u­fac­turer on the ground later in my ca­reer… my life is kinda weird some­times…

SPECULATION WARNING

even still, these aren’t American de­vices, or any­thing like that. Why would Hanwha Vision need any­thing re­motely re­lated to the DoD? Is it pos­si­ble that their CI is pro­vided by some cen­tral­ized team at their par­ent com­pany Hanwha, where the needs of their sis­ter com­pany Hanwha Aerospace cause the shared plat­form to have these en­tries in the CI en­vi­ron­ment vari­ables? Or maybe be­cause of their other sis­ter com­pany, Hanwha Defense USA, where they make other large scary steel ma­chines

I wanted to make sure this was­n’t some kind of fluke, and that there weren’t hun­dreds of other dif­fer­ent github to­kens in their firmware, so I scraped the Hanwha web­site to down­load every firmware for every cam­era i could find, and ended up with around ~500 firmwares (there were like 600 smth cam­eras but not all of them had firmware listed) and I was able to ex­tract 62% of them with the same ap­proach as above, and only three of them had github to­kens, and they were all the same to­ken.

not re­ally sure why the oth­ers did­n’t work, but it’s close enough for me to feel sat­is­fied.

dis­clo­sure

I wrote up a very small email with enough in­for­ma­tion to iden­tify where the to­ken was and sent it over to Hanwha, who have a nice open email for re­port­ing se­cu­rity is­sues, and they re­sponded within 12 hours no­ti­fy­ing me that the to­ken had been re­voked. Sure they should­n’t have ever had a gh to­ken in there but I have never had such a prompt re­sponse and res­o­lu­tion.

we re­ally gotta stop mak­ing these mis­takes so of­ten, how am I sup­posed to be sleep­ing at night?

thanks com­puter, un­til next time

India's first privately developed rocket reaches orbit on dramatic debut launch

arstechnica.com

Think big

On the first at­tempt, reach­ing or­bit, I never thought it was pos­si­ble.”

V. Narayanan, chair­man of the Indian Space Research Organization (second from left), and Pawan Kumar Chandana, CEO of Skyroot Aerospace (second from right), pose with a replica of the Vikram-1 rocket along with other se­nior Indian space of­fi­cials fol­low­ing a suc­cess­ful launch Saturday at the Satish Dhawan Space Center on Sriharikota Island, India.

Credit:

R. Satish Babu/AFP via Getty Images

Indian space of­fi­cials cel­e­brated the de­but flight of Skyroot Aerospace’s Vikram-1 rocket, India’s first fully com­mer­cial satel­lite launcher, as a grand suc­cess” Saturday af­ter an on-tar­get climb into a 280-mile-high or­bit fol­low­ing liftoff from an is­land space­port in the Bay of Bengal.

The Vikram-1 lifted off from India’s pri­mary space­port on Sriharikota Island at 1:35 am EDT (06:35 UTC) Saturday, around mid­day at the launch base along India’s south­east coast. The launch was de­layed more than a half-hour to re­solve a last-minute tech­ni­cal prob­lem. The count­down re­sumed, cul­mi­nat­ing in the com­mand to ig­nite Vikram-1’s solid-fu­eled first stage booster to pro­pel the rocket off the launch pad.

Vikram-1 is mod­est in size com­pared to India’s larger work­horse rock­ets. Skyroot’s rocket stands about 72 feet (22 me­ters) tall, with the ca­pa­bil­ity to place pay­loads of up to 770 pounds (350 kilo­grams) into low-Earth or­bit. This makes Vikram-1 some­what larger than the Electron launch ve­hi­cle de­vel­oped by Rocket Lab, the world’s most suc­cess­ful ded­i­cated small satel­lite launcher.

The flight Saturday went off with­out any ma­jor prob­lems. Three solid-fu­eled rocket mo­tors fired in suc­ces­sion to reach space, then a small liq­uid-fu­eled fourth stage ig­nited and ac­cel­er­ated to or­bital ve­loc­ity, some 17,000 mph. Live views from on­board cam­eras showed each phase of the launch se­quence.

The only sign of any­thing un­usual came dur­ing the sep­a­ra­tion of the rock­et’s third stage from its fourth stage. The spent third stage mo­tor ap­peared to re­main near the fourth stage dur­ing a brief coast, rather than back­ing away to a greater dis­tance. Nevertheless, the fourth stage did its job, fir­ing its 3D-printed en­gine to reach an or­bit ap­prox­i­mately 280 miles (450 kilo­me­ters) high at an in­cli­na­tion of 60 de­grees to the equa­tor, quite close to pre­flight pre­dic­tions, ac­cord­ing to Skyroot Aerospace. US mil­i­tary track­ing data con­firmed the rock­et’s suc­cess­ful ar­rival in or­bit.

Skyroot Aerospace’s Vikram-1 rocket lifts off Saturday from the Satish Dhawan Space Center on Sriharikota Island, India.

Credit: R. Satish Babu/AFP via Getty Images

Skyroot Aerospace’s Vikram-1 rocket lifts off Saturday from the Satish Dhawan Space Center on Sriharikota Island, India.

Credit:

R. Satish Babu/AFP via Getty Images

Beating the odds

We achieved one of the biggest mile­stones ever in India’s space sec­tor—the first pri­vate or­bital rocket reach­ing or­bit on the very first at­tempt,” said Pawan Kumar Chandana, Skyroot’s co­founder and CEO, in re­marks to the com­pa­ny’s launch team. It still feels like a dream, and you all made this dream hap­pen.”

The first flights of new pri­vate or­bital-class rock­ets don’t have a great track record. It took SpaceX four tries be­fore reach­ing or­bit with the Falcon 1 rocket for the first time in 2008. Rocket Lab’s Electron did­n’t make it to or­bit on its first launch in 2017. Blue Origin beat the odds with the in­au­gural flight of its heavy-lift New Glenn rocket in 2025, but the com­pa­ny’s en­gi­neers had pre­vi­ous ex­pe­ri­ence with nu­mer­ous launches of the smaller New Shepard sub­or­bital rocket.

On the first at­tempt, reach­ing or­bit, I never thought it was pos­si­ble,” Chandana said. Skyroot’s team made it pos­si­ble. A big, big, big shoutout to this phe­nom­e­nal team, which made it hap­pen. In fact, this launch was noth­ing short of a sus­pense movie.”

Skyroot of­fi­cials set hum­ble goals for the first launch of Vikram-1. In a press kit re­leased be­fore the flight, the com­pany said its pri­mary ob­jec­tive for the launch was to com­plete a suc­cess­ful liftoff, clear the tower at the launch site, and gather max­i­mum data dur­ing as­cent.

The mis­sion ob­jec­tive was only to lift off and clear the tower,” said Pawan Goenka, chair­man of IN-SPACe, a gov­ern­ment or­ga­ni­za­tion set up in 2020 to pro­mote India’s com­mer­cial space in­dus­try. That was only about 100 me­ters, but what we went to was 450 kilo­me­ters, and it also re­leased all the satel­lites that were sup­posed to re­lease. So the mis­sion was ab­solutely per­fect.”

Skyroot Aerospace’s Vikram-1 rocket on its launch pad.

Credit: ISRO

Skyroot Aerospace’s Vikram-1 rocket on its launch pad.

Credit:

ISRO

In a state­ment, the Indian space agency, ISRO, said it of­fered handholding and sup­port” to the Skyroot ven­ture by pro­vid­ing ac­cess to solid rocket mo­tor cast­ing and test fa­cil­i­ties at ISROs space­port on Sriharikota. ISRO also al­lowed Skyroot to launch from one of its two ac­tive launch pads.

Painted blue and white, the Vikram-1 is made of light­weight car­bon com­pos­ite ma­te­ri­als and is named for the Indian physi­cist Vikram Sarabhai, con­sid­ered the fa­ther of the Indian space pro­gram. Skyroot suc­cess­fully launched a sub­or­bital rocket, Vikram-S, to an al­ti­tude of nearly 300,000 feet (90 kilo­me­ters) in November 2022.

The Vikram-1 builds on lessons learned with Vikram-S. Skyroot’s fu­ture roadmap in­cludes the Vikram-1U, with ad­di­tional strap-on solid rocket boost­ers to haul heav­ier pay­loads, and the Vikram-2, which will de­but a cryo­genic up­per stage to reach a pay­load ca­pac­ity of 2,000 pounds (900 kilo­grams) to low-Earth or­bit. The ini­tial pur­pose of the Vikram rocket fam­ily is to deliver ded­i­cated and re­spon­sive launch ser­vices for small satel­lites,” Skyroot of­fi­cials wrote in the press kit for Saturday’s mis­sion.

But the com­pany has loftier am­bi­tions. In an in­ter­view ahead of the first Vikram-1 launch, Chandana told Ars his as­pi­ra­tion for Skyroot in­volves larger liq­uid-fu­eled fully reusable rock­ets, with a daily ca­dence” from mul­ti­ple coun­tries.

Skyroot will need a lot more fund­ing to re­al­ize that dream, but Saturday’s launch showed the com­pany has in­gre­di­ents re­quired for a suc­cess­ful launch com­pany. Saturday’s launch vaulted Skyroot to a plane above any other space startup in India, or, for that mat­ter, in any coun­try out­side of the United States and China. Skyroot has, so far, raised ap­prox­i­mately $160 mil­lion in cap­i­tal, bring­ing the com­pa­ny’s val­u­a­tion to $1.1 bil­lion. Skyroot now has more than 1,000 em­ploy­ees, mostly work­ing out of the com­pa­ny’s head­quar­ters in Hyderabad. The av­er­age age of Skyroot’s work­force is 28 years old.

This view of the pay­load deck of the Vikram-1 rock­et’s up­per stage was cap­tured mo­ments af­ter or­bital in­ser­tion Saturday. The rocket de­ployed two small CubeSats and hosted sev­eral more pay­loads that re­mained at­tached to the up­per stage.

Credit: Skyroot Aerospace

This view of the pay­load deck of the Vikram-1 rock­et’s up­per stage was cap­tured mo­ments af­ter or­bital in­ser­tion Saturday. The rocket de­ployed two small CubeSats and hosted sev­eral more pay­loads that re­mained at­tached to the up­per stage.

Credit:

Skyroot Aerospace

Skyroot’s break­through launch comes as India’s gov­ern­ment, led by Prime Minister Narendra Modi, seeks to su­per­charge the coun­try’s space in­dus­try. India has long had a ro­bust space pro­gram, with gov­ern­ment-de­vel­oped rock­ets such as the Polar Satellite Launch Vehicle and the larger LVM3 of­ten at­tract­ing com­mer­cial cus­tomers from the United States and Europe. India be­came the fourth coun­try to suc­cess­fully land a space­craft on the Moon in 2023, and is work­ing on an oft-de­layed hu­man-rated crew cap­sule to fly as­tro­nauts to low-Earth or­bit.

Modi has told the Indian space in­dus­try to in­crease its an­nual launch to­tal from about five launches per year to 50 be­fore the end of the decade. The prime min­is­ter called Chandana and con­grat­u­lated the Skyroot team af­ter Saturday’s launch.

This is a defin­ing mo­ment in India’s space jour­ney,” Modi said in a state­ment. The grow­ing par­tic­i­pa­tion of our pri­vate sec­tor is open­ing new fron­tiers and ac­cel­er­at­ing in­no­va­tion. This achieve­ment will en­cour­age count­less young­sters to dream big­ger and in­no­vate fear­lessly.”

Chandana, a for­mer en­gi­neer at India’s space agency, founded Skyroot in 2018 with an­other ISRO sci­en­tist, Naga Bharath Daka. They de­cided to fo­cus on de­vel­op­ing a solid-fu­eled launcher first, op­ti­miz­ing for what Chandana de­scribed as the low­est de­vel­op­ment time and the low­est cost per launch. We wanted to get to an or­bital launch ve­hi­cle in a few years,” Chandana told Ars.

It’s a test launch,” he said at the time. Statistically, the first launch from a pri­vate com­pany al­most al­ways fails. It’s very dif­fi­cult to suc­ceed with all new sys­tems. But I think we have done every­thing we can do to en­sure the first launch goes well.”

Indeed, the first launch went very well, ex­ceed­ing all ex­pec­ta­tions. A sec­ond Vikram-1 launch could hap­pen be­fore the end of the year, Chandana said.

This is a 100 per­cent de­signed in India rocket, a 100 per­cent made in India rocket, built by 100 per­cent Indian peo­ple, for India and for the world,” Chandana said af­ter the launch Saturday. This was a his­toric mo­ment for India, but also a very proud mo­ment for the global space sec­tor be­cause the world needs more ac­cess to space.”

Stephen Clark is a space re­porter at Ars Technica, cov­er­ing pri­vate space com­pa­nies and the world’s space agen­cies. Stephen writes about the nexus of tech­nol­ogy, sci­ence, pol­icy, and busi­ness on and off the planet.

60 Comments

Just a moment...

www.science.org

Nothing Works and Everyone Is Euphoric

ptrchm.com

As I’m writ­ing this, we’re in the mid­dle of an AI-induced mass psy­chosis. People are lit­er­ally to­ken-maxxing them­selves into hos­pi­tal beds, scram­bling to cap­ture some of that mar­ket value be­fore every­thing is au­to­mated away. I can’t blame them. Models keep get­ting bet­ter, pro­gram­mers are be­ing laid off left and right. We’ve been re­peat­edly told that AI will write 100% of the code by the end of the year. Whether that’s true or not, this may not be the best time to sit back.

The wide­spread ex­cite­ment around the Agentic Era comes with the promise of greater pro­duc­tiv­ity and higher qual­ity. There’s no deny­ing that these new tools have al­ready rev­o­lu­tion­ized how we cre­ate and use soft­ware. They have raised up­per man­age­men­t’s ex­pec­ta­tions for team out­put. They may have up­graded the av­er­age skill set of soft­ware teams in a way we have not seen be­fore.

So why does soft­ware keep get­ting worse across the board?

A few ex­am­ples from last week alone:

My bank­ing app re­quires, on av­er­age, three FaceID lo­gins be­fore the 3D Secure con­fir­ma­tion view ap­pears.

I opened Slack on ma­cOS, the icon kept bounc­ing in the dock for a few sec­onds. I got im­pa­tient, switched to Ghostty, and started typ­ing. Just then, the Slack win­dow ap­peared, stole fo­cus from Ghostty and the git pull com­mand was sent to the group chat.

My LG fridge started mak­ing weird sounds, so I tried to file a war­ranty claim, through a multi-step form with count­less fields. It failed with a sub­mis­sion er­ror at the very end. And I only found out be­cause I looked at the JavaScript con­sole.

My car’s in­fo­tain­ment sys­tem got a soft­ware up­date re­cently. It was never great to be­gin with, but at least it did­n’t re­boot it­self dur­ing every drive. Now, it’s rid­dled with bugs: the turn-sig­nal sound ran­domly goes silent un­til I re­boot the OS; I tap the screen to open Google Maps — the ra­dio app shows up; there’s a 1 – 2-second lag be­fore any­thing hap­pens af­ter I tap the screen. This is no longer just a UX prob­lem at this point — those bugs af­fect your abil­ity to fo­cus on dri­ving.

A few months ago, I saw a LinkedIn thread by a PM on the team that re­designed the car’s OS. They were con­grat­u­lat­ing them­selves on what an amaz­ing job they had done. I keep think­ing about that post every time I have to fight their prod­uct.

A few months ago, I saw a LinkedIn thread by a PM on the team that re­designed the car’s OS. They were con­grat­u­lat­ing them­selves on what an amaz­ing job they had done. I keep think­ing about that post every time I have to fight their prod­uct.

While I can’t know the full story, I’m will­ing to bet that most of the teams be­hind those bugs have ac­cess to the lat­est mod­els, with gen­er­ous to­ken bud­gets. LLMs can be re­ally good at squash­ing bugs if given the chance.

Software has al­ways had bugs, and the nos­tal­gia for the good old ma­cOS Snow Leopard era when every­thing was sta­ble is mostly the prod­uct of se­lec­tive mem­ory. Software may have been bet­ter back in the day, but that was mainly be­cause it was much sim­pler. Since then, we have kept com­ing up with new ab­strac­tions, new fron­tend frame­works, and more in­fra­struc­ture com­plex­ity. The bar for user ex­pe­ri­ence” has kept ris­ing, but every­thing has be­come in­creas­ingly frag­ile.

We’ve reached a point where an up­date to ma­cOS — or to any app I rely on, re­ally — is a source of dread rather than ex­cite­ment. I now ex­pect the new ver­sion to be worse.

This is­n’t a rant against AI. Those hum­ming GPU farms have given us su­per­pow­ers, but we still don’t use them to build bet­ter soft­ware.

Software ven­dors have long been KPI-oriented, and mak­ing things more sta­ble does­n’t al­ways have a di­rect ef­fect on the num­bers. It does­n’t look ex­cit­ing in pre­sen­ta­tions:

This quar­ter, we won’t be re­leas­ing any new fea­tures, and we have no plans to re­design any­thing — we will ex­clu­sively fo­cus on fix­ing bugs. — Imaginary PM at a BigCo

This quar­ter, we won’t be re­leas­ing any new fea­tures, and we have no plans to re­design any­thing — we will ex­clu­sively fo­cus on fix­ing bugs.

— Imaginary PM at a BigCo

Until this at­ti­tude changes, the great soft­ware qual­ity de­cay will con­tinue.

That does­n’t sound op­ti­mistic, but I’m ac­tu­ally ex­cited about what comes next. As com­pa­nies col­lec­tively spi­ral into AI debt, in­di­vid­ual de­vel­op­ers have a unique op­por­tu­nity to build soft­ware that would pre­vi­ously have been be­yond their reach.

I have no hope for my car’s Android Auto or LGs stu­pid web­site — but I choose to be­lieve that every­day soft­ware will get bet­ter as a re­sult of this frus­tra­tion. We’re al­ready see­ing acts of re­bel­lion against the cur­rent state of ma­cOS and Windows, and I hope the trend will spread across the stack.

FLUX 3 x mimic: The Next Generation of Video-Action Models

bfl.ai

An early ver­sion of FLUX 3, our new mul­ti­modal foun­da­tion model, is now run­ning on ro­bots. We gave mimic ro­bot­ics early ac­cess to FLUX.3. Their strength in ro­bot learn­ing and de­ploy­ment, com­bined with the mod­el’s world knowl­edge and BFLs foun­da­tion model ex­per­tise, pro­duced FLUX-mimic: the next gen­er­a­tion of video-ac­tion mod­els.

FLUX

FLUX 1 and FLUX 2 gen­er­ate im­ages. FLUX 3 ex­pands into mul­ti­modal­ity and gen­er­ates au­dio-vi­sual con­tent jointly - and, at the same time, pro­vides the foun­da­tion of FLUX-mimic: A video-ac­tion model, de­vel­oped in col­lab­o­ra­tion with mimic, run­ning ro­bots that have been tested and de­ployed at Audi.

At first glance, pro­duc­ing con­vinc­ing vi­sual con­tent and con­trol­ling ro­bots seem to have lit­tle in com­mon. One re­quires gen­er­at­ing pix­els, the other an un­der­stand­ing of how the phys­i­cal world re­sponds when you touch and ma­nip­u­late it. If one model does both, it was never re­ally only a con­tent cre­ation model. It is a model of how the world be­haves, and con­tent cre­ation is one thing one can do with it.

That is what FLUX 3 is.

Video is the hard part

FLUX 3 is one model, jointly trained across im­ages, video and au­dio from the be­gin­ning. The most de­mand­ing part of that train­ing - ac­count­ing for over 95% of the to­tal com­pute costs - is video pre­dic­tion. To gen­er­ate re­al­is­tic videos, a model has no choice but to learn con­tact, mo­tion, weight, cause and ef­fect; get any of them wrong and it looks wrong. Learning to ren­der the world ac­cu­rately means learn­ing how the world be­haves.

Relatively speak­ing, au­dio is the easy modal­ity. Low di­men­sional and far less de­tailed than video, it makes up less than 0.5% of the to­kens in a 720p video with au­dio. Once a model has done the hard work of learn­ing video un­der­stand­ing, it will learn the causal re­la­tion­ship be­tween video and au­dio to pre­dict speech syn­chro­nized to lip move­ment and au­dio ef­fects syn­chro­nized to the phys­i­cal events caus­ing them.

Actions fol­low the same shape: a low di­men­sional rep­re­sen­ta­tion of a ro­bot’s state, tightly cou­pled to vi­sual ob­ser­va­tions. Actions, au­dio and video frames are all par­tial rep­re­sen­ta­tions of a sin­gle un­der­ly­ing phys­i­cal re­al­ity. After the model has learned about the phys­i­cal processes be­hind video and au­dio, ac­tion pre­dic­tion is not a new de­par­ture - it is one more view of the re­al­ity it al­ready mod­els.

A sin­gle back­bone

If that fram­ing is cor­rect, teach­ing FLUX 3 to pre­dict ac­tions should not in­cur last­ing costs: we ex­pect a brief phase of dis­tur­bance as the model has to learn the struc­ture of the ac­tion space and align its in­ter­nal rep­re­sen­ta­tion of the world to it, be­fore re­turn­ing to full per­for­mance. That is ex­actly what we ob­serve.

In a large-scale train­ing run, we added ac­tion pre­dic­tion to the cur­ricu­lum and ob­served the ef­fect on video gen­er­a­tion qual­ity. Human rat­ings on text-to-video and im­age-to-video ini­tially fell by up to 10% as the model started to in­cor­po­rate the new ac­tion modal­ity. After 3500 steps, the model had re­gained its full pre­vi­ous qual­ity on video gen­er­a­tion tasks while now also pre­dict­ing ac­tions.

Each se­ries is nor­mal­ized to its own qual­ity be­fore ac­tion pre­dic­tion was added. Higher is bet­ter.

The model had to in­te­grate ac­tions into its in­puts and out­puts - but do­ing so did­n’t cost it ca­pac­ity per­ma­nently. It merely had to learn how this new modal­ity re­lates to its ex­ist­ing model of the world. Once this was fig­ured out, the per­for­mance penalty on its ex­ist­ing ca­pa­bil­i­ties was gone. Video gen­er­a­tion and ac­tion pre­dic­tion don’t need sep­a­rate foun­da­tions. The same back­bone car­ries both.

This makes Physical AI a nat­ural ex­ten­sion of our roadmap at Black Forest Labs rather than a change in di­rec­tion. Content cre­ation is what our mul­ti­modal FLUX 3 back­bone does with im­age, video and au­dio. Physical AI is what it does with ac­tions. One foun­da­tion model, with vi­sual in­tel­li­gence at its core, en­abling two fam­i­lies of ap­pli­ca­tions. We did­n’t build a sep­a­rate foun­da­tion model. We fo­cused on the hard thing: build­ing a model that un­der­stands the world. Acting in it is what that un­der­stand­ing makes pos­si­ble.

From lab to re­al­ity: FLUX-mimic

What hap­pens when we point the FLUX 3 back­bone at real au­toma­tion tasks on real pro­duc­tion lines?

That’s the ques­tion mimic and BFL cre­ated FLUX-mimic to an­swer. mimic builds their own ro­bots and brings ex­per­tise in ro­bot learn­ing, dex­ter­ous ma­nip­u­la­tion and pro­duc­tion de­ploy­ment; BFL builds vi­sual foun­da­tion mod­els and brings mul­ti­modal train­ing and mod­el­ing ex­per­tise. Together, we built a next-gen­er­a­tion model for gen­eral-pur­pose ma­nip­u­la­tion - adapted to in­dus­try re­quire­ments and in­te­grated into mim­ic’s full-stack de­ploy­ment sys­tem. FLUX-mimic is a video-ac­tion model built on the FLUX 3 back­bone.

Decoding the learned world model

Our the­sis is that FLUX 3 has to learn an in­ter­nal rep­re­sen­ta­tion of the world to be able to gen­er­ate videos. FLUX-mimic fol­lows through on this the­sis and de­codes ac­tions from the learned world rep­re­sen­ta­tion of the FLUX back­bone. This ap­proach, pi­o­neered in mimic-video, trains a light­weight ac­tion de­coder on top of in­ter­me­di­ate fea­tures ex­tracted from the video pre­dic­tion path of FLUX.

Architecture overview how FLUX-mimic is built on top of FLUX 3

The suc­cess of this ap­proach de­pends on two re­lated but dif­fer­ent as­pects: the qual­ity of the world model learned by FLUX and the qual­ity of the fea­ture rep­re­sen­ta­tion of this world model. The qual­ity of the world model is di­rectly re­lated to the gen­er­a­tion qual­ity: if a model does not un­der­stand how the world be­haves, it can­not sim­u­late it. However, even the best world model does not help an ac­tion de­coder if it is in­ac­ces­si­ble: if the fea­ture space keeps the causal re­la­tion­ships be­tween modal­i­ties en­tan­gled non­lin­early, un­der­stand­ing those re­la­tion­ships from the fea­ture rep­re­sen­ta­tion re­mains as dif­fi­cult as un­der­stand­ing them from the raw in­puts - rep­re­sen­ta­tion qual­ity mat­ters.

Generation qual­ity and rep­re­sen­ta­tion qual­ity have long been stud­ied and ap­proached in iso­la­tion from each other. Generative ap­proaches re­sult in high-qual­ity world mod­els that en­able sim­u­la­tions and they ex­hibit scal­ing laws for pre­dictable re­turns on com­pute in­vest­ments. However, com­pared to more spe­cial­ized ap­proaches for rep­re­sen­ta­tion learn­ing they pro­duce less dis­en­tan­gled rep­re­sen­ta­tions, which puts a ceil­ing on their use­ful­ness for tasks that re­quire world un­der­stand­ing.

As gen­er­a­tive mod­els them­selves rely on their own fea­tures, this di­ver­gence in their rep­re­sen­ta­tion qual­ity seems counter-in­tu­itive. Improved rep­re­sen­ta­tions within gen­er­a­tive mod­els should im­prove the qual­ity of their world model and make them more us­able for down­stream tasks. In our work, Self-Flow, we demon­strated how to unify gen­er­a­tion and rep­re­sen­ta­tion learn­ing in a sin­gle frame­work and ob­served ex­actly this rec­i­p­ro­cal im­prove­ment: the world model im­proved - as mea­sured by gen­er­a­tion qual­ity across video, im­age and au­dio - and its rep­re­sen­ta­tion qual­ity im­proved - as mea­sured by suc­cess rate for ro­bot con­trol tasks in sim­u­la­tion.

Self-Flow vs. Flow Matching (FM). Left: gen­er­a­tion er­ror (Fréchet dis­tance) per modal­ity, each nor­mal­ized to FM = 100 (lower is bet­ter). Right: suc­cess rate on ma­nip­u­la­tion tasks av­er­aged over four task groups through fine­tun­ing (higher is bet­ter).

Scaling the world model

Scaling laws re­main true with Self-Flow, and FLUX 3 is the ap­pli­ca­tion of that: the scaled-up ver­sion of Self-Flow. It is trained on tens of mil­lions of hours of gen­eral video con­tent to learn world dy­nam­ics as broadly as pos­si­ble from day one, and on hun­dreds of thou­sands of hours of video con­tent fo­cused on hu­man and ro­bot ma­nip­u­la­tion tasks to be ready as a back­bone for vi­sual in­tel­li­gence. This scal­ing is what trans­lates the suc­cess of Self-Flow from the lab to re­al­ity. mimic de­ployed FLUX-mimic in real fac­tory use cases span­ning the daily re­al­ity of pro­duc­tion and lo­gis­tics work: kit­ting parts into struc­tured trays, in­sert­ing elec­tronic con­trol units into tight-fit­ting fix­tures, as­sem­bling com­po­nents to­gether, and han­dling soft, flex­i­ble ma­te­ri­als like seals and ca­bles that con­ven­tional au­toma­tion has never been able to touch.

Benchmarks demon­strate that the ac­tion de­coder out­per­forms pre­vi­ous vi­sion-lan­guage-ac­tion mod­els, even with a com­pletely frozen FLUX back­bone - a set­ting where pre­vi­ous vi­sion-lan­guage-ac­tion mod­els fail to suc­ceed. This high­lights how scal­ing gives our back­bone strong knowl­edge of the world and how to act in it, and how Self-Flow makes this knowl­edge read­ily de­cod­able from the back­bone’s fea­ture rep­re­sen­ta­tions. When fine­tun­ing the back­bone to­gether with the ac­tion de­coder, FLUX-mimic achieves state-of-the-art suc­cess rates.

Dashed line marks each mod­el’s me­dian suc­cess rate across 20 au­tonomous tri­als. Higher is bet­ter.

From world knowl­edge to a work­ing task

A back­bone ex­pos­ing world knowl­edge in de­cod­able rep­re­sen­ta­tions changes what it takes to teach a ro­bot a new task. If the physics is al­ready in the rep­re­sen­ta­tion and read­ily ac­ces­si­ble, adapt­ing to a task is no longer a mat­ter of teach­ing the model how the world works - it only has to learn how this par­tic­u­lar task maps onto what it al­ready knows. The ex­pen­sive part is done be­fore the ro­bot ever moves.

This shows up di­rectly in how much demon­stra­tion data a new task re­quires. In our Self-Flow ex­per­i­ments, ac­tion pre­dic­tion reached a given suc­cess rate in half the train­ing steps com­pared to a video model with­out Self-Flow - bet­ter rep­re­sen­ta­tions make the world knowl­edge eas­ier to ex­tract, so less data is needed to reach the same ca­pa­bil­ity. The mimic-video pa­per re­ports up to 10x sam­ple ef­fi­ciency for video-ac­tion mod­els over vi­sion-lan­guage-ac­tion mod­els; FLUX-mimic com­bines both ef­fects.

The same ben­e­fit shows up in be­hav­ior. FLUX-mimic nat­u­rally re­cov­ers from fail­ure: a ro­bot that misses a grasp cor­rects it­self, grasps again, and com­pletes the task. No demon­stra­tion set can cover every pos­si­ble way a task can go wrong. Recovery that was never demon­strated has to come from some­where else - from a model that al­ready knows how the world be­haves.

The back­bone’s pre­dicted fu­ture (top), along­side the roll­out the ro­bot pro­duced from the de­coded ac­tions (bottom).

Fast enough to act

Real-world de­ploy­ment sets a hard con­straint: the model has to act as fast as the world moves. The dom­i­nant com­pute cost for FLUX-mimic sits in the back­bone. It is the largest com­po­nent of the model and, in mim­ic’s op­ti­mized de­ploy­ment stack, its la­tency ef­fec­tively sets the ceil­ing for the whole sys­tem.

This is where our method­ol­ogy pays off a sec­ond time. Better rep­re­sen­ta­tions mean more ca­pa­bil­ity per pa­ra­me­ter: a model that has learned a well-struc­tured world model needs less ca­pac­ity to reach a given level of per­for­mance than one that has not. For de­ploy­ment, this trans­lates di­rectly into be­ing able to run a smaller back­bone - and a smaller back­bone is a faster back­bone.

CLIP score at 1.0M train­ing steps; higher is bet­ter. Backbone depth is the dom­i­nant dri­ver of de­ploy­ment la­tency, so fewer lay­ers means a faster model.

As a re­sult, the back­bone of FLUX-mimic can be op­ti­mized to run from in­put to world rep­re­sen­ta­tion in less than 80ms on a sin­gle NVIDIA RTX 5090 GPU - which puts it on the same or­der of mag­ni­tude as hu­man vi­sual re­ac­tion time.

Real-world de­ploy­ments re­quire ad­di­tional op­ti­miza­tions of the full de­ploy­ment stack to avoid adding any ad­di­tional la­tency: mim­ic’s op­ti­miza­tions range from the ac­tion de­coder, through cut­ting in­ter-process la­tency be­tween sen­sors, the model and ac­tu­a­tors, to real-time chunk­ing such that pre­dic­tion and ex­e­cu­tion over­lap and keep the ro­bots run­ning smoothly with­out jit­ter. The end re­sult is a self-con­tained ro­bot sys­tem with re­ac­tion times of 101ms.

On the fac­tory floor

All of this leads back to the place where au­toma­tion mat­ters: the fac­tory. Audi runs one of the most au­to­mated pro­duc­tion net­works in the au­to­mo­tive in­dus­try - which gives it a pre­cise view of where con­ven­tional au­toma­tion still stops. Despite decades of ro­bot­ics in­vest­ment, tasks with flex­i­ble parts and fine ma­nip­u­la­tion have stayed man­ual, largely for eco­nomic rea­sons: the vari­ant di­ver­sity of pre­mium pro­duc­tion makes con­ven­tion­ally pro­grammed ro­bot cells too costly to re-en­gi­neer for each case. Learning-based sys­tems change that math.

In part­ner­ship with mimic, Audi has been test­ing and de­ploy­ing FLUX-mimic. We have seen these ro­bots solve com­plex soft-body ma­nip­u­la­tion work that would have been sim­ply im­pos­si­ble with con­ven­tional ro­bot­ics. This can have a ma­jor im­pact in as­sist­ing our em­ploy­ees, in­creas­ing ef­fi­ciency, and ex­pand­ing flex­i­ble au­toma­tion across pro­duc­tion and lo­gis­tics op­er­a­tions. For us, part­ner­ing with pi­o­neer­ing com­pa­nies such as mimic and Black Forest Labs is es­sen­tial in push­ing the fron­tier of phys­i­cal AI and val­i­dat­ing these in­no­va­tions in real-world pro­duc­tion en­vi­ron­ments.” — Christoph Schneider, Audi Production Lab

Closing

FLUX-mimic is a pur­pose-built ro­bot­ics model - and a proof point for what’s to come. Its sam­ple ef­fi­ciency and ro­bust­ness come from the FLUX 3 back­bone and the qual­ity of the rep­re­sen­ta­tions it ex­poses, not from task-spe­cific en­gi­neer­ing. That is what lets the ap­proach trans­fer across tasks, in­dus­tries, and hard­ware. Read more through mimic.

One model, with vi­sual in­tel­li­gence at its core, gen­er­at­ing im­age, video and au­dio - and dri­ving ro­bots on a pro­duc­tion line. Content cre­ation and phys­i­cal AI are two ap­pli­ca­tions of the same foun­da­tion.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.