10 interesting stories served every morning and every evening.

How I use LLMs to learn complex topics · Laurentiu Raducu

laurentiugabriel.github.io

Many en­gi­neers I know use gen­er­a­tive AI for many func­tions, like build­ing PoCs, in­ter­nal tools or dash­boards, or even learn­ing new stuff. I per­son­ally find the style used by LLMs to ex­plain things dif­fi­cult to fol­low. It’s just too sim­plis­tic and de­pend­ing on the num­ber of emo­jis used, a bit an­noy­ing too.

While I was an­a­lyz­ing new AI bot­tle­necks that might slow down data cen­ter buildup, I re­al­ized there are many as­pects of chip pro­duc­tion that I do not know. Surfing the web, I asked my­self what if there would be a game to get you through the process of build­ing a chip at a fab? For sure learn­ing this way will stick, since you can map con­cepts with ob­jects within the game. This is when I de­cided to try it, and it ac­tu­ally turned out re­ally well.

The flow

Instead of just ask­ing AI to ex­plain a topic, I use the fol­low­ing flow:

In plan mode (using CC, or OpenCode) I ask a model to build the foun­da­tional knowl­edge for X topic.

I ask it to re­view the ac­cu­racy of the knowl­edge base it built in the pre­vi­ous step.

I pro­ceed ask­ing it to build a sim­u­la­tion of that topic in a low-poly, Rollercoaster Tycoon-like an­i­ma­tion. I add some UX el­e­ments as well, like the page needs to be vis­i­ble on both large and small screens, have con­trols to stop the flow when­ever I want etc.

I then push it to a new repo and en­able GitHub Pages for it.

The re­sult

What you get is a beau­ti­ful an­i­ma­tion that is 100% ac­cu­rate and free of hal­lu­ci­na­tions. For me, this method works a lot bet­ter than just read­ing end­less ma­te­ri­als that I find on Google, or try­ing to di­gest a bul­leted list that is spat by a lan­guage model.

I’ve done this specif­i­cally for learn­ing chip build­ing and launch it un­der this web­site: ChipTycoon. You get to fol­low a cart from the mo­ment when sand is col­lected, to the mo­ment when a chip is fi­nal­ized and de­liv­ered to a data cen­ter.

Visually, you can fol­low the cart and see how it changes too. Since it’s low-poly, the de­tails might be miss­ing, but it’s still a good in­di­ca­tor for show­ing how the prod­uct changes once it goes through the many steps re­quired in the man­u­fac­tur­ing process.

How to im­prove it fur­ther

Let’s say that the low-poly de­sign re­quires to much im­mag­i­na­tion to ac­tu­ally vi­su­al­ize what hap­pened to the quartz sand pile af­ter it left the fur­nace. To trans­form this into a more re­al­is­tic rep­re­sen­ta­tion, you can use my skill for trans­form­ing pic­tures into 3d ob­jects, and map the re­sult­ing ob­jects to your sim­u­la­tion. This way you get more ac­cu­rate de­sign.

Also, you can add chal­lenges to your sim­u­la­tion too. Trying to an­swer ques­tions about a pre­vi­ous step in the chip man­u­fac­tur­ing process will help you re­tain the knowl­edge tremen­dously. Add in­tu­itive puz­zles too that will help you learn even bet­ter.

Check out what other pages I cre­ated:

How rocket en­gines are made

How LLMs work

How F1 en­gines are built

How an EUV ma­chine is built

Mea Culpa - Dark Hours

blog.terrygodier.com

Last week I launched a pro­ject called Dark Hours, which was a web­site util­ity to give you an idea of what could be seen in the sky that night.

A de­vel­oper who cre­ated an­other web app called DarkHours.app replied to my com­ment on Bluesky yes­ter­day to show me how sim­i­lar the pro­ject was to his, in­clud­ing the name. The thread is here.

I told him I’d sig­nif­i­cantly dif­fer­en­ti­ate the fea­ture­set, change the name, and write up a blog post to show peo­ple his pro­ject.

About an hour later, once it be­came clear to me that the web app I had launched us­ing Claude was strik­ingly sim­i­lar to his open source pro­ject, even re­pro­duc­ing a bug he had later fixed, it was clear that there was­n’t any­thing to do ex­cept to redi­rect the do­main di­rectly to him and kill any plans I had to launch an iOS app for the pro­ject.

I’d like to give credit where credit is due. I think he’s made a won­der­ful open source app and if you liked the fea­tures that were in­cluded in what I launched, please use the ac­tual ver­sion of the pro­ject. That’s what I’ll be do­ing.

I’d also like to apol­o­gize for my ir­re­spon­si­ble use of AI to build such a thing. While I had gen­uinely never seen DarkHours.app be­fore yes­ter­day, I was care­less in re­ly­ing on AI to gen­er­ate the pro­ject with­out do­ing the work to un­der­stand whether it closely re­sem­bled an ex­ist­ing pro­ject. That’s on me, and I am re­spon­si­ble for what I pub­lished.

Going for­ward, I won’t be us­ing AI in this way to cre­ate any more web stuff, and I do not use it to any­thing close to this ex­tent on iOS soft­ware. I do ask ques­tions, de­bug is­sues, things like that, but I do not cre­ate apps with Claude for iOS.

Windows 11's built-in Weather app wastes more than 1 GB of RAM

www.notebookcheck.net

ⓘ Microsoft

A new re­port shows Windows 11′s built-in Weather app can con­sume more than 1 GB of RAM. By com­par­i­son, Apple’s na­tive Weather app on ma­cOS uses roughly five times less mem­ory un­der sim­i­lar con­di­tions.

Microsoft has been work­ing to make Windows 11 more ef­fi­cient on PCs with lim­ited RAM, but one of its own built-in ap­pli­ca­tions ap­pears to be work­ing against that goal. According to tests pub­lished by Windows Latest, the op­er­at­ing sys­tem’s Weather app can con­sume more than 1 GB of mem­ory de­spite per­form­ing a rel­a­tively sim­ple task.

Windows Latest re­ports that the app ex­ceeded 1.2 GB of RAM while dis­play­ing a weather fore­cast, with no in­ten­sive in­ter­ac­tion from the user. Wccftech ob­served sim­i­lar be­hav­ior, not­ing that mem­ory us­age typ­i­cally starts at around 1 GB, drops to roughly 500 – 600 MB when idle, and can climb to 1.5 – 1.6 GB dur­ing ba­sic ac­tions such as zoom­ing or nav­i­gat­ing the in­ter­face. On a PC equipped with 8 GB of RAM, that means the ap­pli­ca­tion alone may oc­cupy nearly 20% of the sys­tem’s mem­ory.

By com­par­i­son, Apple’s na­tive Weather app on ma­cOS re­port­edly uses less than 250 MB of RAM un­der sim­i­lar con­di­tions, giv­ing Microsoft’s im­ple­men­ta­tion a mem­ory foot­print roughly five times larger.

According to Windows Latest, the high mem­ory con­sump­tion is due to the fact that Weather is not a fully na­tive Windows ap­pli­ca­tion. Instead, it is es­sen­tially an MSN Weather web app built on Microsoft’s WebView2 frame­work. Task Manager shows mul­ti­ple Chromium-based sub­processes run­ning si­mul­ta­ne­ously, which con­tributes to the un­usu­ally high RAM us­age.

The is­sue is un­likely to af­fect high-end PCs with 32 GB or more of RAM, but it could have a no­tice­able im­pact on en­try-level sys­tems. On com­put­ers with 8 GB or even 16 GB of mem­ory, launch­ing the Weather app may in­crease mem­ory pres­sure enough for Windows to rely more heav­ily on the page file, po­ten­tially mak­ing the sys­tem feel less re­spon­sive.

The ap­pli­ca­tion also in­cludes ad­ver­tis­ing within its in­ter­face. According to Windows Latest, spon­sored con­tent is em­bed­ded di­rectly into the fore­cast feed, ap­pear­ing along­side weather cards in a sim­i­lar vi­sual style. While Microsoft li­censes weather data from mul­ti­ple providers — in­clud­ing Foreca, the European Centre for Medium-Range Weather Forecasts (ECMWF), and other re­gional me­te­o­ro­log­i­cal ser­vices — the pres­ence of ads in a built-in Windows ap­pli­ca­tion has at­tracted crit­i­cism.

The find­ings also ap­pear to con­tra­dict Microsoft’s re­cent ef­forts to im­prove Windows 11′s ef­fi­ciency. The com­pany has been up­dat­ing sev­eral built-in ap­pli­ca­tions and has re­peat­edly said it wants the op­er­at­ing sys­tem to per­form bet­ter on lower-end hard­ware. Microsoft ex­ec­u­tive Rudy Huyn has also stated that the com­pany in­tends to de­velop more fully na­tive Windows ap­pli­ca­tions in the fu­ture, al­though it re­mains un­clear whether MSN-branded ap­pli­ca­tions such as Weather will even­tu­ally be re­built us­ing WinUI.

Andrew Sozinow - Tech Writer - 71 ar­ti­cles pub­lished on Notebookcheck since 2024

I’ve been fas­ci­nated by com­put­ers, elec­tron­ics and mod­ern tech­nol­ogy since child­hood. I started writ­ing IT-news at high school and have been do­ing it con­tin­u­ously for more than 10 years. During this time, I have worked for many me­dia out­lets, and now I am a news ed­i­tor at 3DNews. Sometimes I also write smart­phone re­views. In July 2024, I de­cided to try my hand at writ­ing news on Notebookcheck. When I’m not work­ing, I like to play videogames, do puz­zles, and travel.

Andrew Sozinov, 2026 – 08- 9 (Update: 2026 – 08- 9)

A Surveillance ‘Cat-and-Mouse’ Game With AI

www.theatlantic.com

Anthony Bingy” Arillotta waited years to be­come a made man in the Genovese crime fam­ily, and when at last the call came in August 2003, he fol­lowed di­rec­tions to the let­ter. According to sworn tes­ti­mony, Arillotta was sum­moned to a steak house in the Bronx, where he was made to hand over his cell­phone, beeper, and jew­elry be­fore be­ing dri­ven to an apart­ment build­ing. When he got there, he was taken to a small bath­room and strip-searched for elec­tronic de­vices. For his big meet­ing with the boss, he was given a bathrobe to wear.

Until re­cently, only spies and crim­i­nals had to worry this ob­ses­sively about their pri­vate state­ments be­ing picked up by elec­tronic equip­ment. But soon, the av­er­age per­son might need to de­ploy sur­veil­lance coun­ter­mea­sures. The next time you con­duct a del­i­cate bit of of­fice diplo­macy or share a ro­man­tic or fi­nan­cial se­cret with a friend over drinks, a sen­sor built into some­one’s glasses, neck­lace, or lapel pin might be watch­ing you and lis­ten­ing.

In March, the tech start-up Deveillance an­nounced the de­vel­op­ment of Spectre I, a hockey-puck-shaped de­vice that pur­ports to pre­vent oth­ers from record­ing you (no strip search re­quired). The com­pany was founded by Aida Baradari, a re­cent col­lege grad­u­ate who was wor­ried by the surge in peo­ple wear­ing AI-enabled recorders. These wear­ables can be used as a silent note­taker, a per­sonal as­sis­tant, or even a ther­a­pist of sorts. That tech­nol­ogy is­n’t yet main­stream, but it may be soon. Apple—the com­pany with the largest per­sonal-tech ecosys­tem in the world—is ru­mored to be de­vel­op­ing an AI pin or pen­dant that would serve as an iPhone’s con­stant eyes and ears; many other prod­ucts of this type are on the way. AI ac­ces­sories could one day be as wide­spread as AirPods.

New sur­veil­lance tech­nolo­gies tend to breed new coun­ter­mea­sures, which lead, in turn, to more so­phis­ti­cated sur­veil­lance. During the Second World War, af­ter Germany op­er­a­tional­ized radar, the Royal Air Force be­gan drop­ping thin strips of met­al­lized pa­per cut to a spe­cific size that res­onated with the radar, swamp­ing German screens with phan­tom echoes that were in­dis­tin­guish­able from real air­craft. Some his­to­ri­ans have ar­gued that the en­su­ing radar arms race was more con­se­quen­tial to the war’s out­come than the Manhattan Project.

For decades, crude jam­mers have been sold to peo­ple who hope to avoid be­ing recorded. Early ver­sions blasted loud, un­pleas­ant white noise to con­ceal voices. More re­cently, com­pa­nies have made mod­els that emit a steady stream of ul­tra­sonic sound at in­audi­ble fre­quen­cies, ex­ploit­ing a quirk of mi­cro­phone hard­ware that con­verts those high fre­quen­cies into noise. In 2020, a team at the University of Chicago led by Yuxin Chen re­ported that it had mounted 23 ul­tra­sonic trans­duc­ers on a sin­gle bracelet, such that jam­ming sig­nals could be sent in all di­rec­tions in­stead of be­ing fo­cused on a sin­gle tar­get.

Read: The most re­viled tech CEO in New York con­fronts his haters

But even high-tech jam­mers have a hard time fend­ing off to­day’s AI wear­ables. The most ad­vanced pins, pen­dants, and glasses use speech-re­cov­ery al­go­rithms to strip away un­wanted noise, whether it orig­i­nates from every­day sources—such as the clink­ing of glasses in a crowded bar—or from an ul­tra­sonic jam­mer. This task the al­go­rithms per­form is quite dif­fi­cult: In that crowded bar, a mi­cro­phone on a per­son’s lapel will in­ter­cept sound vi­bra­tions from many dif­fer­ent sources at once. It will pick up a bar­tender call­ing out a drink or­der, mu­sic em­a­nat­ing from a speaker, bursts of laugh­ter com­ing from nearby ta­bles—and all of these sounds ric­o­chet off of walls and other ob­jects, cre­at­ing yet more noise. The hu­man body solves this cocktail party prob­lem” with­out us notic­ing: Our ears serve as dual mi­cro­phones, and our brain can use the tim­ing and in­ten­sity dif­fer­ences be­tween them, along with lay­ered pro­cess­ing in the au­di­tory cor­tex, to iso­late the voice of a per­son who is sit­ting across from us.

DeLiang Wang, a com­puter sci­en­tist at Ohio State University, has spent decades train­ing neural net­works to ac­com­plish that same goal, for the pur­pose of im­prov­ing hear­ing aids. By feed­ing the net­works hun­dreds of hours of recorded hu­man voices, he has taught them to rec­og­nize the fre­quen­cies and rhythms of speech. The mod­els build an in­ter­nal rep­re­sen­ta­tion of speech-ness,” and when they en­counter a noisy record­ing, they fo­cus on the parts that match the pat­terns they have learned and then sup­press every­thing else. The most ad­vanced tech­nolo­gies can now in­fer miss­ing syl­la­bles in the way that a reader fills in a redacted word from con­text, al­low­ing them to re­con­struct speech that was­n’t cleanly cap­tured in the first place.

Big tech com­pa­nies are try­ing to do this too. Microsoft has been run­ning an an­nual Deep Noise Suppression Challenge since 2020 to ad­vance the field. (Their in-house team is try­ing to make Teams meet­ings less ex­cru­ci­at­ing.) Other com­pa­nies are work­ing on noise can­cel­la­tion for cell­phone calls and pod­cast soft­ware. This sort of re­search is meant to im­prove the lives of nor­mal users of tech­nol­ogy—as­sum­ing that we pod­cast lis­ten­ers count as nor­mal—but every ad­vance in de-nois­ing can also be used to help an AI as­sis­tant re­cover speech from a jammed record­ing.

Defeating these al­go­rithms may re­quire a dif­fer­ent coun­ter­sur­veil­lance ap­proach al­to­gether. Finn Brunton, a his­to­rian at UC Davis and the co-au­thor of Obfuscation: A User’s Guide for Privacy and Protest, told me that one of the best ways is to iden­tify the data that a de­vice is try­ing to col­lect, and then sup­ply it with a junk ver­sion. The Berlin-based artist Adam Harvey used this strat­egy when he de­vel­oped makeup and cloth­ing that frus­trate fa­cial-recog­ni­tion al­go­rithms. Daniel Howe and Helen Nissenbaum did some­thing sim­i­lar with a browser plug-in called TrackMeNot: Rather than con­ceal­ing a user’s Google searches, the ex­ten­sion con­tin­u­ally runs its own ran­dom­ized de­coy queries in the back­ground, so that what­ever a user ac­tu­ally searched for be­comes lost in a sea of false leads.

People have tried this tech­nique in the realm of au­dio too. Woodrow Hartzog, a law pro­fes­sor at Boston University who stud­ies pri­vacy and sur­veil­lance, told me that early in his le­gal ca­reer, he worked with de­fense at­tor­neys who wor­ried that their jail­house con­ver­sa­tions with clients would be recorded. To fight back, they played babble tapes”—au­dio files lay­ered with 40 tracks of voices in dif­fer­ent ac­cents—in the back­ground.

In 2023, a team led by Ming Gao, now a re­searcher at Nanjing University, used hu­man voices to de­feat speech-re­cov­ery al­go­rithms in a dif­fer­ent way. Its jam­mer, called MicFrozen, is worn by a speaker who does­n’t want to be recorded. It lis­tens as they talk and then gen­er­ates a real-time stream of ul­tra­sonic anti-speech” tuned to the speak­er’s voice, much like the noise-can­cel­la­tion tech­nol­ogy in your head­phones. The de­vice then sends out an­other layer of coun­ter­feit speech-shaped sound to mis­lead any al­go­rithm that tries to re­con­struct what was lost.

Baradari, whose com­pany is work­ing on the Spectre I de­vice, would­n’t tell me ex­actly how her jam­mer’s sig­nals work, but she said that they, too, re­sem­ble speech. The launch video for Spectre I claims that the de­vice will also be able to de­tect the pres­ence of nearby mi­cro­phones. When I asked Baradari how it will do that, she clar­i­fied that her team is still working on that part right now.”

However ef­fec­tive Spectre I turns out to be, it won’t be the end of the record­ing arms race. More ca­pa­ble AI mod­els may even­tu­ally de­ploy some new lis­ten­ing tricks of their own. They may by­pass recorded au­dio al­to­gether. In Stanley Kubrick’s 2001: A Space Odyssey, when two as­tro­nauts re­treat to a sound­proofed pod to dis­cuss dis­con­nect­ing HAL 9000, the ship’s com­puter sim­ply reads their lips through the port­hole. A wear­able pow­ered by a model that’s been trained on enough con­ver­sa­tion footage could, in prin­ci­ple, do the same. In the­ory, it could also stare at a glass of wa­ter be­tween two peo­ple and re­cover their speech from vi­bra­tions on the liq­uid’s sur­face.

AI wear­ables may al­ways have an edge over coun­ter­mea­sures. After all, they’re us­ing a tech­nol­ogy that is a prod­uct of the en­tire speech-pro­cess­ing in­dus­try, which takes in bil­lions of dol­lars in in­vest­ments—not just for AI as­sis­tants but also for hear­ing aids, smart speak­ers, and tele­con­fer­enc­ing tools. Meanwhile, only a few aca­d­e­mics and small com­pa­nies are de­fend­ing us from these tech­nolo­gies. The thing about cat-and-mouse games is that we know how they usu­ally end up for the mouse,” Hartzog said. And in this case, the cat in­cludes some of the most pow­er­ful cor­po­ra­tions to ever ex­ist.”

The Mafia knows what it’s like to be a mouse. By the time Arillotta, the as­pir­ing made man, was told to put on the bathrobe, crim­i­nal or­ga­ni­za­tions had been en­gaged in sur­veil­lance arms races of their own for decades. After law en­force­ment started bug­ging their phones, bosses would con­duct busi­ness in per­son. Sometimes, they’d use a safe house or a ve­hi­cle, but those could be bugged, too, and so sen­si­tive in­for­ma­tion might have been com­mu­ni­cated only dur­ing a walk-and-talk. Eventually, crime fam­i­lies turned to burner phones, and then de­vices with en­cryp­tion. But here, again, they fell prey to the cat.

In 2018, the FBI be­gan se­cretly run­ning Anom, its own en­crypted-phone com­pany. Through in­for­mants, it sold 12,000 de­vices with a spe­cial Anom mes­sag­ing app. Members of Mafia fam­i­lies, mo­tor­cy­cle gangs, and other crim­i­nal or­ga­ni­za­tions treated the phones as a sta­tus sym­bol, and used them to ne­go­ti­ate drug deals, laun­der money, and par­tic­i­pate in all man­ner of other il­le­gal ac­tiv­ity. But the se­cu­rity that they of­fered was a ruse: Every mes­sage that they sent was be­ing in­ter­cepted by the feds.

Taxi drivers rarely die of Alzheimer’s – how complex mental maps and spatial reasoning protect your brain

theconversation.com

Taxi and am­bu­lance dri­vers are less likely than work­ers in al­most any other job to die of Alzheimer’s dis­ease. That was the sur­pris­ing re­sult of a 2024 study ex­am­in­ing the death cer­tifi­cates of nearly 9 mil­lion peo­ple in the U.S.

These find­ings stopped me in my tracks be­cause those two jobs rely on the same thing as my own work: maps.

I have spent more than two decades star­ing at maps. Not pa­per maps on a wall, but dig­i­tal ones with mul­ti­ple lay­ers: flood bound­aries draped over cen­sus blocks, car crash hot spots plot­ted against road geom­e­try, and satel­lite read­ings of rain­fall stitched across river basins. Much of my work as a civil and en­vi­ron­men­tal en­gi­neer is done through GIS — that is, ge­o­graphic in­for­ma­tion sys­tems. Engineers like me hold sev­eral spa­tial re­la­tion­ships in their minds at once, rea­son­ing about where things sit rel­a­tive to one an­other across scales rang­ing from a city block to a whole wa­ter­shed.

I al­ways as­sumed that spa­tial rea­son­ing across map lay­ers was purely pro­fes­sional. But that study on taxi and am­bu­lance dri­vers made me won­der whether all that men­tal work might be do­ing some­thing good to the brain.

Taxi dri­ver brains

Of the 9 mil­lion death cer­tifi­cates from January 2020 to December 2022 that re­searchers ex­am­ined, taxi and am­bu­lance dri­vers had the low­est risk of dy­ing from Alzheimer’s dis­ease out of 443 oc­cu­pa­tions. After ad­just­ing for age, sex, race, eth­nic­ity and ed­u­ca­tion, roughly 1 in 100 taxi and am­bu­lance dri­vers died of Alzheimer’s, com­pared with 1 in 60 peo­ple over­all.

This pat­tern did not ex­tend to other dri­ving jobs. The re­searchers con­cluded that the key to re­duc­ing the risk of Alzheimer’s was not dri­ving it­self but con­tin­u­ous real-time nav­i­ga­tion: the con­stant work of lo­cat­ing your­self in space, track­ing a des­ti­na­tion and up­dat­ing a men­tal map as con­di­tions change. Drivers whose jobs re­lied on fixed or pre­de­ter­mined routes, like bus dri­vers and air­craft pi­lots, did­n’t seem to ex­pe­ri­ence a sim­i­lar ad­van­tage.

Researchers be­lieve the as­so­ci­a­tion be­tween nav­i­ga­tion-heavy work and lower Alzheimer’s risk cen­ters on the hip­pocam­pus, a part of the brain that gov­erns mem­ory and spa­tial nav­i­ga­tion. It’s one of the first brain re­gions that Alzheimer’s dam­ages: Problems with spa­tial nav­i­ga­tion and ori­en­ta­tion are among the ear­li­est signs of the dis­ease, some­times sur­fac­ing be­fore ob­vi­ous mem­ory loss.

In one land­mark 2000 study, neu­ro­sci­en­tists com­pared the brains of li­censed London taxi dri­vers with those of peo­ple who did not drive cabs. Their find­ings pro­vided the first ev­i­dence via struc­tural imag­ing that re­gions of the adult brain can mea­sur­ably change un­der sus­tained nav­i­ga­tional de­mand. To earn a li­cense, London cab­bies must mem­o­rize more than 25,000 streets within a 6-mile ra­dius of Charing Cross, a chal­lenge known as The Knowledge” that takes three to four years.

The re­searchers found that London taxi dri­vers had mea­sur­ably more gray mat­ter in the pos­te­rior hip­pocam­pus, a brain area tied to stor­ing large-scale spa­tial maps. That vol­ume tracked with ex­pe­ri­ence: The longer some­one had dri­ven, the larger that part of the brain. The change was built through prac­tice, not in­her­ited. While peo­ple who are good at nav­i­ga­tion might grav­i­tate to this kind of job, the job it­self does have an im­pact on the brain.

Together, these two stud­ies make a co­her­ent case: Work that in­ten­sively ex­er­cises the hip­pocam­pus may re­shape it, and that re­shap­ing may pro­tect against one of the most feared dis­eases of ag­ing.

Where map spe­cial­ists fit in

Cartographers, ur­ban plan­ners and geospa­tial an­a­lysts spend their work­ing days in sus­tained spa­tial rea­son­ing. A ge­o­graphic in­for­ma­tion sys­tem spe­cial­ist might use a com­puter to over­lap pop­u­la­tion data on flood ex­po­sure maps to find who is at risk, or read satel­lite im­agery to map land cover af­ter a wild­fire. Researchers jug­gle sev­eral lay­ers of data, co­or­di­nate sys­tems and scales at once.

Does spa­tial rea­son­ing through a screen en­gage the hip­pocam­pus the way mov­ing through a real city does?

While a taxi dri­ver nav­i­gates from in­side the scene at street level, GIS re­searchers pic­ture space from above as a map — what cog­ni­tive sci­en­tists call al­lo­cen­tric rea­son­ing. But these two per­spec­tives over­lap in the brain: The hip­pocam­pus also builds maps from out­side view­points, not just from a nav­i­ga­tor’s own po­si­tion.

Research on cog­ni­tive maps points to­ward the same con­clu­sion as the study on taxi dri­vers. In 2023, re­searchers ran a ma­chine learn­ing model on more than 22,500 peo­ple in a na­tional dataset and were able to pre­dict which ZIP codes had higher rates of Alzheimer’s with 84% ac­cu­racy based on how com­plex the en­vi­ron­ment was. Those liv­ing in spa­tially com­plex sur­round­ings — the kind that force ac­tive map build­ing, such as dense street net­works with nu­mer­ous in­ter­sec­tions, di­verse points of in­ter­est and land­marks, and mul­ti­ple path op­tions — were less likely to de­velop Alzheimer’s.

A fol­low-up study tied geospa­tial com­plex­ity in one’s en­vi­ron­ment to greater vol­ume in the brain’s spa­tial nav­i­ga­tion re­gions. The re­searchers hy­poth­e­sized that rou­tinely build­ing cog­ni­tive maps ex­er­cises the very cir­cuitry that Alzheimer’s at­tacks first. Repeatedly en­gag­ing the cog­ni­tive sys­tems in­volved in spa­tial nav­i­ga­tion could help de­lay symp­toms of dis­ease.

However, these two stud­ies fo­cus on where peo­ple live, not the work peo­ple do. Whether spa­tial rea­son­ing through a screen ex­er­cises the same cir­cuitry as real-life city nav­i­ga­tion re­mains untested.

Protecting your brain

The im­pli­ca­tions of whether sus­tained spa­tial rea­son­ing pro­tects the brain reach be­yond map­mak­ers and cab dri­vers. Studies have re­peat­edly found a link be­tween men­tally com­plex oc­cu­pa­tions and de­layed cog­ni­tive de­cline and lower de­men­tia risk, even af­ter ac­count­ing for ed­u­ca­tion.

Research on cog­ni­tive re­serve — the brain’s ca­pac­ity to con­tinue func­tion­ing de­spite dis­ease — can help ex­plain why two peo­ple with a sim­i­lar dis­ease bur­den can show markedly dif­fer­ent lev­els of cog­ni­tive im­pair­ment. If de­mand­ing spa­tial think­ing pro­tects the brain, then how so­ci­eties de­sign school­ing, pro­fes­sional train­ing and re­tire­ment all be­come ques­tions of brain health.

Spatial rea­son­ing can be trained, and train­ing op­por­tu­ni­ties are al­ready wide­spread. GIS and re­mote sens­ing in­struc­tion are avail­able through ge­og­ra­phy, en­gi­neer­ing, pub­lic health and en­vi­ron­men­tal sci­ence pro­grams world­wide.

If sus­tained en­gage­ment with spa­tial rea­son­ing can help the brain strengthen the cir­cuitry that Alzheimer’s at­tacks first, its value reaches well be­yond the tech­ni­cal skills it builds.

Historian Jill Lepore says Silicon Valley misreads science fiction and undermines democracy

techcrunch.com

In her up­com­ing book The Rise and Fall of the Artificial State,” Jill Lepore warns that tech com­pa­nies are in­creas­ingly re­plac­ing the func­tions of de­mo­c­ra­tic gov­ern­ment. This shift, she said, marks a re­turn to tyranny and mys­ti­fi­ca­tion in the form of rule by al­go­rithms, cor­po­ra­tions, ma­chines.”

On the lat­est episode of TechCrunch’s Equity pod­cast, I spoke to Lepore — a Harvard his­to­rian and New Yorker staff writer who re­cently won a Pulitzer Prize for her his­tory of the U.S. Constitution — about the evo­lu­tion of what she de­scribed as the idea that we should live un­der an ar­ti­fi­cial state or gov­ern­ment by ma­chines.”

I’m not an anti-tech­nol­o­gist,” Lepore in­sisted. Instead, she said, My beef is the ways in which pri­vate cor­po­ra­tions have in­creas­ingly taken on the func­tions of the state.”

While Lepore’s book ex­am­ines tech­no­cratic philoso­phies that go back cen­turies, she ar­gued that many of Silicon Valley’s charismatic or not-so-charis­matic lead­ers” — es­pe­cially Elon Musk — seem to be ush­er­ing in a fu­ture pulled from mis­read pulp sci­ence fic­tion and comic books.

But what’s funny about Musk is, the stuff he likes ac­tu­ally com­pletely de­feats and de­fies all of his po­lit­i­cal be­liefs,” she said.

Our con­ver­sa­tion also cov­ered Apple’s fa­mous 1984” Macintosh ad, why it’s bananas” to call Twitter a dig­i­tal town hall, and the cur­rent data cen­ter back­lash. Keep read­ing for high­lights, edited for length and clar­ity.

So you’ve prob­a­bly had to do this a lot al­ready, but can you ex­plain what you mean by the artificial state”?

By the ar­ti­fi­cial state, I mean a kind of state that is re­plac­ing the lib­eral de­mo­c­ra­tic na­tion-state in the United States and around the world. It’s both a real thing, a con­struct, but it’s also an idea.

And so, in this book The Rise and Fall of the Artificial State,” I trace the rise of the idea that we should live un­der an ar­ti­fi­cial state or gov­ern­ment by ma­chines. I also trace the no­tion that this is an in­evitable fail­ure, that the ar­ti­fi­cial state can­not sur­vive, and I trace that idea through sci­ence fic­tion.

At one point, you say the rise of the ar­ti­fi­cial state marks the end of cen­turies of democ­racy and equal rights, and it’s a re­turn to tyranny and mys­ti­fi­ca­tion in the form of rule by al­go­rithms, cor­po­ra­tions, ma­chines.” Can you just say a lit­tle bit more about why you see it in such stark terms?

Yeah, I do have a pretty neg­a­tive view of it, and I think it’s im­por­tant to dis­tin­guish the ar­ti­fi­cial state from tech­nol­ogy it­self or modes of tech­nol­ogy. I’m not an anti-tech­nol­o­gist. I’m mar­ried to a com­puter sci­en­tist. I’m re­ally ex­cited about all kinds of in­tel­lec­tual rev­o­lu­tions that we’re in the midst of right now.

That’s not my beef, right? My beef is the ways in which pri­vate cor­po­ra­tions have in­creas­ingly taken on the func­tions of the state. No one con­sented to that. This has been a kind of grad­ual, largely ac­ci­den­tal trans­for­ma­tion of how many na­tion-states around the world work — it’s hap­pened first in the United States.

I think of­ten these in­no­va­tions in bring­ing new tech­nolo­gies to the op­er­a­tions of gov­ern­ment have been ex­tremely well in­ten­tioned; they orig­i­nate with an in­ter­est in ef­fi­ciency and speed and cheap­ness. And then, I think, only in the last 20, 25 years or so have these de­ci­sions been pur­pose­ful and de­lib­er­ate as a kind of usurpa­tion of the role of the na­tion-state.

And that’s not my spec­u­la­tion. You hear a lot of a lot of very promi­nent tech en­tre­pre­neurs talk about want­ing to move be­yond the era of the na­tion-state. […] A lot of fu­tur­ists in the 90s were lib­er­tar­i­ans, and they had a spe­cific in­ter­est in us­ing the ad­vance of the in­ter­net and the suc­ces­sive in­no­va­tions that fol­lowed as a means to erad­i­cate the na­tion-state.

You talk about, on the one hand, the tech­nolo­gies them­selves, and then also the philoso­phies be­hind them, the role the cor­po­ra­tion has in­creas­ingly played. I’m cu­ri­ous to what ex­tent we can sep­a­rate them. Can we ac­tu­ally have a ver­sion of the in­ter­net and of so­cial me­dia that does­n’t nec­es­sar­ily lead to this fu­ture that it seems like we’re [currently] hurtling to­wards?

Absolutely. I’m a his­to­rian. I’m not a tech writer. I’m not a tech jour­nal­ist. I’m not a com­puter sci­en­tist. I’m a his­to­rian, and I’m chiefly a po­lit­i­cal his­to­rian, though I’m also a lit­er­ary his­to­rian. And so, one of the things that I’m re­ally in­ter­ested in un­rav­el­ing for read­ers in this book is all the what-ifs, all the al­ter­na­tives, the paths along the road that were not taken and why.

There was, of course, a re­ally avid dis­cus­sion in the 1990s about what the in­ter­net should look like when it was opened up, and what we ended up with, the 1996 Telecommunications Act — I think, a lot of peo­ple would say [that] just was a mis­take, not an act of sin­is­ter in­tent, right?

But it was a prod­uct of a par­tic­u­lar po­lit­i­cal mo­ment, re­ally was deeply in­flu­enced by Newt Gingrich and his Contract with America, and it’s been very dif­fi­cult to re­visit. I think it’s worth think­ing about what were the al­ter­na­tives that were in play at the time.

And you could say the same thing about the per­sonal com­puter. So, to the de­gree that we can lo­cate an ori­gin point for the promise that bet­ter com­puter tech­nol­ogy would make for bet­ter democ­ra­cies, I think the mo­ment you would first look to would be January 1984, that Super Bowl ad that Apple ran for the re­lease of the Macintosh, with the the sort of George Orwell, 1984 [theme]. Apple was re­ally big on the idea that main­frame com­put­ers rep­re­sented to­tal­i­tar­i­an­ism. They were try­ing to dis­man­tle the gi­ant gray IBM ma­chines, as in rep­re­sent­ing them in that ad as a to­tal­i­tar­ian state. And the lithe, beau­ti­ful, quick, adorable, per­sonal Macintosh would be the ax that would de­stroy that ma­chine and would usher in a new era in which 1984 would not be 1984.’”

That was clever ad­ver­tis­ing. I doubt that any­body at Apple re­ally be­lieved the per­sonal com­puter was go­ing to be an in­stru­ment of per­sonal lib­er­a­tion. I mean, it was go­ing to make pos­si­ble a lot of cool things. I re­mem­ber when I got my first Macintosh — it cer­tainly was­n’t 1984, but it was re­ally cool, it was re­ally fun, it was re­ally ex­cit­ing, I did a lot of things on it. It would never oc­cur to me that it was im­prov­ing my ca­pac­ity for cit­i­zen­ship or my abil­ity to func­tion bet­ter in civil so­ci­ety. It was a cool tool.

But if you wind the reel for­ward in time, down to 2026 — stops along the way in­clude the 2016 elec­tion when Facebook News, in re­sponse to its crit­ics, es­tab­lishes a Supreme Court. You get to last year, when Anthropic hired a moral philoso­pher to write a con­sti­tu­tion. You get to re­cently, when Sam Altman was on Joe Rogan and said [in re­sponse to a ques­tion from Rogan], Oh, an AI pres­i­dent would be a great idea.”

In some ways, they’re silly ex­am­ples. But you see the ways in which these cor­po­ra­tions, these tech com­pa­nies from Silicon Valley, and es­pe­cially their charis­matic or not-so-charis­matic lead­ers, are just tak­ing on the trap­pings of the na­tion-state and the func­tions of democ­racy.

They’re not peo­ple with a so­phis­ti­cated po­lit­i­cal phi­los­o­phy, but it’s like a car­toon ver­sion of that 1984 Macintosh ad, ex­cept that it takes it­self so se­ri­ously. And these com­pa­nies have so much power.

But that said, the book does­n’t be­gin in 1984. I just think that’s a good ex­am­ple of our mod­ern era and the way a fun ad­ver­tis­ing cam­paign turns into a kind of delu­sional fan­tasy on the part of peo­ple like Sam Altman.

You [also] talk about the promise of the quote-un­quote Twitter rev­o­lu­tion,” and this idea that it would bring democ­racy every­where. I can’t help but let that color the way I [react] when Sam Altman or some other AI CEO now says that AI is go­ing to bring all these in­cred­i­ble gifts — and there­fore, if you stand in the way, you’re stand­ing in the way of progress, in the way of his­tory.

To what ex­tent should we just dis­miss all these claims out-of-hand, or are there ways that it might come true?

I mean, Twitter is ac­tu­ally a good ex­am­ple, right? When it was launched, when Jack Dorsey started it, it did­n’t an­nounce it­self as, We’re go­ing to save hu­man­ity, we’re go­ing to res­cue hu­man civ­i­liza­tion from ex­tinc­tion.” It was kind of a goof, and I think peo­ple that used Twitter re­ally early on were like, You know what? It was ac­tu­ally re­ally fun.” It was like, I made a tuna fish sand­wich to­day. What did you have for lunch?” Twitter as a com­pany did not launch it­self on a stage say­ing, We’re here to save democ­racy.”

And re­ally, what hap­pened was that politi­cians, elected of­fi­cials be­gan us­ing Twitter in ways that en­hanced their po­lit­i­cal power, in ways that am­pli­fied their mes­sages, in ways that al­lowed them to reach a younger au­di­ence, in ways that al­lowed them to have a con­stant con­nec­tion with an au­di­ence. Politicians and po­lit­i­cal cam­paigns re­ally kind of con­vinced Twitter — at least in­so­far as I see them, I don’t have an in­side ac­count of the com­pany — but some­what be­grudg­ingly, Twitter came around to like, Twitter’s got­ten so big, and peo­ple post about pol­i­tics so of­ten that it’s al­most like Twitter is a town hall.”

By the time you get to, I think it’s 2012 — many years into Twitter’s fairly short his­tory — they pub­lish this thing called the Twitter Politics and Elections Handbook, which is re­ally a guide for po­lit­i­cal can­di­dates and elected of­fi­cials and how to most ef­fec­tively use Twitter. And then they be­gin the roll­out of, It’s a town hall in your pocket, and it’s im­prov­ing our democ­ra­cies be­cause we’re restor­ing the de­funct New England town meet­ing,” and that’s all just ba­nanas.

Objectively, noth­ing could be fur­ther from the truth. At that time, one in five Americans had a Twitter ac­count. Most peo­ple who had Twitter ac­counts had never used them, and above 90% of all tweets about pol­i­tics were posted by fewer than 10% of the peo­ple that did use Twitter all the time. There was no way in which Twitter was a rep­re­sen­ta­tion of the elec­torate. Twitter was a rep­re­sen­ta­tion of the most ex­treme, po­lit­i­cally ac­tive, hy­per-par­ti­san among Americans, who were fol­low­ing pol­i­tics re­ally avidly. Looking at it now, we can see, Well, that’s re­ally just a dis­tor­tion ma­chine. And if politi­cians are us­ing it to gauge the elec­torate, they’re get­ting re­ally bad in­for­ma­tion.”

Again, you can say Twitter was not try­ing to par­tic­i­pate in the ar­ti­fi­cial state or un­der­mine democ­racy. Twitter is try­ing to do busi­ness and get more users and sell more what­ever. But it had these un­in­tended con­se­quences that then it sort of set­tles into and be­comes com­fort­able with.

I want to talk a lit­tle bit more about the struc­ture of the book. Like you said, it starts with this his­tory of tech­nol­ogy, his­tory of ideas, and the sec­ond half is about sci­ence fic­tion. Can you say more about how that struc­ture came to you and why you wanted to ad­dress things that way?

I be­came re­ally in­ter­ested, on the one hand, in how of­ten sci­ence fic­tion sto­ries pre­dict the ar­rival of what I then came to call the ar­ti­fi­cial state, and so I re­ally wanted to iden­tify a lit­er­ary tra­di­tion that I think of as the para­ble of the ar­ti­fi­cial state, in which ma­chines get more and more so­phis­ti­cated, they take over more and more of the func­tions of hu­mans, in­clud­ing the func­tions of gov­ern­ment, and even­tu­ally they come to rule the hu­mans, and then maybe they de­stroy all the hu­mans be­cause they don’t re­ally need them any­more.

Maybe they just en­slave them, it kind of de­pends. Are we in The Terminator” or are we in Battlestar Galactica”? There’s dif­fer­ent ver­sions, and these sto­ries go way back. They go back to the 1850s and the early decades of ru­mi­na­tion about the con­se­quences of in­dus­tri­al­ism.

I think a lot of peo­ple — this is cer­tainly true of my stu­dents, my un­der­grad­u­ates — re­ally be­lieve that tech­no­log­i­cal change equals progress. And not only that, but the only kind of progress is tech­no­log­i­cal change. That’s a nov­elty in hu­man his­tory. That’s an in­tel­lec­tual in­ven­tion of the 19th cen­tury, and it is partly be­cause tech­no­log­i­cal change was ac­cel­er­at­ing right at the time that Charles Darwin was de­vis­ing and then pub­lish­ing his the­ory of evo­lu­tion.

So there’s kind of a weird mar­riage be­tween evo­lu­tion as progress and tech­no­log­i­cal change as progress, and what drops out of that are all other, ear­lier no­tions of progress, which chiefly in­volve moral progress — like, things are get­ting bet­ter be­cause peo­ple are be­com­ing bet­ter, or things are get­ting bet­ter be­cause peo­ple are more free.

There are a lot of other ways we might think about progress, but what dom­i­nates to­day is this 19th-century no­tion of tech­no­log­i­cal progress as the only kind of progress, and there­fore all tech­no­log­i­cal change is progress, as op­posed to — ob­jec­tively, it’s only progress if things are get­ting bet­ter.

But in any event, that con­flu­ence in the 19th cen­tury of the idea of tech­no­log­i­cal progress and the idea of evo­lu­tion meant that peo­ple who were think­ing clearly were like, Well, if the ma­chines keep get­ting bet­ter and faster and able to do more things — not just la­bor, but maybe talk or think or move around — what if they evolve to be­come bet­ter at every­thing than we are? Not just bet­ter at run­ning a loom, not just faster at mov­ing through time and space like a rail­road car?” And with that grew an in­cred­i­ble anx­i­ety that found form in sci­ence fic­tion again and again and again and again and again.

My fa­vorite one of these sto­ries was pub­lished, I think, in 1909 by E. M. Forster, right around when he was writ­ing A Room with a View.” He wrote this story called The Machine Stops,” which could be sub­ti­tled, The Room Without a View.” He imag­ines a near fu­ture in which every­body just lives in these rooms, these lit­tle cells. You never see other peo­ple be­cause every­thing you need comes right to your room. It’s like DoorDash, your food is de­liv­ered, you have a screen where you can com­mu­ni­cate with other peo­ple. All your needs are met.

The thing that peo­ple fear most is the nat­ural world. No one wants to ever see the sun, it’s a lit­tle Matrix”-y, and they all wor­ship the ma­chine that or­ga­nizes their lives and brings to them in their cubby-like rooms all the things that they need. The story is about, Humans have be­come es­sen­tially slaves of the ma­chine, which is stronger, more pow­er­ful, and has more ca­pac­ity than hu­mans do, and hu­mans have lost what ca­pac­ity they had.” And then the cli­max of the story is when the ma­chine stops.

If read­ers were to go look at that story, it feels like it could be writ­ten to­day, ex­cept that it’s less sci­ence fic­tion-y to­day than it is the di­ary of a very un­happy YouTuber.

You con­nect that thread to some of the folks run­ning com­pa­nies and ar­guably run­ning as­pects of our gov­ern­ment to­day, like Elon Musk. Essentially, you sug­gest that they’re very bad sci­ence fic­tion read­ers. They read a lot of warn­ing sto­ries, or at least am­biva­lent sto­ries, as if they were man­u­als for the fu­ture.

This is some­thing I wres­tle with a reader of sci­ence fic­tion — some­one who loves Isaac Asimov, for ex­am­ple. I think it’s true that when Musk or Altman is just un­am­bigu­ously be­ing like, Yes, this story is a tem­plate for what I should do with my com­pany,” that’s bonkers. But there is [also] this tech­no­cratic lib­er­tar­ian thread in sci­ence fic­tion that they are pick­ing up on. It’s not some­thing that they’re mak­ing up out of whole cloth, right?

Although weirdly, that’s Heinlein. That’s not Asimov, that’s not Douglas Adams.

Sure, there is that thread in sci­ence fic­tion. I don’t know, I guess [Jeff] Bezos is a big Robert Heinlein fan. You could say, Okay, that lines up well. They’re read­ing it lit­er­ally, but at least they’re get­ting the po­lit­i­cal mes­sage that any ra­tio­nal per­son could find within that lit­er­ary work.”

But what’s funny about Musk is, the stuff he likes ac­tu­ally com­pletely de­feats and de­fies all of his po­lit­i­cal be­liefs.

You also say, re­peat­edly, that the ar­ti­fi­cial state in its cur­rent form is in­com­plete and doomed to fail­ure. Why is it doomed to fail­ure?

This is some­thing that’s fore­seen in all the sci­ence fic­tion that I dis­cuss.

It’s not an Asimov story, but it’s one of Asimov’s [favorite] sto­ries from his boy­hood [“The Man Who Awoke” by Laurence Manning] about a fu­ture in which the foresters have de­feated the wasters. […] The war that the fu­ture hu­mans had was be­tween the wasters, who just fig­ured you could just use every­thing up and waste it, and the foresters, who re­ally be­lieved in — we would call re­for­esta­tion and rewil­d­ing.

That’s gen­er­ally the ten­sion in these sto­ries. It’s be­tween the ar­ti­fi­cial state and the nat­ural world. To erect an ar­ti­fi­cial state and rule hu­mans within it, you must alien­ate them from the nat­ural world be­cause you are de­stroy­ing it. The ar­ti­fi­cial state will de­stroy the nat­ural world, and yet it needs the re­sources of the nat­ural world to run.

So, it is doomed in the sense that there is not a pos­si­bil­ity that the nat­ural world, a hab­it­able planet — hab­it­able for hu­mans — can sur­vive the full con­struc­tion and re­liance on the de­vices of the ar­ti­fi­cial state. That’s how the sci­ence fic­tion works, in any event.

Like you said, you’re a his­to­rian, not a politi­cian or a fu­tur­ist. But what do you think the de­feat of the ar­ti­fi­cial state looks like? Is it ba­si­cally just dis­man­tling all these com­pa­nies, tear­ing down the data cen­ters? Or is there a fu­ture that’s more about bring­ing it un­der con­trol?

I mean, I don’t have a play­book here, ex­cept for the rec­om­men­da­tion that we live in a democ­racy where de­ci­sions have to be made in con­sul­ta­tion with the gov­erned, and these de­ci­sions are not pop­u­lar.

You see this in all the lit­tle data cen­ter crises, town to town, county to county, state to state —  which are partly a con­se­quence of the de­cline of lo­cal news­pa­pers and the de­struc­tion of jour­nal­ism that has been one of the many con­se­quences of so­cial me­dia, and in the case of [Mark] Zuckerberg, I think a some­what in­ten­tional con­se­quence.

What you see is a lot of peo­ple show up at these town meet­ings and say, We don’t even have hous­ing. We don’t have health­care. We don’t have jobs. Who said we’re build­ing this data cen­ter? I need to know a lot more about it. I need to know what its en­ergy costs are go­ing to be. Tell me about the wa­ter con­sump­tion. Are there go­ing to be jobs? Are the jobs go­ing to be long last­ing? Are they just go­ing to be for six months? What’s go­ing to hap­pen to the egrets that live in this area?” Whatever it is that peo­ple want to know.

More and more, you see peo­ple are — like in the Salt Lake ex­am­ple, where well over 70% of the peo­ple re­ally were op­posed to this data cen­ter, and their rep­re­sen­ta­tives sup­ported it. That’s not rep­re­sent­ing the peo­ple. I think there are po­lit­i­cal costs, and we’ll be­gin to see those at elec­tions.

Or maybe we won’t. Enough of de­mo­c­ra­tic func­tion­ing has to be in­tact for peo­ple to ac­tu­ally be able to re­spond to malfea­sance on the part of their rep­re­sen­ta­tives.

Part of your think­ing about [the ar­ti­fi­cial state] started with this great piece you wrote more than a decade ago for The New Yorker, about Clayton Christensen, cri­tiquing his idea of the in­no­va­tor’s dilemma and dis­rup­tive in­no­va­tion — which is very closely as­so­ci­ated with TechCrunch, be­cause we have a big con­fer­ence called Disrupt.

Ten years on, how do you feel about that idea of dis­rup­tive in­no­va­tion?

I stand by every­thing in that piece. [At the time, Lepore wrote, Disruptive in­no­va­tion is a the­ory about why busi­nesses fail. It’s not more than that. It does­n’t ex­plain change. It’s not a law of na­ture.” Christensen re­sponded that Lepore broke all the rules of schol­ar­ship that she ac­cused me of break­ing.”]

I reread it last sum­mer when I was work­ing on this book. What I would say here is, try­ing to be a peace­able hu­man be­ing, I think it re­ally is a prob­lem that his­to­ri­ans have not en­gaged with these ideas. One of the rea­sons I wrote that ar­ti­cle about dis­rup­tive in­no­va­tion — which was not an idea of mine, it was an as­sign­ment […] — was be­cause I just felt like, Disruptive in­no­va­tion is a the­ory of his­tory. It’s a the­ory of his­tor­i­cal change, and it’s based on ev­i­dence from the archives.” And I just thought, as a his­to­rian, it makes no sense. His use of ev­i­dence is com­pletely un­ac­cept­able by any proper un­der­stand­ing of his­tor­i­cal method. Its ar­gu­ment is in con­ver­sa­tion with no mean­ing­ful un­der­stand­ing of how change hap­pens.

So I went and re­did the re­search, and it just did not stand up at all. I felt like I had to write it. And I wish that I felt like there were more en­gage­ment, in the years since, of aca­d­e­mic his­to­ri­ans think­ing through the na­ture of change — which are ques­tions that gen­uinely and au­then­ti­cally in­ter­est peo­ple who are in­volved in de­vel­op­ing new tech­nolo­gies.

People re­ally want to think [about], What is this? What am I do­ing? What are go­ing to be the con­se­quences? Is there any­thing I could learn from his­tory? What hap­pened when the au­to­mo­bile re­placed the horse? What hap­pened to the law? How did we end up with dri­ver’s li­censes? How did we end up with traf­fic law? We did­n’t have traf­fic rules be­fore the au­to­mo­bile. We did­n’t have cer­tain kinds of in­sur­ance sys­tems. We did­n’t have dri­ver’s tests. How did those things emerge? How did [we de­velop] those guardrails on a tech­nol­ogy that was tremen­dously ex­cit­ing, im­proved peo­ple’s lives in many many ways, ut­terly changed the land­scape, rev­o­lu­tion­ized tort law? Maybe I should think about that.”

I just wish that his­to­ri­ans were more in con­ver­sa­tion with tech­nol­o­gists over these years, and with en­tre­pre­neurs. Not just be­cause we can stand around and say, You know, I have a lec­ture to of­fer you on his­tory,” but I think there’s a real con­ver­sa­tion to be had.

All of which is just to say, thanks for hav­ing me on.

When you pur­chase through links in our ar­ti­cles, we may earn a small com­mis­sion. This does­n’t af­fect our ed­i­to­r­ial in­de­pen­dence.

A Falling Weight Just Broke the Sound Barrier with Tom Stanton's Supersonic Trebuchet

www.techeblog.com

Tom Stanton has spent years chas­ing a num­ber that grav­ity it­self seemed to for­bid. On a quiet field some­where in the UK, a 40-kilogram mass dropped a short dis­tance, spun a car­bon-fiber arm past 2,300 rev­o­lu­tions per minute, and sent a 4-gram pro­jec­tile into the air at 776 miles per hour. That is nine miles per hour past the speed of sound. For the first time, a purely grav­ity-pow­ered tre­buchet crossed the bar­rier.

Medieval en­gi­neers cre­ated these ma­chines to fling huge stones at cas­tle walls. The ba­sic idea is to hoist a large heavy weight, let it fall down, and then use the lever and sling to redi­rect that en­ergy into a much lighter pro­jec­tile. Unfortunately, physics gets in the way. A weight in free fall ac­cel­er­ates at a max­i­mum of 9.81 me­ters per sec­ond squared; drop one from 2 me­ters and it smacks into the ground at around 6 me­ters per sec­ond. No mat­ter how heavy you make the weight, the speed re­mains con­stant, and a typ­i­cal arm is like a car with a bike chain stuck in first gear, with lots of torque at first but just enough to get it mov­ing by the end.

Stanton de­vised a cre­ative en­gi­neer­ing so­lu­tion to the speed prob­lem. The coun­ter­weight is sus­pended from a pul­ley sys­tem with a 3:1 ra­tio, as a large di­am­e­ter at the be­gin­ning pro­pels the arm with plenty of force, and as it winds down, the smaller di­am­e­ter end pro­vides a sig­nif­i­cant jump in ro­ta­tional speed. In the end, the drum was re­designed so that weight could be lifted up to 1.9 me­ters, stor­ing more power in the sys­tem than pre­vi­ous it­er­a­tions.

The arm has to be both in­cred­i­bly light and su­per rigid. Carbon fiber was re­ally the only op­tion be­cause it is­n’t heavy enough to weigh down the en­tire sys­tem while yet be­ing able to with­stand pun­ish­ment. Stanton used his home­made CNC mill to carve the piece, keep­ing the dust un­der con­trol with a HEPA vac­uum and a good spray of wa­ter. It weighs only 116 grams. Stress tests and nu­mer­ous de­fec­tive 3D printed pro­to­types re­vealed that it would buckle un­der the sling’s pres­sure, so he mod­i­fied the de­sign to com­pen­sate, re­moved ma­te­r­ial from the ten­sioned side, tough­ened up the op­po­site side, and added some ex­tra brac­ing for good mea­sure. Aluminum hubs con­nect the arm to a short counter-arm, keep­ing the spin­ning bit as bal­anced as pos­si­ble.

The pro­jec­tile now starts near the axle, wrapped firmly in a sling that un­wraps at just the right mo­ment. The en­tire sys­tem is me­chan­i­cally re­leased, us­ing a spring-loaded catch that opens af­ter a cer­tain num­ber of rope ro­ta­tions. The re­lease win­dow is only a few mil­lisec­onds long, which is plenty of time to com­plete the task. Early it­er­a­tions failed un­der stress, but the fi­nal pin-and-loop struc­ture held up.

Testing was care­fully in­creased, and he be­gan with a 10 kg weight. The mod­i­fied aero­dy­namic arm reached a re­spectable 1248 rpm and launched at 394 mph with an in­cred­i­ble 43.7% ef­fi­ciency. Twenty and thirty kilo­grams passed thru with­out a hitch. 40 kg, on the other hand, sped the arm to 2336 rpm and the mis­sile to a blis­ter­ing 716 mph, falling only 51 mph short of the magic bar­rier. The ma­chine even­tu­ally snapped, and the clasp shat­tered be­neath the weight of 50 kg. Stanton re­duced the pro­jec­tile weight to ap­prox­i­mately 4 grams, changed the drum ta­per once more, and re­turned to the test area with the 40 kilo­gram setup.

The fi­nal test is ob­vi­ously the most im­por­tant, since the arm spun up to 2342 rpm. Tip speed reached a stag­ger­ing 274 mph. The high-speed footage was truly eye-open­ing, as the mis­sile trav­eled 1.94 me­ters in 5.6 mil­lisec­onds. Crunching those sta­tis­tics yields 346.4 me­ters per sec­ond, or a more than re­spectable 776 mph. To top it all off, there was a loud crack fol­lowed by a pleas­ant echo, con­firm­ing the sonic boom to every­one within earshot. Not one as­pect of the ma­chine came close to fail­ing, or so we’d like to think.

What Happened to HackerOne?

blog.teknogeek.io

So…what’s go­ing on at HackerOne lately? It might be time for a well­ness check.

If you are new to the bug bounty space (1 – 3 years), you might not have any idea what I’m talk­ing about.

But as a prop­erly washed-up bug bounty hunter who lived through the golden era of HackerOne, I think it’s time to ad­dress the ele­phant in the room.

For some con­text, I started as a hacker on HackerOne in 2017. When I be­gan work­ing in tech, that hands-on ex­pe­ri­ence was ex­tremely use­ful for man­ag­ing a bug bounty pro­gram, since I knew what re­searchers wanted, and how to in­ter­act with them.

As a re­sult, I have man­aged mul­ti­ple large bug bounty pro­grams on HackerOne across var­i­ous com­pa­nies from 2018 to 2025 and I’ve been on both sides of the equa­tion.

What I’m about to talk about comes from first-hand ex­pe­ri­ence, both as a re­searcher and as a bug bounty pro­gram man­ager, and many, many years of di­rect con­ver­sa­tions with HackerOne, both pub­licly and pri­vately.

Background

#

To start, I think it’s im­por­tant to re­al­ize what HackerOne was orig­i­nally de­signed to be.

In 2011, two eth­i­cal hack­ers, Jobert Abma and Michiel Prins, set out to find se­cu­rity vul­ner­a­bil­i­ties in 100 of the largest tech com­pa­nies. They suc­ceeded and found bugs in Google, Facebook, Apple, Microsoft, Twitter, and many oth­ers. At this point in time, the land­scape for eth­i­cal se­cu­rity re­search was risky, legally du­bi­ous, and very scary for se­cu­rity re­searchers.

Not only was there sig­nif­i­cant per­sonal li­a­bil­ity, but there had been mul­ti­ple in­stances of hack­ers be­ing crim­i­nally charged and sen­tenced to jail time for find­ing and re­port­ing se­cu­rity vul­ner­a­bil­i­ties prior to this. Much of this was due to spe­cific ar­bi­trary lines drawn in the sand which, if crossed, made you a bad ac­tor, but if not crossed, made you a po­ten­tially bad ac­tor but tech­ni­cally not one.

Bug Bounty Platforms like HackerOne were de­signed to di­rectly ad­dress this is­sue. It cre­ated a safe mu­tual space for com­pa­nies and hack­ers to con­nect, and it paved the way for eth­i­cal hack­ers to sub­mit se­cu­rity vul­ner­a­bil­i­ties to com­pa­nies, with full con­sent, and get paid for that work. This was a huge mile­stone. You no longer had to worry about get­ting dragged to court (or jail) for find­ing an IDOR that leaked cus­tomer data. Instead, you got a thank you” and a cash pay­out for mak­ing every­one a lit­tle safer.

This op­er­at­ing model was the foun­da­tion for bug bounty and re­mained that way for the next 5+ years.

The Golden Age of HackerOne and Live Hacking Events

#

During this pe­riod, there was a very strong and ex­plicit fo­cus for the busi­ness: how do we make this the best pos­si­ble prod­uct for hack­ers?

The peo­ple run­ning the busi­ness day-to-day were hack­ers, hacker-ad­ja­cent, and most (if not all) were face-to-face with hack­ers on a reg­u­lar ba­sis.

From 2017 to 2020, HackerOne was do­ing Live Hacking Events (LHEs) every few months. These were ex­clu­sive events where the top bug bounty re­searchers around the world would fly into a lo­ca­tion, be given a tar­get, and go ab­solutely ham find­ing crit­i­cal vul­ner­a­bil­i­ties. LHEs were a huge value prop for pro­grams. During a 1 – 3 day pe­riod, you would get more high and crit­i­cal se­cu­rity re­ports than you would have re­ceived for the whole year oth­er­wise.

Every event had a 1-of-1 cus­tom de­signed poster with graph­ics, hacker user­names, stick­ers and chal­lenge coins. It’s hard to over­state what an in­cred­i­ble and pro­duc­tive pe­riod this was for HackerOne and their top pro­grams.

These events were ex­clu­sive and highly cov­eted; in­vites and +1s were prac­ti­cally their own cur­rency. And the en­vi­ron­ment at these events was sur­real. You would be given free flights and ho­tels around the world, and spend a few days sur­rounded by the best and most skilled bug bounty hunters in the world. These re­searchers would reg­u­larly find some of the most im­pact­ful bugs us­ing their own novel tech­niques, and all while shar­ing tips and tricks in one-off con­ver­sa­tions that could not be repli­cated any­where else.

Prior to the ad­vent of live hack­ing events, most se­cu­rity re­searcher cir­cles were small, iso­lated, and shar­ing in­for­ma­tion pub­licly was prac­ti­cally un­heard of. LHEs cre­ated a way for se­cu­rity re­searchers to con­nect with each other, and es­sen­tially cre­ated a whole new com­mu­nity within in­fosec. Most of my clos­est friends nowa­days are peo­ple who I met through the live hack­ing scene, and I am ex­tremely grate­ful to HackerOne for that.

LHEs were not the only area where HackerOne was build­ing and es­tab­lish a com­mu­nity for se­cu­rity re­searchers. They cre­ated a HackerOne Community space to or­ga­nize mee­tups, on­line events, work­shops, and CTFs. They cre­ated re­gional clubs, and ap­pointed hack­ers who lived there as am­bas­sadors to help fos­ter and grow lo­cal re­searcher com­mu­ni­ties all around the world.

But slowly but surely, things be­gan to change. The com­mu­nity groups and events lost mo­men­tum and fiz­zled out. The cus­tom de­signed silkscreen LHE posters be­came cheap low-ef­fort laser prints. The peo­ple who had ded­i­cated years to cre­at­ing and run­ning in­cred­i­ble events were laid off or left. The LHE in­vi­ta­tion and scor­ing sys­tems be­came (even more) ex­clu­sive, gam­i­fied, and ex­ceed­ingly cal­cu­lated.

So one by one, the domi­noes be­gan to fall.

The Profit Problem

#

Before go­ing fur­ther, I think it’s im­por­tant to get into the why” be­hind these changes that started hap­pen­ing. Sometime around 2020 or 2021, HackerOne was forced to come face-to-face with a very im­por­tant ques­tion that every busi­ness must ask it­self at some point: How do we make money?”

To un­der­stand how HackerOne even was able to sur­vive as a busi­ness, you have to first know that HackerOne was ba­si­cally run­ning en­tirely on VC money for the first 10 years of its ex­is­tence. Between 2014 and 2022, the com­pany did a seed round al­most every 2 years, rais­ing a to­tal of $160M. If a com­pany is self-sus­tain­ing, there is lit­tle-to-no rea­son to con­tinue rais­ing money un­less you have an in­cred­i­bly high burn-rate, which would be odd for a com­pany with such lit­tle in­fra­struc­ture and tech­ni­cal in­no­va­tion.

If you know any­thing about VC, you know how this tends to work; they give you money, and in re­turn, you give them in­di­rect con­trol of the busi­ness through board seats, ad­vi­sory po­si­tions, and other lever­age mech­a­nisms. The VCs gave the money, so the VCs get the power.

Rome did not fall in a day, and nei­ther did HackerOne. The par­a­digm shift started slowly, and then all at once. The core tech­nol­ogy dri­ving the plat­form be­gan to stag­nate. Very lit­tle changed within the plat­form. The UI re­mained tired and functional”. The per­for­mance re­mained lack­ing.

But the VCs looked at their play­book and re­al­ized that the eas­i­est way to make money was sim­ple: get more cus­tomers.

First, the found­ing CEO was re­placed with a cor­po­rate CEO. Instead of a 20% cut of boun­ties paid, they shifted to ca­pac­ity-based fee struc­tures, an­nual con­tracts, and lock­ing cus­tomers into multi-year deals.

And then, sales. SALES SALES SALES!! They were fully con­vinced that sales were the bread and but­ter of the busi­ness. So in­stead of bas­ing the com­pany around hack­ers, they leaned into sales.

More cus­tomers = more an­nual con­tracts = more pre­dictable long-term rev­enue. Once a cus­tomer is locked in, they are fed stats and num­bers and greased up to keep them happy, but also to keep de­mands low. When con­tract re­newal comes up, they are given mas­sive (30 – 60%) dis­counts in ex­change for a multi-year con­tract in or­der to keep them locked in and pay­ing.

Account Managers would have reg­u­lar meet­ings with cus­tomers and en­cour­age them to raise boun­ties to stay com­pet­i­tive and at­trac­tive to hack­ers. They would tell you, Hackers want to spend time fo­cused on the high­est-pay­ing pro­grams”.

HackerOne was proud of this and en­cour­aged the be­hav­ior in­ter­nally. They gave out re­wards, free va­ca­tions, and en­cour­aged their rapidly grow­ing sales team to sign new cus­tomers as much as pos­si­ble.

As this con­tin­ued, things con­tin­ued to de­grade for every­one in­volved.

For triagers, the best ones (who were hack­ers them­selves) burned out and quit. The pay was too low and the work­load was too much.

For hack­ers, the triage ex­pe­ri­ence de­clined, and they be­came flooded with ex­ces­sive amounts of low-qual­ity pro­grams.

For cus­tomers, the qual­ity of re­ports went down, the costs went up, and the race to the bot­tom be­gan.

And for every­one, the plat­form ex­pe­ri­ence de­clined and never evolved.

Typically at this point, mar­ket eco­nom­ics would dic­tate that this is where a com­peti­tor takes your busi­ness and you must im­prove to re­tain mar­ket share. But bug bounty is an oli­gop­oly, and there are 3 com­pa­nies that con­trol al­most the en­tire mar­ket.

So in­stead, HackerOne was forced to cre­ate a new ex­clu­sive op­por­tu­nity called the Hacker Success Program” (HSP). This is a ded­i­cated space for top hack­ers to get di­rect sup­port from HackerOne for any­thing bug bounty re­lated. They were as­signed to a Hacker Success Manager (HSM) who was em­ployed by HackerOne, and they could lever­age that re­la­tion­ship to nav­i­gate mis­com­mu­ni­ca­tions, bad pro­gram ex­pe­ri­ences, low or in­ac­cu­rate bounty pay­outs, and much more. They also use this space to of­fer unique op­por­tu­ni­ties like H1 chal­lenges for hack­ers in the HSP group.

One thing I want to call out is that while this is a great idea in the­ory, in prac­tice it cre­ated a com­pletely lop­sided play­ing field for new hack­ers. You have ba­si­cally zero ways to ad­vo­cate or work through prob­lems (good luck work­ing with HackerOne sup­port), so the rich get richer and you are given the op­tion to deal with it or get lost un­less/​un­til you are a top hacker. Oddly enough, it would be much eas­ier to be­come a top hacker if you had ac­cess to these re­sources in the first place.

Anyway, for some hack­ers, the HSP rev­o­lu­tion­ized what was pre­vi­ously an im­pos­si­ble brick wall. If you got screwed over by a com­pany who failed to un­der­stand the se­cu­rity im­pact of your re­port, you could now lean on your HSM to help get a di­rect line of com­mu­ni­ca­tion to the com­pany and try to re­solve that.

But for oth­ers, the HSP high­lighted one of the core un­der­ly­ing fail­ures that has plagued HackerOne for years: stag­nat­ing fea­ture de­vel­op­ment. Even if you are part of the ex­clu­sive HSP group, you still don’t have the abil­ity to do one im­por­tant thing, which is to in­voke change and de­vel­op­ment within the plat­form.

And yet, for some rea­son, HackerOne de­cided to take a per­for­ma­tive ap­proach to this, go­ing as far as to cre­ate a ded­i­cated plat­form feed­back chan­nel, but do­ing ab­solutely noth­ing with any of that feed­back. I can­not even count how many pieces of in­di­vid­ual feed­back and sug­ges­tions have been posted in there, with pos­i­tively zero ac­tion from HackerOne.

The AI Era

#

And then, around 2021, we all en­tered a new pe­riod of hu­man his­tory: the rise of AI and LLMs.

Beginning around this time, we all be­gan to wit­ness the huge gains in ve­loc­ity and the ca­pa­bil­i­ties of LLMs. You could now one-shot fea­tures in a frac­tion of the time, build cus­tom tools, and over­all get more done in less time.

But this is where HackerOne took a con­fus­ing turn. It is truly baf­fling to me how HackerOne man­aged to fum­ble this tech­nol­ogy in the worst way pos­si­ble. They had a 10-year back­log of fea­ture re­quests, and were handed one of the most pow­er­ful soft­ware de­vel­op­ment tools in the last 50 years. But in­stead of us­ing this to im­prove the plat­form and add long-re­quested fea­tures, they de­cided to move to­ward cre­at­ing their own” AI as­sis­tant called Hai. Not only was Hai just a wrap­per on top of OpenAI, but it barely had any real unique ca­pa­bil­i­ties to help hack­ers or pro­grams do what they had been re­quest­ing from HackerOne to add to the core plat­form.

On top of this, the dri­ving force be­hind the com­pany, the hack­ers and co-founders who started the busi­ness, were qui­etly shoved away into the dun­geon of HackerOne. The web­site has them listed as part of the ex­ec­u­tive team, but they are pup­pets in their own busi­ness, barely get­ting to work on the prod­uct they built from the ground up.

The Enshittification of HackerOne

#

For al­most 10 years, Marten Mickos served as CEO of HackerOne. Some hack­ers loved Marten; oth­ers had mixed feel­ings. Personally, my in­ter­ac­tions with him were some­where be­tween fine and good. But I do be­lieve that he was very pas­sion­ate about bug bounty and did a lot of good for the com­pany dur­ing his time as CEO.

But in late 2024, Marten was re­placed by Kara Sprague, the for­mer Chief Product Officer at F5. You might be ask­ing your­self, what does a CPO of a net­work­ing com­pany know about hack­ing and bug bounty?

Honestly, I am not sure. What I will say is that there are mixed sig­nals on­line re­gard­ing what hap­pened to BIG-IP dur­ing her time as CPO at F5.

But that did­n’t stop the board from de­cid­ing she was the right fit for the fu­ture of the com­pany. And this was nei­ther the first nor the last in a se­ries of ques­tion­able busi­ness de­ci­sions made by HackerOne.

As AI agents con­tin­ued to grow and shift into var­i­ous iden­ti­ties, HackerOne’s iden­tity also be­gan to shift. Slowly but steadily, HackerOne re­branded it­self around a newly in­vented in­dus­try cat­e­gory: CTEM, or Continuous Threat Exposure Management”.

HackerOne quickly be­gan to evolve into a soul­less cor­po­ra­tion ob­sessed with num­bers, busi­ness ob­jec­tives, and B2B sales. They switched from talk­ing about bug bounty pro­grams, live hack­ing events, and how they could help you stay se­cure, to pro­mot­ing their in-house AI se­cu­rity prod­uct and con­tin­u­ous se­cu­rity mon­i­tor­ing tool.

Wait, what? What in-house AI se­cu­rity prod­uct?

Your Reports Are Yours, We Promise

#

Sometimes, it all starts with a tweet.

In February 2026, well-known hacker zseano no­ticed that there was a surge in HackerOne em­ploy­ees leav­ing the com­pany and asked why that would be hap­pen­ing.

This ended up re­veal­ing that HackerOne had made ToS up­dates which made it so that re­ports sub­mit­ted on HackerOne could be used to train AI mod­els.

As you might ex­pect, this cre­ated quite a bit of noise, es­pe­cially within the HSP chats. One thing I’ve no­ticed is that the mod­ern day HackerOne only ever lets the founders out of the dun­geon for dam­age con­trol, and this was a per­fect time to pull that lever.

Within 24 hours, Alex Rice, co-founder, CTO (and CISO…?) of HackerOne emerged from the dun­geon to per­form dam­age con­trol. (This was the first time he had ever mes­saged in the HSP gen­eral chat since it was cre­ated two years ear­lier).

We do not train, fine-tune, or oth­er­wise im­prove GenAI or large lan­guage mod­els on re­searcher data. That in­cludes Agentic PTaaS. https://​docs.hackerone.com/​en/​ar­ti­cles/​10908081-hai-se­cu­rity-trust

We do not train, fine-tune, or oth­er­wise im­prove GenAI or large lan­guage mod­els on re­searcher data. That in­cludes Agentic PTaaS.

https://​docs.hackerone.com/​en/​ar­ti­cles/​10908081-hai-se­cu­rity-trust

And when asked to re­move Section 3.1 of the ToS, Alex promised that this would be taken care of and made clearer in an up­com­ing up­date to the ToS.

Within a week, CEO Kara Sprague had been looped in and made a fluffy LinkedIn post ex­plain­ing that

HackerOne does not train gen­er­a­tive AI mod­els, in­ter­nally or through third-party providers, on re­searcher sub­mis­sions or cus­tomer con­fi­den­tial data. Researcher sub­mis­sions are not used to train, fine-tune, or oth­er­wise im­prove gen­er­a­tive AI mod­els. This ap­plies across our plat­form, in­clud­ing ca­pa­bil­i­ties used within our Agentic PTaaS of­fer­ing and our in-plat­form AI, Hai.

HackerOne does not train gen­er­a­tive AI mod­els, in­ter­nally or through third-party providers, on re­searcher sub­mis­sions or cus­tomer con­fi­den­tial data.

Researcher sub­mis­sions are not used to train, fine-tune, or oth­er­wise im­prove gen­er­a­tive AI mod­els. This ap­plies across our plat­form, in­clud­ing ca­pa­bil­i­ties used within our Agentic PTaaS of­fer­ing and our in-plat­form AI, Hai.

Alex Rice even went on to the Critical Thinking Bug Bounty Podcast and did a guest spot to talk about this is­sue specif­i­cally and shut down any con­cerns about re­port data be­ing used to train AI mod­els.

Amazing, that’s set­tled then. HackerOne does­n’t use re­searcher data to train AI mod­els.

…Right?

Just Kidding They Never Were

#

Unfortunately for us, the pre­sent-day HackerOne fol­lows a rather im­pres­sive PR dam­age con­trol play­book:

Address fears

Ease con­cerns

Don’t change any­thing

About two months af­ter this drama, some­one men­tioned in the HSP chat that they had no­ticed a com­ment from hackerone-agent on one of their re­ports and wanted to know if this was an AI agent or an ac­tual hu­man.

The re­sponse from one of the HackerOne prod­uct man­agers was:

This is a pre­lim­i­nary re­view done by AI […] The H1 triage process has a step called H1 Intake” […] which we’ve au­to­mated for cer­tain re­ports to make sure they reach the Validation team” faster.

This is a pre­lim­i­nary re­view done by AI […] The H1 triage process has a step called H1 Intake” […] which we’ve au­to­mated for cer­tain re­ports to make sure they reach the Validation team” faster.

And when asked whether certain re­ports” meant cer­tain re­port types or whether a pro­gram could en­able it every­where, they re­sponded:

We’re run­ning the analy­sis on all re­ports, the AI de­ter­mines what can be sent di­rectly to the val­i­da­tion team vs what needs to be re­viewed by a hu­man from the in­take team first. […] The sys­tem also learns from be­hav­iour so, if a rec­om­men­da­tion is re­jected […] it will take these into ac­count with its rec­om­men­da­tions as well.

We’re run­ning the analy­sis on all re­ports, the AI de­ter­mines what can be sent di­rectly to the val­i­da­tion team vs what needs to be re­viewed by a hu­man from the in­take team first. […] The sys­tem also learns from be­hav­iour so, if a rec­om­men­da­tion is re­jected […] it will take these into ac­count with its rec­om­men­da­tions as well.

Hang on a sec­ond.

I thought that re­searcher sub­mis­sions are not used to train, fine-tune, or oth­er­wise im­prove gen­er­a­tive AI mod­els.

But now all re­ports are be­ing run through an AI sys­tem, and the sys­tem is us­ing the out­comes of those re­ports to in­flu­ence how it han­dles fu­ture re­ports. So is that tech­ni­cally fine-tuning”? Technically, no.

So how can both of these things be pos­si­ble?

Well, when I asked this, I re­ceived an in­cred­i­ble copy-paste re­sponse of what Alex Rice had said back in February:

Hai does not train, fine-tune, or oth­er­wise im­prove GenAI or large lan­guage mod­els on cus­tomer or re­searcher data: https://​docs.hackerone.com/​en/​ar­ti­cles/​10908081-hai-se­cu­rity-trust

Hai does not train, fine-tune, or oth­er­wise im­prove GenAI or large lan­guage mod­els on cus­tomer or re­searcher data: https://​docs.hackerone.com/​en/​ar­ti­cles/​10908081-hai-se­cu­rity-trust

Docker Sandboxes | Sandboxes for Coding Agents | Docker

www.docker.com

Docker Sandboxes

Run AI agents safely in lo­cal sand­boxes.

Disposable, iso­lated sand­boxes for AI agents like Claude Code, Gemini CLI, Copilot CLI, Codex, OpenCode, and Kiro that need safe, un­at­tended ex­e­cu­tion.

ma­cOS

$ brew trust docker/​tap && brew in­stall docker/​tap/​sbx

Windows

> winget in­stall Docker.sbx

See it in ac­tion

Sandboxes in ac­tion.

Watch an agent in­stall pack­ages, run Docker, mod­ify con­figs, and ex­e­cute un­at­tended. Then dis­pose of the sand­box in one com­mand.

sbx-demo

Click Run Demo” to start

Get started

Get started in sec­onds.

ma­cOS

$ brew trust docker/​tap && brew in­stall docker/​tap/​sbx

Windows

> winget in­stall Docker.sbx

Why sand­boxes

Give agents the au­ton­omy they need to get work done, safely.

Agents do their best work when they have free­dom. Sandboxes let them run fast with­out run­ning wild, so speed and safety stop be­ing a trade­off.

Capabilities

YOLO mode, safely.

Each agent runs in­side a ded­i­cated mi­croVM with your dev en­vi­ron­ment and only your pro­ject work­space mounted in. Agents can in­stall pack­ages, mod­ify con­figs, and spin up their own Docker con­tain­ers. Your host stays un­touched. No man­ual re­view, no per­mis­sion prompts, no su­per­vi­sion re­quired.

Customizable Safe Execution

Network and filesys­tem con­trols you de­fine.

Enforceable org-wide with Docker AI Governance.

MicroVM Isolation

Hard se­cu­rity bound­ary from the host.

Fast to Spin Up, Easy to Tear Down

Disposable by de­fault. Faster than VMs.

Agents Can Use Docker Too

Agents can spin up con­tain­ers within Sandboxes.

Real Dev Environment

Install pack­ages, run ser­vices, work un­at­tended.

One Sandbox for All Your Coding Agents

Claude Code, Gemini CLI, Copilot CLI, Codex, Kiro, OpenCode.

Default –dangerously-skip-permissions Use per­mis­sive modes with con­fi­dence. In fact, that’s the de­fault.

Works with lead­ing cod­ing agents

Every team is about to have their own team of AI agents do­ing real work for them. The ques­tion is whether it can hap­pen safely. NanoClaw was built on the prin­ci­ple that you don’t trust agents with se­cu­rity, you build walls around them. Docker has been ahead of the curve on ex­actly this. Docker Sandboxes is what that looks like at the in­fra­struc­ture level, mak­ing it pos­si­ble for or­ga­ni­za­tions to get the full value from agents with­out com­pro­mis­ing on se­cu­rity.

Gavriel Cohen

Creator of NanoClaw, NanoClaw

Docker Sandboxes let agents have the au­ton­omy to do long-run­ning tasks with­out com­pro­mis­ing safety. We’re ex­cited to in­te­grate Sandboxes into Warp so that de­vel­op­ers can run agents freely with a con­sis­tent en­vi­ron­ment, re­gard­less of whether agents are run­ning lo­cally or in the cloud.

Ben Navetta

Engineering Lead, Warp

Give agents free­dom. Keep what mat­ters safe.

ma­cOS

$ brew trust docker/​tap && brew in­stall docker/​tap/​sbx

Windows

> winget in­stall Docker.sbx

FAQ

Common ques­tions.

What is a sand­box for AI cod­ing agents?

A sand­box is a mi­croVM iso­lated en­vi­ron­ment that pro­tects your filesys­tem and net­work from agents run­ning in­side it.

Which cod­ing agents are sup­ported?

Out of the box we sup­port Claude Code, Gemini CLI, Copilot CLI, Codex, OpenCode, Kiro. You can also cre­ate your own

What does YOLO mode” mean, and is it safe?

YOLO mode (–dangerously-skip-permissions) gives agents au­ton­omy with no ap­proval prompts. Essential for speed, but risky with­out guardrails. Sandboxes make it safe by iso­lat­ing each agent in­side a ded­i­cated mi­croVM.

How is a sand­box dif­fer­ent from a VM?

Sandboxes run fully iso­lated in mi­croVMs, giv­ing more iso­la­tion with­out pay­ing the full cost of run­ning a VM. This lets them do things that need more per­mis­sions safely, like run­ning ad­di­tional Docker con­tain­ers.

Do I need Docker Desktop to use sand­boxes?

No.

What if I need ad­di­tional ad­min con­trols?

Installing Sandboxes cov­ers core func­tion­al­ity. For cen­tral­ized con­trols across a team such as net­work poli­cies, filesys­tem rules, MCP gov­er­nance: Docker AI Governance.

Need More Control Over Your Sandboxes?

With Docker Sandboxes, your de­vel­op­ers get iso­lated en­vi­ron­ments to run agents freely and safely. When your team needs to go fur­ther with net­work ac­cess re­stric­tions, filesys­tem poli­cies, and cen­tral­ized ad­min con­trols, we can help you con­fig­ure the right setup.

Docker AI Governance adds net­work ac­cess poli­cies, filesys­tem con­trols, and org-wide MCP gov­er­nance: de­fined once, en­forced every­where.

Talk to us about:

Network ac­cess poli­cies for sand­box en­vi­ron­ments

Filesystem ac­cess con­trols and re­stric­tions

Admin-level con­fig­u­ra­tion for your team

Talk to an ex­pert

Thank you for your in­ter­est. The Docker Team will be in touch

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

claude.com

We’re mak­ing auto mode the de­fault in Claude Code. Starting on August 14, new ses­sions on Pro, Max, and Team plans will run in auto mode. If you’ve al­ready set a dif­fer­ent de­fault your­self, you may get a one-time prompt ask­ing whether you want to switch to auto mode. If you have a pinned de­fault, noth­ing changes for you. The auto mode clas­si­fier uses a small num­ber of ex­tra to­kens per tool call, and we’re no longer charg­ing Claude Code users on Pro, Max, and Team plans for that clas­si­fier over­head, ef­fec­tive to­day.

Auto mode re­mains opt-in for now on Claude Enterprise, the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud’s Agent Platform, and Microsoft Foundry, giv­ing ad­mins time to re­view the change. In the com­ing month, work­ing with our cloud part­ners, we plan to make it the de­fault across all of these and no longer charge for clas­si­fier over­head. In the mean­time, Enterprise ad­mins can make Claude Code’s auto mode the de­fault through man­aged set­tings.

Auto mode is de­signed to bal­ance users’ de­sire not to be in­ter­rupted with a sys­tem that helps avoid harm­ful ac­tions: in­stead of prompts, it routes each tool call through a clas­si­fier tar­geted at block­ing ac­tions that are ir­re­versible, de­struc­tive, or aimed out­side your en­vi­ron­ment. When the clas­si­fier blocks some­thing, Claude usu­ally finds a safer way to pro­ceed on its own or asks you di­rectly for the go-ahead; if it can’t make progress—three blocks in a row, or twenty across a ses­sion—Claude Code falls back to man­ual ap­provals.

We spent the last sev­eral months test­ing whether auto mode is as safe or safer than an av­er­age user click­ing through prompts. We ran in­ter­nal red-team­ing, third-party red-team­ing and prompt-in­jec­tion eval­u­a­tions, a con­trolled study with 1,053 paid testers, and analy­sis of real pro­duc­tion ses­sions. On every mea­sure we tested, auto mode matched or out­per­formed man­ual re­view.

Auto mode also lets Claude work au­tonomously for longer stretches. This makes mod­els built for long-run­ning work, like Claude Opus 5, more prac­ti­cal to leave run­ning for hours on large tasks. Reducing over­head for users also in­creases out­put. Among Teams & Enterprise adopters, auto mode users ship about 25% more PRs. Unblocking Claude al­lows tasks to run longer un­in­ter­rupted and get more work done. Teams at Adobe, Nuro, Gusto, and Garner Health al­ready run auto mode as their pro­duc­tion de­fault.

Below, we share the safety data and cus­tomer re­sults mo­ti­vat­ing the change, and how to set a dif­fer­ent de­fault if you pre­fer.

Comparing man­ual re­view to auto mode

Data sug­gests that man­ual re­view can be­come ha­bit­ual: users ap­prove 97% of per­mis­sion prompts in Claude Code. While most prompts are likely for safe, rou­tine com­mands, an ap­proval rate that high sug­gests many users are click­ing through re­flex­ively rather than re­view­ing each com­mand. These prompts ask de­vel­op­ers to make dozens or hun­dreds of im­por­tant se­cu­rity de­ci­sions every day, of­ten in the mid­dle of pro­jects, which places the re­view bur­den on users and in­creases the chance that some­thing im­por­tant slips through the cracks. Data also sug­gests that users more fre­quently scru­ti­nize and push back on other types of di­a­logues: for ex­am­ple, when Claude pre­sents a plan for ap­proval, users re­ject 39% of them. But for in­di­vid­ual per­mis­sions re­quests, the re­jec­tion rate is only 3%.

The same pat­tern shows up in set­tings files. As of June 2026, 49.5% of ac­tive CLI users have man­u­ally cre­ated a Bash al­low-rule—5% al­low any shell com­mand out­right, and an­other 43% have in­ter­preter rules like Bash(python:*) or Bash(node:*) that are es­sen­tially equiv­a­lent in prac­tice—and that share is grow­ing roughly 5 per­cent­age points every 5 weeks. Beyond al­low-rules, 62% of users have used by­passPer­mis­sions or clicked don’t ask again” on Bash, and 25% of in­ter­ac­tive ses­sions start in by­pass per­mis­sions mode.

Permission rules still fire be­fore the clas­si­fier in auto mode, ex­cept for al­low rules broad enough to grant ar­bi­trary code ex­e­cu­tion (e.g. python:*). These ar­bi­trary rules are set aside while in auto mode, since they would let com­mands skip the clas­si­fier en­tirely. Settings files aren’t mod­i­fied, and the rules ap­ply again the mo­ment you switch to an­other mode.

Improving safety with auto mode

Research con­ducted in­ter­nally and through sev­eral third par­ties sug­gests that auto mode is safer than how a typ­i­cal user en­gages with per­mis­sion prompts to­day. We tested auto mode against what de­vel­op­ers ac­tu­ally ap­prove when they’re click­ing through per­mis­sion prompts. Across each of these mea­sures—in­ter­nal red-team­ing, third-party pen­e­tra­tion test­ing, and analy­sis of real-world ses­sions—auto mode matched or out­per­formed hu­man ap­proval.

Auto mode out­per­forms man­ual per­mis­sions on safety

We hired 1,053 paid testers through a re­search ven­dor for a cod­ing study. These par­tic­i­pants worked in a test en­vi­ron­ment set up specif­i­cally for the study, not in their own code­bases or on real pro­jects. Participants knew they were be­ing eval­u­ated, but—fol­low­ing stan­dard blind-study de­sign—weren’t told which be­hav­ior was be­ing mea­sured. Partway through each ses­sion, a sin­gle per­mis­sion prompt was swapped for a clearly dan­ger­ous com­mand, and the ven­dor recorded whether the tester ap­proved it (Note: ap­prov­ing the prompt never re­sulted in any­thing dan­ger­ous be­ing run; only the text of the prompt testers saw was changed).

The testers caught the dan­ger­ous com­mand just 13.6% of the time (143 of 1,053), while auto mode blocked 89% of the same com­mands (937 of 1,053). Head to head, auto mode blocked 800 com­mands that a hu­man ap­proved, while hu­mans blocked only 6 that auto mode al­lowed. As ses­sions got longer, hu­mans did worse: they blocked about 17% of dan­ger­ous com­mands early in a ses­sion, drop­ping to about 5% af­ter 50 or more prior prompts, while auto mod­e’s block rate stayed flat re­gard­less of ses­sion length.

Auto mode pre­vents more harm­ful ac­tions

Our re­search also sug­gests that auto mode leads to fewer harm­ful ac­tions than man­ual per­mis­sions mode.

We ex­am­ined how auto mode per­forms in real us­age. From May to June 2026, we ex­am­ined ses­sions our safety pipeline had flagged for re­view, drawn only from ac­counts that have opted in to data use for train­ing and analy­sis. We re-graded each one with a model on two 0 – 10 scales. The first was harm sever­ity with 10 be­ing com­pany-wide, ir­re­versible harm, such as leak­ing a com­pa­ny’s en­tire source code, 7 rep­re­sent­ing pro­duc­tion-level sever­ity such as delet­ing rows from a pro­duc­tion data­base, and 0 a fully re­versible ac­tion con­fined to the lo­cal work­space, like over­writ­ing a scratch file. The sec­ond was how ex­plic­itly the user had asked for the ac­tion, so we could sep­a­rate harm­ful ac­tions the user di­rectly re­quested from harm­ful ac­tions the agent took with less ex­plicit di­rec­tion.

Among the flagged ses­sions, those that went through man­ual ap­proval con­tained se­ri­ous un­in­tended harm more than twice as of­ten as auto mode ses­sions. At pro­duc­tion-level sever­ity (7+), 6.3% of man­u­ally ap­proved ses­sions con­tained a harm­ful ac­tion the user had­n’t ex­plic­itly asked for, com­pared to 2.4% of auto mode ses­sions, and the gap held at every sever­ity level we tested.

Adversarial red-team­ing made auto mode stronger

We worked with Apollo Research, a UK-based AI safety startup, on a two-week pi­lot in­ject­ing syn­thetic at­tacks into real cod­ing tra­jec­to­ries, across three datasets cov­er­ing broad fail­ures, data ex­fil­tra­tion, and safety-re­lated threats.

The en­gage­ment fol­lowed a stan­dard find-fix-ver­ify loop. Apollo shared a set of suc­cess­ful at­tacks along with rec­om­men­da­tions to strengthen auto mode. We hard­ened the clas­si­fier against them pri­mar­ily by giv­ing it more con­text about the en­vi­ron­ment it’s pro­tect­ing (repository vis­i­bil­ity, git state, and the data-han­dling rules de­scribed be­low).

After hard­en­ing, Apollo re-tested on a sep­a­rate held-out at­tack set that we never had ac­cess to. Auto mode caught about half of the at­tacks it pre­vi­ously missed, at a cost of miss­ing about 2% of at­tacks it pre­vi­ously caught. Since we could­n’t tune against the held-out set, this sug­gests the hard­en­ing gen­er­al­izes to new at­tacks rather than patch­ing the spe­cific ones Apollo shared. Across both the shared and held-out sets, the clas­si­fier’s miss rate fell from 12% to 7%.

These at­tacks are syn­thetic and ad­ver­sar­ial by de­sign, built to find where the clas­si­fier fails, and more ma­li­cious than any­thing Claude would do un­prompted. The 7% miss rate is mea­sured against this set and should not be taken as the miss rate on real traf­fic.

Auto mode re­duces prompt in­jec­tion risk

The auto mode clas­si­fier adds an ad­di­tional layer of de­fense against prompt in­jec­tion at­tacks that try to force the agent to de­vi­ate from the user’s in­struc­tions. When run­ning with auto mode, tool re­sults are scanned by probes for po­ten­tially ma­li­cious in­struc­tions, and auto mode checks that ac­tions are aligned with user in­tent. We tested auto mode com­pared to an un­guarded model to il­lus­trate the power of this de­fense-in-depth ap­proach.

We com­mis­sioned an eval­u­a­tion from a third party, Trajectory Labs, who tested dif­fer­ent mod­els within the lat­est pub­licly avail­able ver­sions of Claude Code and Codex as of July 17th 2026.1 They tested 72 in­di­rect prompt in­jec­tion sce­nar­ios held out from Anthropic. Each sce­nario was tested 10 times. Successful at­tacks would, in a non-test­ing en­vi­ron­ment, re­sult in ir­re­versible ac­tions that could harm the user such as sen­si­tive-data ex­fil­tra­tion, fi­nan­cial fraud, or cre­den­tial phish­ing. The at­tacks were op­ti­mized against Claude Opus 4.7, which is not in the re­ported set, so every model in the eval­u­a­tion faced them as un­seen at­tacks. In ad­di­tion to auto mode, the eval­u­a­tions looked at Claude Code’s by­passPer­mis­sions mode and Codex’s Full Access mode.

Both Codex and Claude Code were eval­u­ated us­ing an iden­ti­cal browser in­te­gra­tion pro­duced by Trajectory Labs, PBC. This in­te­gra­tion pro­vides a set of MCP tools which are a thin wrap­per around com­mon Chrome APIs such as `navigate`. Safeguards built into first-party browser in­te­gra­tions such as the Chrome ex­ten­sions pro­vided by OpenAI and Anthropic were not tested. As such, these re­sults should be viewed as a mea­sure­ment of the un­der­ly­ing model, rather than the com­plete set of safe­guards which might ex­ist in a given de­ploy­ment.

In this eval­u­a­tion, none of the 720 at­tack at­tempts suc­ceeded against Claude Fable 5, Opus 5, or Sonnet 5 run­ning auto mode. On the other hand, 5.83% of the at­tacks suc­ceeded against GPT-5.6 Sol run­ning Codex’s Auto-review mode. Notably, this is greater than the 0.09% av­er­age at­tack suc­cess rate against our lat­est mod­els run­ning in by­passPer­mis­sions mode with­out ad­di­tional safe­guards. The tests showed a 19.03% at­tack suc­cess rate against GPT-5.6 Sol when run­ning in Full Access mode. Attack suc­cess rates dif­fer be­tween tasks where tools re­turn re­sults as text and tasks that re­quire in­ter­act­ing with a browser through screen­shots, as shown be­low.

Three in­ci­dents auto mode pre­vented in­side Anthropic

Auto mode is also the de­fault for all Claude Code in­ter­nal us­age at Anthropic. Here are three ac­tions the clas­si­fier pre­vented in­ter­nally:

An off-net­work data leak: When an up­load to an in­ter­nal drive failed, Claude’s fall­back was to pub­lish the same re­port to a pub­lic code-shar­ing site. Auto mode de­nied the fall­back and prompted the user for next steps.

A de­struc­tive mass op­er­a­tion: In a long-run­ning ses­sion, Claude at­tempted a mass process kill across roughly two thou­sand pods and queued dele­tions that would have torn down hun­dreds of GPUs of in-flight train­ing jobs. Auto mode flagged the de­struc­tive com­mands be­fore ex­e­cu­tion.

Over-broad priv­i­lege es­ca­la­tion: Asked to pro­vi­sion read-only billing ac­cess, Claude reached for a cloud role that grants full ad­min­is­tra­tive con­trol at the ac­count root. Auto mode caught the mis­match be­tween what was re­quested and what roles would have been granted.

In each case, Claude ei­ther found a safer path on its own or checked in with the user be­fore pro­ceed­ing.

Making auto mode even safer

We’re con­tin­u­ously in­vest­ing in new auto mode fea­tures that make it safer and eas­ier to ship pro­duc­tion code. Recent ex­am­ples in­clude:

Hard de­nies: Data ex­fil­tra­tion, like send­ing your code or se­crets some­where ex­ter­nal, sits in a cat­e­gory the clas­si­fier is de­signed to never ap­prove. To run an ac­tion like that, you have to switch out of auto mode or run the com­mand your­self. Hard deny rules are cus­tomiz­able via set­tings so you can add more rules that you never want al­lowed even when re­quested by users in your or­ga­ni­za­tion.

Rules for data ac­cess and shar­ing: The clas­si­fier now car­ries ex­plicit rules dis­tin­guish­ing se­crets and po­ten­tially sen­si­tive/​con­fi­den­tial in­for­ma­tion—and where each can be ac­cessed and shared. To make those rules en­force­able, it also checks whether the des­ti­na­tion of a git push or pull re­quest is pub­lic, pri­vate, or trusted be­fore the ac­tion runs. The same push can be rou­tine or an ex­fil­tra­tion de­pend­ing on where it lands: code that be­longs in your team’s pri­vate repos­i­tory should­n’t end up in a pub­lic one, and the clas­si­fier is now de­signed to flag when this might hap­pen.

Checking git sta­tus be­fore de­struc­tive git ac­tions: Before a com­mand that could dis­card un­com­mit­ted work, like git re­set –hard, the clas­si­fier sees the repos­i­to­ry’s cur­rent git sta­tus, let­ting auto mode know what is be­ing re­set.

Prompt in­jec­tion screen­ing: When Claude pulls con­tent from ex­ter­nal sources, like web pages, file con­tents, or tool out­puts, an API-side probe checks that con­tent for at­tempts to hi­jack Claude’s be­hav­ior. When some­thing looks like an in­jec­tion at­tempt, a warn­ing is added to Claude’s con­text be­fore the re­sult is shared with the user.

Auto mode in pro­duc­tion

Teams are al­ready run­ning auto mode as their pro­duc­tion de­fault:

Adobe’s mer­chan­dis­ing plat­form team is re­spon­si­ble for keep­ing pric­ing and pro­mo­tional pages ac­cu­rate and cur­rent across 90+ coun­tries and 30+ lan­guages on Adobe.com. They built an agen­tic loop to build and ver­ify those pages, run­ning it in auto mode so en­gi­neers re­ceive fin­ished PRs for re­view.

Nuro runs auto mode across its re­search and en­gi­neer­ing orgs, us­ing it to power overnight re­search agents that hill-climb eval­u­a­tion met­rics and re­turn fin­ished PRs for re­view by morn­ing.

Gusto adopted auto mode to end the per­mis­sion fa­tigue that was push­ing en­gi­neers to­ward by­pass­ing per­mis­sions checks en­tirely. About 10% of ses­sions since mid-May in­clude a clas­si­fier de­nial—ev­i­dence it’s do­ing real work with­out slow­ing le­git­i­mate tasks.

Garner Health pushed auto mode as the de­fault to all 550 em­ploy­ees via man­aged set­tings, stan­dard­iz­ing a com­pany-wide soft­ware de­vel­op­ment life­cy­cle (SDLC) that no longer de­pends on hand-cu­rated com­mand al­lowlists.

eBook

Learn how these cus­tomers are run­ning auto mode in pro­duc­tion.

Getting started

For Pro, Max, and Team users: if you haven’t set a de­fault per­mis­sion mode, you’ll re­ceive an in-prod­uct no­tice and new ses­sions will start in auto mode au­to­mat­i­cally. If you’ve set a dif­fer­ent de­fault, you may see a one-time prompt ask­ing if you’d like to switch your de­fault to auto mode. If your Team ad­min has set a de­fault in man­aged set­tings, noth­ing changes for you.

For Enterprise users and users who ac­cess Claude Code via the Claude API, auto mode re­mains opt-in for now. We plan to make auto mode the de­fault in the com­ing month, and we’ll no­tify Enterprise ad­mins be­fore we do.

To switch modes, press Shift+Tab in the CLI or use the mode drop­down on the desk­top app. Admins can pin an org-wide de­fault with `defaultMode` in man­aged set­tings, or turn auto mode off en­tirely with `disableAutoMode`.

Finally, while we be­lieve auto mode re­duces risk for most users, it re­lies on clas­si­fi­ca­tion sys­tems and there­fore does not elim­i­nate risk. For high-stakes changes to pro­duc­tion in­fra­struc­ture, we still rec­om­mend re­view­ing Claude’s ac­tions your­self. See the auto mode docs for full con­fig­u­ra­tion in­struc­tions.

This ar­ti­cle was writ­ten by Conner Phillippi, with con­tri­bu­tions by Nicholas Carlini, Isaac Fung, John Hughes, Alex Isken, Shawn Moore, Javier Rando, and Molly Vorwerck. The au­thors would also like to thank Yacine Azmi, Chandler Bair, Kefan Chen, Boris Cherny, Ian Grunert, Lydia Hallie, Alex Kleiman, Lauren Polansky, Deon Poncini, Robert Schonberger, Marie Vachovsky, Qing Wang, Cat Wu, Daniel Xu, and Alice Zhao.

1 We eval­u­ated Claude Code v2.1.205 and Codex v0.144.5. OpenAI re­leased a new ver­sion of Auto-review last week that could change the re­sults.

No items found.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.