10 interesting stories served every morning and every evening.

Pi, Minimal and Performant | EARENDIL

earendil.com

Pi’s Minimalism Is Its Advantage

AI has made code cheap, and as a re­sult many com­pa­nies are build­ing big­ger tools in pur­suit of bet­ter per­for­mance. Larger prompts, more or­ches­tra­tion, more lay­ers, more com­plex­ity. This also makes these tools in­trin­si­cally more ex­pen­sive to use. Pi takes the op­po­site ap­proach.

Pi is the cod­ing har­ness that chooses min­i­mal­ism on pur­pose. It comes out of the box with only 4 tools, and its sys­tem prompt and tool de­f­i­n­i­tions come in be­low 1,000 to­kens. The idea be­ing that most work can be done with the ba­sics, and if you want more, build it.

Evidence in­creas­ingly sug­gests that Pi’s de­sign is not just cleaner; it’s cheaper and more per­for­mant. Users are find­ing that vanilla Pi pro­duces in­dus­try lead­ing re­sults, even be­fore adding on ex­ten­sions to match user spe­cific work­flows and needs. As we’ll see in case stud­ies of Databricks and Shopify, Pi pro­duced ideal out­comes for both.

Case Studies

Databricks Study: Cost Per Task

Databricks re­cently shared their find­ings Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase.” The goal of their re­search was to un­der­stand which cod­ing agents of­fer the best per­for­mance on real-world cod­ing tasks, and how task-per­for­mance varies with price.

To avoid bias from ex­ter­nal bench­marks that have be­come over­sat­u­rated, they cre­ated their own based on tasks their team of en­gi­neers reg­u­larly per­forms. The re­sults match what we would ex­pect, but what many in the in­dus­try may have been sur­prised to learn. In their words, …the har­ness a model is called from dra­mat­i­cally im­pacts cost and qual­ity,” and, in many cases, sim­ple har­nesses like Pi per­formed best on our work­loads.”

When com­bined with Opus 4.8, xhigh, Pi had the high­est over­all pass-rate, at a sig­nif­i­cantly lower cost than both Claude Code and Codex.

Minimal har­ness, mea­sur­able ef­fect

Pi shines be­cause it does­n’t try to wrap the model in a bunch of de­faults and in­struc­tions that get lost in the in­struc­tion hi­er­ar­chy. Instead, Pi stays out of the mod­el’s way, and the team is able to add what they ac­tu­ally need for their work­flow.

Databricks’ study is in­sight­ful be­cause it sep­a­rates model from har­ness.

They re­ported that when they ran the same model with the same think­ing ef­fort through dif­fer­ent har­nesses, the cost per task dif­fered sig­nif­i­cantly (more than 2x in some cases), while qual­ity re­mained the same”. We call this Pi’s context dis­ci­pline”. Pi sent about 3x less con­text per turn. It man­aged con­text bet­ter, keep­ing a tighter work­ing set and fin­ish­ing the tasks in fewer runs.”

We agree that one must take into ac­count end-to-end en­gi­neer­ing eco­nom­ics, and not just price per to­ken. And this is also true at the model level; we have ob­served, for in­stance, that run­ning com­plex work­flows on Haiku 4.5 was of­ten more ex­pen­sive than Sonnet 4.6, es­pe­cially when code ex­e­cu­tion was in­volved, sim­ply be­cause the agent re­quired more turns to com­plete the task suc­cess­fully.

Now we see this at the har­ness level too; stronger, more ex­pen­sive mod­els with a per­for­mant har­ness can be cheaper than the con­verse.

Shopify builds Pi Autoresearch: Extensible beats bloat

Minimalism is part of Pi’s core phi­los­o­phy. What makes this work is that min­i­mal does not mean in­flex­i­ble. In fact, it is the first widely used agen­tic in­fra­struc­ture cre­ated for ex­ten­si­bil­ity and self-ed­itabil­ity.

Another in­sight­ful ex­ter­nal val­i­da­tion of Pi’s de­sign comes from Shopify. In this post from Shopify Engineering, David Cortés de­scribes build­ing pi-au­tore­search di­rectly as a Pi ex­ten­sion, by sim­ply ask­ing Pi, [to] cre­ate an ex­ten­sion for Autoresearch…”. Pi reads its own ex­ten­sion doc­u­men­ta­tion and starts build­ing a new work­flow from there.

Autoresearch is an au­tonomous loop for op­ti­miza­tion with cod­ing agents. When you ask for a change, it runs ex­per­i­ments to find out what works and what causes re­gres­sions. For as long as the tar­get is mea­sur­able, it can throw out these re­gres­sions and keep self-im­prov­ing.

For Shopify and oth­ers, the Autoresearch ex­ten­sion quickly be­came a se­ri­ous in­ter­nal pro­duc­tiv­ity tool. Shopify re­ported cases in­clud­ing unit tests run­ning 300 times faster,” React com­po­nent mount­ing 20% faster,” re­duced build times across mul­ti­ple pro­jects, and even im­prove­ments to pnpm per­for­mance.

The im­por­tant point here is that Pi does­n’t ship any of these tools out of the box. Instead, it makes it ridicu­lously sim­ple for you to build them. Instead of as­sum­ing the ven­dor knows your work­flow and try­ing to ship every tool un­der the sun, Pi as­sumes you know best, and gifts you ex­ten­si­bil­ity to wield and craft your own work­flow.

Why min­i­mal wins now

About a year ago, an ar­gu­ment could be made for na­tive har­nesses hav­ing a struc­tural ad­van­tage over all oth­ers, be­cause mod­els were built around them. However, this ar­gu­ment has got­ten weaker.

Frontier mod­els are now gen­er­ally very com­pe­tent at un­der­stand­ing a ter­mi­nal (or ter­mi­nal-style) cod­ing en­vi­ron­ment, and act­ing within it. Anthropic re­cently cut­ting down Claude Code’s sys­tem prompt by 80% is a clear sign of this. So the ques­tion is be­com­ing less about how na­tive the har­ness is, and more about how it han­dles con­text to avoid re­dun­dancy and act with clean prim­i­tives. Models need a clean in­ter­face to the en­vi­ron­ment, and a har­ness that does not waste con­text.

Pi pro­vides this: less prompt over­head and re­peated con­text, cheaper runs, fewer un­nec­es­sary ab­strac­tions. Because it is ex­ten­si­ble, you do not lose power, but gain se­lec­tiv­ity. You add com­plex­ity only when it earns its keep”.

We are also see­ing lo­cal mod­els de­vel­op­ing fast, and at Earendil we find them very promis­ing. Pi’s con­text dis­ci­pline is es­pe­cially an as­set here. Local mod­els usu­ally have lower con­text win­dows, and pre­fill can take a long time, so pre­serv­ing a sta­ble prompt pre­fix mat­ters. Context dis­ci­pline means we do not change the con­text with­out the user ex­plic­itly ask­ing for it, avoid­ing minute-long re-pre­fill­ing. Combined with the min­i­mal de­fault sys­tem prompt and tool set, this makes pi an ideal har­ness for lo­cal mod­els.

Pi is prov­ing that it can man­age it all. To be cheaper, min­i­mal, and more per­for­mant.

A Civilian Plane Crashed in New Mexico. Was the Military’s Tech to Blame?

www.wired.com

Jul 30, 2026 6:00 AM

Drone war­fare is mak­ing the skies more dan­ger­ous, even for air­planes far from the bat­tle­field.

ANIMATION: Gabriel Gabriel Garble

This past May, a twin-en­gine Beechcraft King Air mede­vac plane took off from Roswell, New Mexico, and headed west to the town of Ruidoso to pick up a pa­tient. It should­n’t have been a chal­leng­ing flight for the two pi­lots and two nurses aboard. The tem­per­a­ture was 69 de­grees; the sky was clear. The 60-mile jour­ney nor­mally takes a half hour, at most.

But once air­borne, the plane ran into trou­ble. At the White Sands Missile Range that night, US mil­i­tary per­son­nel were con­duct­ing a GPS jam­ming ex­er­cise that left the King Air pi­lots—and any­one else within hun­dreds of miles—un­able to use mod­ern nav­i­ga­tion sys­tems. Forced to re­vert to older tech­nol­ogy, ones that they rarely if ever use, the mede­vac pi­lots got dis­ori­ented and crashed into the side of a moun­tain. There were no sur­vivors.

The ac­ci­dent marked the first time that GPS jam­ming had con­tributed to the crash of a civil­ian plane in the United States. But it was just one of a string of re­cent dis­rup­tions across the world. The skies are more con­tested than ever, whether it’s civil­ian drones wan­der­ing out of the ap­proved zone or US agen­cies get­ting their sig­nals crossed, as hap­pened ear­lier this year when New Mexico and Texas scared the pub­lic by tem­porar­ily clos­ing their air­space. (It turned out that US Customs and Border Patrol were us­ing anti-drone lasers in that area.) The GPS jam­ming ex­er­cise that led to this lat­est crash is not a sin­gu­lar event. In the past year, the US mil­i­tary ap­peared to have sent out no­tices for at least 10 such ex­er­cises. As drone war­fare and elec­tronic war­fare ex­pand, air­lines are in­creas­ingly en­coun­ter­ing nav­i­ga­tion dis­rup­tions hun­dreds of miles be­yond the ac­tual con­flict zone,” says Eliran Almog, CEO of the cy­ber­se­cu­rity firm Cyviation.

It’s worth tak­ing a closer look at what hap­pened last May. While the Ruidoso crash was the first fa­tal ac­ci­dent in the US known to be linked to elec­tronic war­fare, there’s no rea­son to think it will be the last.

Even ab­sent GPS jam­ming, mede­vac is one of the most dan­ger­ous cat­e­gories of civil avi­a­tion. (Kreindler, a law firm spe­cial­iz­ing in air crash lit­i­ga­tion, says that mede­vac flights have an ac­ci­dent rate more sim­i­lar to com­bat fly­ing than to civil avi­a­tion.) Flights are of­ten or­ga­nized on short no­tice, and they fly into airstrips that might be un­fa­mil­iar to the flight crew, and be­cause hu­man lives are at stake, there is an in­cen­tive to fly when weather con­di­tions are mar­ginal.

Some of those fac­tors were at play on the night of May 13. At 11 pm, the crew was no­ti­fied that they had to fly to Ruidoso to pick up a pa­tient and bring them to Albuquerque. The pi­lots were cap­tain Keelan Clark, aged 30, and first of­fi­cer Ali Kawsara, aged 23. Clark had got­ten his com­mer­cial pi­lot’s li­cense just a year and a half be­fore; he’d been pro­moted from first of­fi­cer to cap­tain the pre­vi­ous month. Kawsara had just two months on the job. He’d only worked cargo jobs be­fore this one.

Both men had demon­strated pro­fi­ciency fly­ing in low-vis­i­bil­ity con­di­tions us­ing what’s called in­stru­ment flight rules, or IFR. There are two ba­sic ways to nav­i­gate in bad weather. Modern cock­pits are equipped with GPS-enabled equip­ment that shows where the plane is on a com­puter screen and por­trays a ma­genta-col­ored line that shows pi­lots where they need to go. This is called RNAV fly­ing; an RNAV ap­proach” brings planes all the way to the thresh­old of a run­way for land­ing in low-vis­i­bil­ity con­di­tions.

This method of fly­ing is much eas­ier than the pre­vi­ous it­er­a­tion. Before GPS be­came widely avail­able in the 2000s, air­lin­ers nav­i­gated us­ing a com­bi­na­tion of mag­netic com­passes, ground-based ra­dio bea­cons, and in­er­tial sys­tems de­rived from old-fash­ioned gy­ro­scopes. Flying by ra­dio bea­cons re­quires pi­lots to form a 3D men­tal map of their lo­ca­tion rel­a­tive to the bea­cons. They have to prac­tice un­til they be­come so ef­fi­cient that they can stay calm un­der pres­sure, lest they panic, lose their sit­u­a­tional aware­ness, and spi­ral out of con­trol. Just ask Kennedy,” says Kenneth Krentsa, a re­tired air­line pi­lot, re­fer­ring to JFK Jr.’s 1999 night­time crash.

Clark and Kawsara took off at eight min­utes to mid­night and ini­tially headed due west, to­ward Ruidoso. The weather was clear, but be­cause the night was nearly moon­less and the area is rural, the only vi­sual ref­er­ences avail­able were the lights of scat­tered set­tle­ments. It’s a black hole out there,” says Juan Browne, an air­line pi­lot who hosts a crash-in­ves­ti­ga­tion pod­cast.

Unable to ori­ent them­selves with­out vi­sual cues, the pi­lots called up Albuquerque Air Route Traffic Control Center—Albuquerque Center, for short—and re­quested per­mis­sion to fly in­stru­ments-only to Ruidoso. The re­quest was ap­proved.

Under nor­mal cir­cum­stances, the flight that fol­lowed would have been un­event­ful. The pi­lots would have fol­lowed the ma­genta line, and the GPS nav­i­ga­tion equip­ment would have lined them up for a smooth land­ing.

But 100 miles to the west, an Air Force Unit called the 746th Test Squadron of the 704th Test Group was hold­ing its an­nual NAVFEST event at the White Sands Missile Range. The event draws to­gether elec­tronic war­fare units from across the armed ser­vices for two weeks of ex­er­cises, in which units test dif­fer­ent tech­nolo­gies for dis­rupt­ing GPS and deal­ing with ad­ver­saries’ dis­rup­tion.

Courtesy of AirNavRadar.com

The event is held at White Sands be­cause it’s among the most re­mote and sparsely set­tled ar­eas of the con­ti­nen­tal United States. (Not co­in­ci­den­tally, the first atomic bomb was det­o­nated there.) But in the run-up to NAVFEST, the Federal Aviation Administration warned air­craft op­er­a­tors that GPS could be af­fected up to 400 miles away be­tween May 12 and May 18.

At mid­night on May 14, eight min­utes af­ter Clark and Kawsara took off from Roswell, they told Albuquerque Center that they’d lost their GPS. Unable to nav­i­gate on their own, they asked that the con­troller give them a head­ing—a mag­netic di­rec­tion to fly in. The con­troller gave them a head­ing to fly west, and then, a minute later, to turn north.

The pi­lots said that they wanted to fly an RNAV ap­proach to Ruidoso. This be­ing ruled out while GPS is jammed, the King Air pi­lots changed their re­quest and asked to use an al­ter­nate form of land­ing sys­tem called Instrument Landing System, or ILS, that does­n’t re­quire GPS re­cep­tion. At 12:05 am, the con­troller as­sured the King Air that they would pro­vide them with vec­tors to guide them in a cou­ple of min­utes.” In the mean­time, they kept fly­ing north.

In ret­ro­spect, tragedy might have been avoided if Albuquerque air traf­fic con­trol had been able to pay closer at­ten­tion to the young pi­lots in the King Air. But tonight they were busy. Three other air­craft also re­ported that they’d lost their GPS and needed help. One was strug­gling to get a bear­ing on a ra­dio bea­con.

While ATC helped other planes, the King Air con­tin­ued north. By 12:08 am, they had over­shot the land­ing pat­tern by 10 miles.

At this point the pi­lots had three op­tions. They could stick to the cur­rent plan and wait for the busy con­troller to give them the next vec­tor to­ward the land­ing. Or, now that GPS was work­ing again, they could ask to switch to the RNAV ap­proach and fly it them­selves. Or they could ditch the in­stru­ment ap­proach al­to­gether and fly what’s called a vi­sual ap­proach. You see a run­way, and you fly to it.

As the King Air flew north, they were high enough to see the lights of Ruidoso’s air­port 31 miles to the south­west. To the pi­lots in the cock­pit of the King Air, a vi­sual ap­proach must have seemed a tan­ta­liz­ing prospect. Why hang around wait­ing for ATC to give them vec­tors, why go through the men­tal ac­ro­bat­ics of try­ing to fig­ure out where they were rel­a­tive to the ILS bea­con? All they had to do was fly to­ward the lights that they could clearly see through their wind­shield.

The King Air called Albuquerque Center and asked to go vi­sual.” The re­quest was granted.

As they turned and de­scended to­ward Ruidoso’s lights, what the pi­lots could­n’t see was the 10,000-foot-high mass of the Capitan Mountains ly­ing across their path. As they drew closer, the dark mass of rock ap­peared to rise up, swip­ing away the lights of the val­ley. This could have cre­ated confusion and a loss of sit­u­a­tional aware­ness,” Browne says. When those lights go out, man, you know you are in big trou­ble.”

The pi­lots slowed their de­scent, even climb­ing a lit­tle, but it was­n’t enough. They kept fly­ing straight to­ward the moun­tain. When you are in that state of mind, you climb as high as you can,” Krentsa says. And you cir­cle. You stay in one place un­til you fig­ure out where you are. You don’t just keep press­ing for­ward, blindly.”

But that’s what the King Air pi­lots did. They flew straight into the ris­ing slope and hit it at full speed. All four oc­cu­pants died in­stantly.

The na­ture of war­fare is chang­ing pro­foundly, and quickly, as drones be­come cheaper, more nu­mer­ous, and more deadly. Hard to spot, and hard to shoot down, they pro­vide an ef­fec­tive way for smaller, less re­sourced na­tions to level the play­ing field against more pow­er­ful ad­ver­saries. Ukraine, nearly over­whelmed by Russia’s con­ven­tional war­fare might at the be­gin­ning of 2022, has rapidly de­vel­oped its drone force to gain what ap­pears to be an up­per hand in the con­flict. And while the US achieved to­tal air su­pe­ri­or­ity over Iran af­ter at­tack­ing the coun­try this February, it has found it­self help­less to stop Iran from us­ing drones and mis­siles to ef­fec­tively shut down traf­fic through the Strait of Hormuz.

To fight back, de­fend­ers can try to ex­ploit a drone’s nav­i­ga­tion. A cheap and sim­ple way for drones to reach their tar­gets is by GPS, which uses ra­dio sig­nals re­ceived from a con­stel­la­tion of satel­lites to cal­cu­late a po­si­tion. When those sig­nals are blocked, an en­e­my’s drones can be ren­dered blind. But the en­emy, too, can take coun­ter­mea­sures. Drone and anti-drone tech­nolo­gies find them­selves in an end­less cat-and-mouse bat­tle, each con­tin­u­ously try­ing to outdo the other. Exercises like NAVFEST of­fer a way for the US mil­i­tary to stay on top of the game.

Civilian GPS has be­come col­lat­eral dam­age, and air travel most of all. Since 2023, planes fly­ing over large swaths of the Middle East, the Black Sea, and the Baltic Sea re­gions have en­dured waves of GPS jam­ming. Airlines have learned to adapt, but a price is still be­ing paid. GPS was a ma­jor boost for air­line safety, and while re­mov­ing it may not in­stantly cause planes to fall from the sky, it re­moves a layer of pro­tec­tion from pas­sen­gers and crew. Add in other stres­sors—a dark night, an in­ex­pe­ri­enced crew, a lack of pro­fi­ciency in the backup tech­nol­ogy—and the sum to­tal is enough to yield dis­as­ter.

There is no ques­tion that in­creased lev­els of GPS jam­ming and spoof­ing around the world pose a safety risk for com­mer­cial avi­a­tion. When alarms go off rou­tinely in the cock­pit, and when pi­lots learn to dis­re­gard key read­ings from their in­stru­ments be­cause the read­ings can’t be trusted, we’re a long way from nor­mal op­er­a­tion,” says Todd Humphreys, a pro­fes­sor of aero­space en­gi­neer­ing at the University of Texas at Austin who has been a lead­ing re­searcher into GPS dis­rup­tion. Air travel is still very safe, but it may be stuck for the next five years or more in a mild-and-in­creas­ing risk sit­u­a­tion as we con­front ever more GPS in­ter­fer­ence within the very-slow-to-adapt avi­a­tion in­dus­try.”

For a few years, US avi­a­tion was spared the dis­rup­tions of anti-drone elec­tronic war­fare. Then it started to hap­pen here, too. In March 2025, air­lin­ers fly­ing into Ronald Reagan National Airport in Washington, DC, re­ceived spu­ri­ous alarms from a col­li­sion-avoid­ance sys­tem, and sev­eral had to abort their land­ings. It later turned out that the Secret Service was test­ing elec­tronic war­fare equip­ment at the vice pres­i­den­t’s res­i­dence. This year, two sep­a­rate in­ci­dents in West Texas in­volv­ing US Army and CBP drone op­er­a­tions led to air­space clo­sures and the dis­rup­tion of com­mer­cial flights.

The avi­a­tion in­dus­try has been slow to grap­ple with the pro­lif­er­a­tion of counter-drone mea­sures and their po­ten­tial ef­fects on flight safety. Airlines and other com­mer­cial op­er­a­tors are still heav­ily re­liant on GPS for nav­i­ga­tion, and other cru­cial tech­nolo­gies, like col­li­sion avoid­ance, are also vul­ner­a­ble. The harder prob­lem with drones is­n’t de­feat­ing them. It’s do­ing it with­out cre­at­ing a sys­tem that neg­a­tively im­pacts civil avi­a­tion,” says Kris Brost, gen­eral man­ager of Robin Radar Systems, a drone de­fense com­pany. Counter-drone tech­nol­ogy has to be sur­gi­cal, not a sledge­ham­mer, be­cause the air­space we’re try­ing to pro­tect is the same air­space the econ­omy runs on.”

Living in a world with drones of both the friendly and un­friendly va­ri­ety is go­ing to take a lot of ad­just­ing. Historically, ma­jor changes in avi­a­tion take place only af­ter crashes that kill a large num­ber of peo­ple. But a suf­fi­ciently mo­ti­vat­ing cat­a­stro­phe may not be far off.

On July 7, a 737 freighter, op­er­ated by a tiny Pakistani cargo air­line called K2 Airways, took off in the late af­ter­noon from Sharjah in the United Arab Emirates and flew east to­ward Karachi with a five-per­son crew. Its route took it just south of the Strait of Hormuz, an area that had been ex­pe­ri­enc­ing in­tense GPS jam­ming due to the US-Iran con­flict. Later, af­ter night­fall, the flight crew called Karachi air traf­fic con­trol and re­ported a navigational sys­tem is­sue,” ac­cord­ing to the Pakistan Civil Aviation Authority. In the three min­utes that fol­lowed, the plane dove 5,000 feet, climbed 6,000 feet, and then plunged 36,000 feet into the ocean in a near-ver­ti­cal dive, killing every­one aboard. It’s too early to know what caused the crash. But it won’t be any sur­prise if the elec­tronic war­fare made an­other pi­lot fly into dark­ness.

Let us know what you think about this ar­ti­cle. Submit a let­ter to the ed­i­tor at [email protected].

Jeff Wise is a New York–based sci­ence jour­nal­ist spe­cial­iz­ing in avi­a­tion, tech­nol­ogy, and psy­chol­ogy who con­tributes fre­quently to New York magazine. A co­host of the pod­cast Find­ing MH370, he was widely seen in the Netflix doc­u­men­tary se­ries MH370: The Plane That Disappeared and ex­ec­u­tive-pro­duced the Showtime doc­u­men­tary fea­ture Gringo: The Dangerous Life of John McAfee. He is the … Read More

Just a moment...

www.axios.com

Cloudflare OS: an open platform for agents, apps, and work

blog.cloudflare.com

Every or­ga­ni­za­tion has a mis­sion, a rea­son for be­ing. Organizations pass that mis­sion — along with their ter­mi­nol­ogy, pro­ce­dures, sys­tems, stan­dards, and ways of work­ing — to their peo­ple. People, in turn, take this con­text to­gether with their own ex­pe­ri­ence and work to­wards the mis­sion.

Work can take many forms, from code, to doc­u­ments and slides, to re­la­tion­ships, to out­comes in the phys­i­cal world.

Some of these are straight­for­ward: code ei­ther runs or it does­n’t. Agents have been us­ing this feed­back loop to pro­duce code that works” for de­vel­op­ers over the last cou­ple of years. But what about the rest of us?

Bringing the same lever­age to the rest of the or­ga­ni­za­tion is a harder prob­lem. Agents need to un­der­stand the con­text of the com­pany and be able to reach the sys­tems peo­ple use to do their jobs. They need to turn that con­text and ac­cess into work that moves the or­ga­ni­za­tion to­wards its mis­sion.

That’s why we cre­ated Cloudflare OS. It gives every per­son an agent and work­space built around their com­pany: how it works, what it knows, and the sys­tems it re­lies on.

In May of this year, we gave every per­son at Cloudflare ac­cess to the first ver­sion of Cloudflare OS. Thousands of peo­ple across every func­tion, many of them out­side of en­gi­neer­ing, use it every day to cre­ate doc­u­ments and slides, au­to­mate re­peat­able tasks, and build small apps to vi­su­al­ize data and help them do their work.

Cloudflare OS also gave every­one a shared li­brary of con­text and skills built by teams at Cloudflare. It cap­tures our ter­mi­nol­ogy, pro­ce­dures, and best-known ways of do­ing re­cur­ring work as in­struc­tions an agent can fol­low. When one per­son fig­ures out a bet­ter way to do some­thing, every­one else can use it.

Today, we are open sourc­ing a new ver­sion of Cloudflare OS. Any or­ga­ni­za­tion can de­ploy it, con­nect it to in­ter­nal sys­tems, and make it their own.

What we learned from the first ver­sion

The Cloudflare OS we are open sourc­ing to­day is based on what we learned from run­ning the first ver­sion in­ter­nally, a jour­ney our CIO, Sam Rhea, cov­ers in his blog post.

The first ver­sion cen­tered on in­di­vid­u­als work­ing with agents through pri­vate work­spaces. Apps were sta­tic rather than live soft­ware con­nected to in­ter­nal sys­tems, and mostly de­ter­min­is­tic jobs still re­quired run­ning an agent skill again and con­sum­ing more model to­kens.

Collaboration ex­posed a more fun­da­men­tal chal­lenge. Access to an MCP server told us which tools an agent could call, but not which un­der­ly­ing re­sources the agent had ob­served. Once peo­ple be­gan shar­ing work­spaces, apps, and out­puts, we needed to en­sure that col­lab­o­ra­tion could not ex­pose in­for­ma­tion some­one was not per­mit­ted to see.

We re­built Cloudflare OS on a new foun­da­tion to solve these prob­lems. Security had to be part of the plat­form, not some­thing every per­son build­ing an app or us­ing an agent has to im­ple­ment cor­rectly.

The re­sult is a plat­form de­signed to be­long to the com­pany run­ning it. You can cus­tomize the in­ter­faces, con­nect your tools, and add the skills and con­text that cap­ture how your or­ga­ni­za­tion works.

Introducing Cloudflare OS

Cloudflare OS starts with a con­ver­sa­tion in your browser, like many other AI tools. What makes it dif­fer­ent is that each con­ver­sa­tion is grounded in the con­text and skills your or­ga­ni­za­tion has cu­rated. Give your work­space a goal, and it can draw on that knowl­edge and work with the tools and data your or­ga­ni­za­tion al­ready uses to achieve it.

Cloudflare OS com­bines three parts:

An agent work­space grounded in con­text and skills your com­pany cu­rates, with an iso­lated run­time where agents can write and run code.

A new se­cu­rity and gov­er­nance frame­work for safe ac­cess to in­ter­nal data and ser­vices.

A plat­form for per­sonal, mod­i­fi­able apps that peo­ple can build, share, and con­tinue chang­ing.

What be­gins as a con­ver­sa­tion can be­come a doc, an app, or a work­flow that con­tin­ues do­ing the work.

An agent work­space for every­one in your com­pany

Agent work­spaces were de­signed for every­one in your or­ga­ni­za­tion to use. You in­ter­act with them in your browser, so you don’t have to be a de­vel­oper or know how to use a ter­mi­nal.

A work­space com­bines agent ses­sions, per­sis­tent state, out­puts and files, re­source ac­cess, and an iso­lated run­time where the agent can write and run code.

They come loaded with the cu­rated con­text and skills your team or com­pany has col­lected. No more rein­vent­ing the wheel for every task — if some­one on your team has fig­ured out the best way to do some­thing, every­one ben­e­fits. People no longer have to ex­plain the same process, ter­mi­nol­ogy, and best prac­tices to a model every time they start a task.

A few things you can do:

Research and ask ques­tions

Ask a work­space to re­search a topic us­ing com­pany con­text and the re­sources you make avail­able to it. The agent can write code to search, fil­ter, join, and an­a­lyze in­for­ma­tion in­stead of pulling an en­tire dataset into the mod­el’s con­text win­dow.

Create docs, slides, and spread­sheets

A work­space can turn its re­search into a doc­u­ment, pre­sen­ta­tion, or spread­sheet that you can con­tinue edit­ing. These out­puts do not have to be sta­tic files. They can re­main con­nected to live data, be up­dated as their sources change, and still be ex­ported to fa­mil­iar for­mats or ser­vices such as Google Drive.

Create col­lab­o­ra­tive, con­nected apps for your team

When a doc­u­ment or spread­sheet is not enough, the agent can build an app with its own in­ter­face, logic, and state. The app can use con­nected com­pany re­sources and sup­port mul­ti­ple peo­ple work­ing to­gether.

Run de­ter­min­is­tic work­flows

Not every job needs a full agent ses­sion. Many are a known se­quence of steps with one or two places where judg­ment is use­ful. A work­space can turn those jobs into mostly de­ter­min­is­tic work­flows, us­ing code for the pre­dictable steps and a model only where it adds value. Workflows can run on de­mand, on a sched­ule, or when an event oc­curs in a con­nected sys­tem.

Cloudflare OS gives agents and apps gov­erned ac­cess to sys­tems of record through Gatekeepers (more on this in the se­cu­rity sec­tion be­low). It also sup­ports ex­ist­ing Model Context Protocol (MCP) servers your or­ga­ni­za­tion al­ready uses via MCP Server Portals.

A new se­cu­rity and gov­er­nance frame­work for safe ac­cess to in­ter­nal data and ser­vices

As peo­ple be­gin ex­per­i­ment­ing with AI at work, one of their first re­quests is of­ten for API keys to com­pany sys­tems. This makes sense: AI is­n’t much use at work if it does­n’t have ac­cess to the sys­tems peo­ple use to do their jobs.

But hand­ing over API keys to peo­ple and agents is dan­ger­ous and does not scale. Keys of­ten pro­vide broad, long-lived ac­cess that is dif­fi­cult to con­strain, share safely, and au­dit.

MCP gives agents a bet­ter way to use these sys­tems. An MCP server can hold the cre­den­tial and ex­pose a de­fined set of tools in­stead of hand­ing the key di­rectly to the agent. But con­trol­ling which tools an agent can call is only the first step. MCP alone does not tell us which un­der­ly­ing re­sources an agent has ob­served. The agent can com­bine in­for­ma­tion across sys­tems, send it some­where less re­stricted, or ex­pose it through apps and out­puts to peo­ple who may not be al­lowed to see the orig­i­nal re­sources. Authorization has to ac­count for where the data can go next.

Agents start with no ac­cess

Cloudflare Access con­trols who can en­ter Cloudflare OS. Inside, every agent and app starts with ac­cess to noth­ing. An agent can ask for ac­cess to a spe­cific re­source, which you can grant or deny. Generated code re­ceives that re­source as a typed bind­ing:

const is­sues = await env.PRO­JECT.lis­tIs­sues({ teamId: ENG, state: open”, });

env.PRO­JECT is a ca­pa­bil­ity rep­re­sent­ing per­mis­sion to use a spe­cific re­source un­der a spe­cific pol­icy. The cre­den­tial re­mains com­pletely iso­lated from the agent and any gen­er­ated code.

Server code runs in a Dynamic Worker with global out­bound net­work­ing dis­abled. Client code runs in a sand­boxed frame in the browser. Neither can reach the Internet ex­cept through ca­pa­bil­i­ties you ex­plic­itly pro­vide.

Gatekeepers gov­ern re­sources and ac­tions

A Gatekeeper is a ser­vice-spe­cific Worker that sits be­tween Cloudflare OS and an ex­ter­nal ser­vice. It un­der­stands the ser­vice’s API, its re­sources, and the op­er­a­tions that can be per­formed on them.

Giving an agent ac­cess to your en­tire GitHub ac­count is likely too broad. A Gatekeeper can give it ac­cess to a sin­gle repos­i­tory, al­low it to read is­sues but not source code, mask par­tic­u­lar fields, ap­ply rate lim­its, and re­quire ap­proval be­fore merg­ing a pull re­quest.

The agent and its apps see a small TypeScript API. The Gatekeeper han­dles OAuth, holds the cre­den­tial, en­forces pol­icy, records what was read, and me­di­ates any­thing with an ex­ter­nally vis­i­ble side ef­fect.

Policy fol­lows what the agent has seen

Controlling the ini­tial read is not enough. Take, for ex­am­ple, the case where an agent reads a sen­si­tive table in a data ware­house and uses it to pro­duce a live dash­board. Sharing the dash­board must not be­come a way to share the table with peo­ple who could not ac­cess it di­rectly.

Cloudflare OS records every re­source agents ob­serve. These ob­ser­va­tions re­main at­tached to the agent and its work. When an­other per­son tries to open the work­space, in­ter­act with the agent, or view what it pro­duced, Gatekeepers ver­ify that per­son’s ac­cess to the ob­served re­sources.

The same ob­ser­va­tion log is used to in­form poli­cies that de­ter­mine when agents can make ex­ter­nal re­quests. A read of sen­si­tive data can pre­vent the agent from writ­ing data to cer­tain sources, invit­ing new col­lab­o­ra­tors, hand­ing work to an­other agent, or mak­ing an out­bound re­quest.

People us­ing agents or build­ing apps do not have to worry about mak­ing these mis­takes. The plat­form can now be used to han­dle this.

A plat­form for build­ing and shar­ing per­sonal, mod­i­fi­able apps

Most pro­duc­tiv­ity suites give you a fixed set of ap­pli­ca­tions: doc­u­ments, spread­sheets, and pre­sen­ta­tions. In Cloudflare OS, each file” can be its own ap­pli­ca­tion, writ­ten by an agent for one per­son, one pro­ject, or one team.

These are not pro­to­types that you have to ex­port and de­ploy some­where else. Each one is a full-stack ap­pli­ca­tion with client code, server code, an API, and durable state. Apps are pri­vate by de­fault, but can be shared like doc­u­ments.

Every app is a Worker

When you ask your work­space to build an app, the agent writes two parts:

Client code that ren­ders the ap­p’s UI in the browser

Server code that stores state and im­ple­ments the ap­p’s be­hav­ior

The server is loaded on de­mand as a Dynamic Worker and in­stan­ti­ated as a Durable Object Facet (both are fea­tures we built for this pro­ject). The facet gives the app its own SQLite data­base, sep­a­rate from the Cloudflare OS run­time man­ag­ing it. Dynamic Workers use light­weight V8 iso­lates, so every app can have its own iso­lated run­time with­out need­ing a ded­i­cated server or con­tainer sit­ting around.

The browser client talks to the server us­ing Cap’n Web, Cloudflare’s open source ob­ject-ca­pa­bil­ity Remote Procedure Call (RPC) sys­tem. A server method can be called from the client like a nor­mal JavaScript func­tion:

const is­sues = await app.lis­tIs­sues({ sta­tus: done”, });

The spe­cial part is that the agent can also call the same method.

So if you can build a tool to do a job your­self, agents can use your tool to do the job when you’re not there.

Share the app, or share how it was built

When you build an app in Cloudflare OS, you have two ways to share them:

Sharing your app it­self lets other peo­ple col­lab­o­rate in real time us­ing the same state.

Sharing a blue­print of your app lets other peo­ple cre­ate their own copy of your app.

An app in­stan­ti­ated from a blue­print con­tains the orig­i­nal ap­p’s code. But it does not con­tain its SQLite data, con­ver­sa­tion his­tory, cre­den­tials, or con­nected re­sources. Each new app starts with in­de­pen­dent state and re­sources.

This means when you share apps with your team, they can mod­ify them them­selves with AI in­stead of fil­ing a fea­ture re­quest and as­sign­ing you.

Use any model, and con­trol what it costs

Cloudflare OS can be used with any model. Every in­fer­ence call runs through Cloudflare AI Gateway, giv­ing your or­ga­ni­za­tion one place to de­cide which mod­els are avail­able and which model should han­dle each job.

Not every task needs the most ex­pen­sive model. You may not want to run the most ex­pen­sive fron­tier model to sum­ma­rize your un­read emails every morn­ing. AI Gateway gives you the con­trol needed to make sure ex­pen­sive mod­els are only be­ing used for the hard­est work.

Every re­quest is at­trib­uted to the per­son, team, or work­space that made it. Administrators can see where in­fer­ence spend is go­ing, set bud­gets and rate lim­its, and de­cide what hap­pens when a limit is reached.

Open source, so you can make it yours

Cloudflare OS is avail­able to­day and is open source. Check out the cloud­flare-os GitHub repos­i­tory. You can de­ploy it into your own Cloudflare ac­count and use your own Access poli­cies, AI Gateway con­fig­u­ra­tion, data, and in­te­gra­tions.

Our in­ter­nal de­ploy­ment re­flects Cloudflare’s sys­tems, ter­mi­nol­ogy, poli­cies, and ways of work­ing. Yours should re­flect your or­ga­ni­za­tion.

Cloudflare OS is de­signed so you can cus­tomize the in­ter­face, add in­ter­nal Gatekeepers, and build or­ga­ni­za­tion-spe­cific fea­tures with­out chang­ing the core prod­uct.

We are re­leas­ing two repos­i­to­ries: the Cloudflare OS core and an ex­am­ple de­ploy­ment based on how we run it in­ter­nally at Cloudflare. The de­ploy­ment repos­i­tory con­sumes the core with­out patch­ing it, pro­vid­ing a place for con­fig­u­ra­tion, cus­tom UI, in­ter­nal in­te­gra­tions, an­a­lyt­ics, and de­ploy­ment pipelines.

Delivered together with our part­ners

The source code is only the start­ing point. The con­text, skills, work­flows, in­ter­nal sys­tems, and poli­cies are what make Cloudflare OS even more use­ful for your or­ga­ni­za­tion.

Cloudflare’s strate­gic part­ners, Presidio and Happy Cog, will work with you to cus­tomize Cloudflare OS around how your or­ga­ni­za­tion op­er­ates and roll it out across your work­force.

Partners can help you cu­rate shared skills and in­sti­tu­tional con­text, build cus­tom in­ter­faces, con­nect in­ter­nal sys­tems through Gatekeepers and MCP Server Portals, and con­fig­ure se­cu­rity, model, and cost con­trols.

You get your own branded Cloudflare OS, con­nected to your sys­tems, run­ning on Cloudflare, and shaped around how your peo­ple ac­tu­ally work.

Get started

Cloudflare OS is avail­able to­day on GitHub. You can ex­plore the source code, try the demo, or de­ploy it into your own Cloudflare ac­count in a few min­utes us­ing our starter repos­i­tory.

We’re just get­ting started. We’re work­ing on bring­ing Cloudflare OS to the Cloudflare dash­board as a fully man­aged prod­uct, adding con­tain­ers for de­vel­op­ment work­flows, and bring­ing work­spaces into Slack and other chat tools.

If you’re in­ter­ested in talk­ing with our team, we would love to chat. Use this form to reach out!

Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)

simonwillison.net

31st July 2026

Tuesday was Stateless MCP day—the roll­out of MCP 2.0, or the 2026 – 07-28 Model Context Protocol spec­i­fi­ca­tion to use the more for­mal but less mem­o­rable name. This is the most sig­nif­i­cant change to the MCP spec since it first launched, and has also served to reignite my per­sonal in­ter­est in the pro­to­col.

For back­ground: MCP is the Model Context Protocol, which de­scribes a stan­dard way to ex­pose new tools to LLM-powered agent frame­works. It was in­tro­duced by Anthropic back in November 2024, had a huge spike of in­ter­est through much of 2025, and then be­came some­what eclipsed by Skills (another Anthropic in­ven­tion) when it be­came ap­par­ent that an agent har­ness with ac­cess to a ter­mi­nal and curl could do most of what MCP did in a more flex­i­ble way. I wrote about that in my re­view of 2025.

I’m com­ing back around to MCP now. Giving an agent a shell en­vi­ron­ment with the abil­ity to ac­cess the in­ter­net is fraught with risk, and re­quires a strong model that is ca­pa­ble of ef­fec­tively dri­ving such an en­vi­ron­ment. MCP tools are eas­ier to au­dit and con­trol, and sim­ple enough that smaller mod­els that run on a lap­top can still drive them rea­son­ably well.

The new state­less MCP spec­i­fi­ca­tion also greatly de­creases the com­plex­ity of im­ple­ment­ing both clients and servers for the pro­to­col. I built three of those this week!

What’s eas­ier with state­less MCP

The best demon­stra­tion of the dif­fer­ence be­tween state­ful and state­less MCP is in this May 21st blog post that in­tro­duced the RC for the new spec­i­fi­ca­tion. It in­cluded a clear be­fore-and-af­ter ex­am­ple.

The older state­ful MCP (I’m go­ing to call it legacy MCP) re­quired two HTTP re­quests—the first to ini­tial­ize a ses­sion and ob­tain a Mcp-Session-Id, and the sec­ond to ac­tu­ally call the tool:

POST /mcp HTTP/1.1 Content-Type: ap­pli­ca­tion/​json

{ jsonrpc”: 2.0″, id”: 1, method”: initialize”, params”: { protocolVersion”: 2025 – 11-25″, capabilities”: { }, clientInfo”: { name”: my-app”, version”: 1.0″ } } }

POST /mcp HTTP/1.1 Mcp-Session-Id: 1868a90c-3a3f-4f5b Content-Type: ap­pli­ca­tion/​json

{ jsonrpc”: 2.0″, id”: 2, method”: tools/call”, params”: { name”: search”, arguments”: { q”: otters” } } }

The new state­less way uses a sin­gle HTTP re­quest which looks like this:

POST /mcp HTTP/1.1 MCP-Protocol-Version: 2026 – 07-28 Mcp-Method: tools/​call Mcp-Name: search Content-Type: ap­pli­ca­tion/​json

{ jsonrpc”: 2.0″, id”: 1, method”: tools/call”, params”: { name”: search”, arguments”: { q”: otters” }, _meta”: { io.modelcontextprotocol/clientInfo”: { name”: my-app”, version”: 1.0” } } } }

This is so much cleaner from both a client- and server-side im­ple­men­ta­tion per­spec­tive. It’s also a bet­ter fit for build­ing scal­able web ap­pli­ca­tions, since now you don’t need to main­tain server-side state to keep track of those ses­sion IDs, or worry about rout­ing the same ses­sion to the same back­end ma­chine.

mcp-ex­plorer

I could­n’t find a great CLI tool for in­ter­ac­tively prob­ing an MCP server, so I had Codex help build my own.

mcp-ex­plorer is the re­sult. It’s a state­less Python CLI tool, so you don’t even need to in­stall it to try it out—it works with uvx like this:

uvx mcp-ex­plorer list https://​agen­tic-mer­maid.dev/​mcp

This queries Ade Oshineye’s agen­tic-mer­maid.dev demo MCP. The above com­mand re­turns the fol­low­ing list of tools:

ex­e­cute(code: string, time­outMs?: in­te­ger) - Execute Mermaid SDK code Run JavaScript in an iso­lated sand­box; re­turn a value.

de­scribe_sdk(fam­ily: string, de­tail?: string) - Describe Mermaid SDK op­er­a­tions Return ver­sion-matched mu­ta­tion op­er­a­tions for one di­a­gram fam­ily.

ren­der_svg(source: string, op­tions?: ob­ject) - Render Mermaid as SVG Render a Mermaid source string to the­me­able SVG. Returns { ok, svg }.

ren­der_ascii(source: string, use­Ascii?: boolean, tar­getWidth?: in­te­ger, op­tions?: ob­ject) - Render Mermaid as text Render a Mermaid source string to text. Returns { ok, text }.

ren­der_png(source: string, scale?: num­ber, back­ground?: string, fitTo?: ob­ject, op­tions?: ob­ject) - Render Mermaid as PNG Rasterize a Mermaid source string to PNG. Returns { ok, png_base64 }. …

Then to in­spect a tool:

uvx mcp-ex­plorer in­spect ren­der_svg

This out­puts a whole bunch of in­for­ma­tion, in­clud­ing the JSON schema of the in­puts and out­puts.

To call that tool and pass ar­gu­ments to it:

uvx mcp-ex­plorer call \ https://​agen­tic-mer­maid.dev/​mcp \ ren­der_svg \ -a source graph TD; A–>B’ \ -a op­tions {“padding”:24}’

Which re­turns:

{“ok”:true,“svg”:“<svg xmlns="hhttp://​www.w3.org/​2000/​svg\ width=…

To get just the raw SVG try adding | jq .svg -r to that com­mand. I got back this im­age:

There are a few more com­mands in the README, but you get the gen­eral idea. I find build­ing CLI tools like this to be a re­ally pro­duc­tive way to get fa­mil­iar with a spec­i­fi­ca­tion, even if an agent writes most of the ac­tual code.

datasette-mcp

The sec­ond pro­ject is datasette-mcp, a Datasette plu­gin which adds a /-/mcp end­point to any Datasette in­stance.

This is prob­a­bly the fourth time I’ve tried build­ing this plu­gin, but thanks to the new state­less MCP spec­i­fi­ca­tion I fi­nally have a ver­sion that feels good to re­lease.

It pro­vides just three tools: list_­data­bases(), get_­data­base_schema(data­base_­name), and ex­e­cute_sql(data­base_­name, sql). They do ex­actly what you would ex­pect them to do—though ex­e­cute_sql() is read-only for the mo­ment.

Wire these into an agent, or a chat tool like ChatGPT or Claude, and they’ll gain the abil­ity to run SQL queries against your hosted Datasette in­stance.

So far I’m run­ning it on the Datasette mir­ror of my blog, at datasette.si­mon­willi­son.net/-/​mcp. It took a bit of fid­dling to fig­ure out how to at­tach that to ChatGPT and Claude, but I got there in the end. Here’s a new TIL show­ing ex­actly how to do that.

Here’s a shared Claude ses­sion where I asked it:

list ta­bles in si­mon­willi­son.net

list ta­bles in si­mon­willi­son.net

And then:

what has Simon said re­cently about MCP?

what has Simon said re­cently about MCP?

It ran 7 sep­a­rate SQL queries to fig­ure out the an­swer.

llm-mcp-client

My LLM tool is long over­due for an of­fi­cial MCP in­te­gra­tion. The new al­pha llm-mcp-client plu­gin is my at­tempt at ex­actly that:

llm in­stall llm-mcp-client llm -T MCP(“https://​datasette.si­mon­willi­son.net/-/​mcp)′ count the notes’

Here’s the out­put (including rea­son­ing trace, I’m us­ing LLM 0.32rc2):

Considering note count I see the ques­tion count the notes” is prob­a­bly ask­ing me to tally up blog notes. It could also mean pub­lished notes or drafts, so there’s some am­bi­gu­ity there. I’ll need to fig­ure out the to­tal num­ber of notes, likely by query­ing the count for both pub­lished notes and drafts to get a clear an­swer. Let’s ex­e­cute that count! There are 151 notes.

Considering note count

I see the ques­tion count the notes” is prob­a­bly ask­ing me to tally up blog notes. It could also mean pub­lished notes or drafts, so there’s some am­bi­gu­ity there. I’ll need to fig­ure out the to­tal num­ber of notes, likely by query­ing the count for both pub­lished notes and drafts to get a clear an­swer. Let’s ex­e­cute that count!

There are 151 notes.

And the out­put of llm logs for that prompt.

Once this is fully baked, I’m con­sid­er­ing bring­ing it di­rectly into LLM core. I’m ex­cited to ex­per­i­ment with MCP in Datasette Agent and llm-cod­ing-agent as well.

MCP is a safer way to build with agents

A few months af­ter MCP was first re­leased, I wrote Model Context Protocol has prompt in­jec­tion se­cu­rity prob­lems, where I noted that the pat­tern of hav­ing end users mix and match tools pushed re­spon­si­bil­ity for avoid­ing data ex­fil­tra­tion at­tacks out to the users them­selves. I had­n’t coined the Lethal Trifecta yet, but that was ab­solutely what I had in mind.

Then gen­eral agents with ar­bi­trary shell and curl ac­cess came along, and that’s so much harder to keep se­cure!

Something I’ve come to ap­pre­ci­ate about MCP is that it’s much eas­ier to rea­son about agent ca­pa­bil­i­ties and what might go wrong than with ar­bi­trary com­mand ex­e­cu­tion in an open net­work en­vi­ron­ment—the de­fault for most of to­day’s gen­eral and cod­ing agent tools.

I plan to lean into MCP a whole lot more when I’m build­ing sen­si­tive ap­pli­ca­tions on top of LLMs.

Discovery Loop — Continuous Exploration

www.discoveryloop.com

Continuous Exploration

Automating dis­cov­ery to ac­cel­er­ate sci­ence and en­gi­neer­ing for the world.

Scientific dis­cov­ery is bot­tle­necked.

The sci­en­tific method is one of the great­est tools hu­man­ity has ever de­vised, yet ex­e­cu­tion en­tails repet­i­tive ex­per­i­men­tal loops that are hard to scale with to­day’s man­ual ef­forts: you pro­pose an ex­per­i­ment, im­ple­ment and run it, ex­am­ine the re­sults, then it­er­ate to re­fine your ap­proach.

Historically, sci­en­tific progress has re­lied on these se­quen­tial hu­man it­er­a­tions. In many do­mains, this process re­mains in­cred­i­bly slow and la­bor-in­ten­sive.

01 — The Approach

Automating the ex­per­i­men­tal loop.

At Discovery Loop, we are build­ing sys­tems to au­to­mate these en­tire ex­per­i­men­tal loops. By uti­liz­ing fron­tier AI mod­els and large-scale com­pu­ta­tional in­fra­struc­ture, our sys­tems will be able to rapidly pro­pose, run, and learn from eval­u­a­tions.

This ap­proach al­lows for the par­al­lel ex­e­cu­tion of thou­sands of ex­per­i­ments, dras­ti­cally com­press­ing it­er­a­tion time and dri­ving up the quan­tity and qual­ity of sci­en­tific and en­gi­neer­ing out­put.

Start with Machine Learning

We will ini­tially fo­cus on au­tomat­ing the process of ma­chine learn­ing re­search and en­gi­neer­ing.

Act as Our Own First Customer

We will use these au­to­mated ML ca­pa­bil­i­ties to rapidly op­ti­mize our own tech­nol­ogy stack be­fore ex­pand­ing to other do­mains.

Grand Challenges

We be­lieve our ap­proach will be able to solve any learn­ing loop with mea­sur­able out­comes within the do­mains of sci­ence and en­gi­neer­ing. Ultimately, we are build­ing sys­tems ca­pa­ble of tak­ing on National Academy of Engineering (NAE) Grand Challenges—such as en­gi­neer­ing bet­ter med­i­cines, ad­vanc­ing health in­for­mat­ics, mak­ing so­lar en­ergy eco­nom­i­cal, pro­vid­ing ac­cess to clean wa­ter, se­cur­ing cy­ber­space, and en­gi­neer­ing the tools of sci­en­tific dis­cov­ery.

02 — Mission

Our mis­sion is straight­for­ward: we are build­ing AI so­lu­tions that can au­to­mat­i­cally solve im­por­tant prob­lems in ma­chine learn­ing, sci­ence, and en­gi­neer­ing. By ad­vanc­ing the pace at which we con­duct en­gi­neer­ing and sci­en­tific dis­cov­ery, we can bring the ben­e­fits of sci­ence and tech­nol­ogy to the world much faster. Ultimately, our goal is to build AI sys­tems that act as a deeply pos­i­tive, em­pow­er­ing force for hu­man­ity, de­liv­er­ing tech­nol­ogy so­lu­tions that im­prove peo­ple’s lives on a global scale.

04 — The Team

The brain trust.

Our found­ing team — Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals — has a shared his­tory of deep friend­ship and decades of close and im­pact­ful col­lab­o­ra­tion.

From left Oriol Vinyals  ·  Sanjay Ghemawat  ·  Jeff Dean  ·  Quoc Le

Collectively, we rep­re­sent three of the most-cited re­searchers in ar­ti­fi­cial in­tel­li­gence and two of the most-cited re­searchers in dis­trib­uted sys­tems.

Between us, we have pi­o­neered mas­sive scale com­put­ing and led the cre­ation of crit­i­cal in­fra­struc­ture, prod­ucts, and foun­da­tional AI ad­vances that the world re­lies on, in­clud­ing mul­ti­ple gen­er­a­tions of Google Search, Google Ads, Google News, Google Translate, Google File System, MapReduce, BigTable, Spanner, TensorFlow, Pathways, TPUs, AlphaChip, AlphaStar, AlphaCode, AlphaFold, Gemini, model dis­til­la­tion, mix­ture-of-ex­perts model ar­chi­tec­tures, word2vec, se­quence-to-se­quence mod­els, chain of thought rea­son­ing, neural ar­chi­tec­ture search, and mul­ti­ple gen­er­a­tions of Large Language Models (LLMs) among oth­ers.

Our rel­a­tive ad­van­tage is­n’t just our tech­ni­cal abil­ity; it is the un­prece­dented scale of the sys­tems we have pre­vi­ously built. We pos­sess true full-stack depth that spans chips, hard­ware in­fra­struc­ture, soft­ware in­fra­struc­ture, ML mod­els, and prod­ucts.

04 — What’s Next

Imagine a fu­ture where a hand­ful of peo­ple can con­duct sci­en­tific re­search and en­gi­neer­ing tasks much more rapidly, and with higher qual­ity, than mas­sive teams of sci­en­tists and en­gi­neers do to­day. By au­tomat­ing the loops of dis­cov­ery, the world will be able to make much more rapid ad­vances across count­less fields of sci­ence.

We are build­ing a lean, in-per­son team to ex­e­cute this trans­for­ma­tive vi­sion.

Hartwork Blog · libexpat now funded by the City of Munich for up to 6 months

blog.hartwork.org

For read­ers new to Expat:

lib­ex­pat is a fast stream­ing XML parser. Alongside libxml2, Expat is one of the most widely used soft­ware li­bre XML parsers writ­ten in C, specif­i­cally C99. It is cross-plat­form and li­censed un­der the MIT li­cense.

lib­ex­pat is a fast stream­ing XML parser. Alongside libxml2, Expat is one of the most widely used soft­ware li­bre XML parsers writ­ten in C, specif­i­cally C99. It is cross-plat­form and li­censed un­der the MIT li­cense.

Starting 2026 – 08-01, the security va­ca­tion” of the pro­ject has ended and(!) I will be be paid to work on main­tain­ing lib­ex­pat for up to 6 months thanks to the City of Munich un­der the um­brella of their Open Source Sabbatical pro­gram. What does that mean?

For much of the past 10 years, work­ing on lib­ex­pat has been com­pet­ing with my reg­u­lar oc­cu­pa­tion as a soft­ware en­gi­neer, chores, so­cial life and re-cre­ation. For the first time, I am now be­ing em­ployed to work on main­tain­ing lib­ex­pat as my regular job” for a lim­ited pe­riod of time. My top pri­or­i­ties will be:

Fixing the cur­rently 5 known un­fixed vul­ner­a­bil­i­ties

Fixing the cur­rently 5 known un­fixed vul­ner­a­bil­i­ties

Adding sup­port for XML 1.0r5

Adding sup­port for XML 1.0r5

Further im­prov­ing the ro­bust­ness and main­tain­abil­ity of the pro­ject

Further im­prov­ing the ro­bust­ness and main­tain­abil­ity of the pro­ject

Yesterday and to­day most of my time went into fix­ing a vul­ner­a­bil­ity un­cov­ered by Mozilla.

Technically, I am be­ing em­ployed by digi­tial@M now for of up 6 months with a reg­u­lar work­ing con­tract, in­clud­ing can­cel­la­tion by ei­ther party, re­motely from home. There is plenty to do.

Unvalidated AI slop sub­mis­sions will still not be ap­pre­cated, but for every­thing else: if you want to throw in­tel­li­gence at find­ing fur­ther vul­ner­a­bil­i­ties in lib­ex­pat and send them my way, the com­ing months will be the best chance at get­ting things fixed in rea­son­able time. Queueing the­ory and laws of physics still ap­ply.

Wish me luck!

PS: If any­one man­aged to com­bine Clang-based MinGW with AddressSanitizer and Wine with­out crash­ing at launch, please show me how and drop me an e-mail. Thank you!

Best, Sebastian

nytimes.com

www.nytimes.com

Please en­able JS and dis­able any ad blocker

Cops Used Flock to Track a Man Across State Lines to Create Pretext to Search His Car for Weed

www.404media.co

Police in Wisconsin used Flock to de­ter­mine that a man travels to Michigan fre­quently,” where mar­i­juana is le­gal, then back to Wisconsin, where it is il­le­gal. They then used his travel across state lines as tracked by Flock as part of the prob­a­ble cause jus­ti­fi­ca­tion to search his car for weed; he was even­tu­ally ar­rested on mar­i­juana pos­ses­sion charges, ac­cord­ing to court records re­viewed by 404 Media.

The searches came to light in a Wisconsin crim­i­nal com­plaint against Edward Abrams-Phillips, who was wanted for bail jump­ing on do­mes­tic vi­o­lence charges. But the crim­i­nal com­plaint makes clear that be­yond the bail jump­ing and do­mes­tic vi­o­lence charges, po­lice specif­i­cally stud­ied Abrams-Phillips’ in­ter­state travel to cre­ate the pre­text for search­ing his car for mar­i­juana. The bail jump­ing charge was dis­missed; Abrams-Phillips was found guilty only of weed pos­ses­sion in the case, ac­cord­ing to the court records.

The com­plaint ex­plains that Abrams-Phillips was tracked via Flock’s net­work over the course of the day to de­ter­mine that he drove from Wisconsin to Michigan, a known source state for mar­i­juana as it is le­gal there,” the com­plaint states, adding that pre­vi­ous Flock hits in­di­cated that he travels to Michigan fre­quently.” Police note that, us­ing Flock, they were able to track Abrams-Phillips dri­ving from Wisconsin to Michigan, then back to Wisconsin over the course of sev­eral hours, where he was pulled over and ar­rested. The Flock searches and ar­rests hap­pened in April 2025.

The ve­hi­cle was ob­served hit­ting flock on sev­eral oc­ca­sions to in­clude 41 north­bound from Brown Rd, 41NB and County Line in Marinette [Wisconsin], and 41 NB on Bridge St. go­ing into Michigan. Based on prior flock hits, the ve­hi­cle trav­els to Michigan fre­quently which is a known source State for Marijuana as it is le­gal there,” the charg­ing doc­u­ment notes. Around 3:56 p.m., the ve­hi­cle was seen on Flock head­ing south­bound on in­ter­state 41 to­wards Green Bay [Wisconsin]. Deputies made a co­or­di­nated ef­fort to in­ter­cept the ve­hi­cle on 41 from Brown Rd. Deputy Kowalski ini­ti­ated a traf­fic stop on the ve­hi­cle as the dri­ver matched the de­scrip­tion of Edward.”

This post is for paid mem­bers only

Become a paid mem­ber for un­lim­ited ad-free ac­cess to ar­ti­cles, bonus pod­cast con­tent, and more.

Subscribe

Sign up for free ac­cess to this post

Free mem­bers get ac­cess to posts like this one along with an email round-up of our week’s sto­ries.

Subscribe

Already have an ac­count? Sign in

AI fuels more than half of cybercrime in Africa as digital scams surge, INTERPOL

www.africanews.com

Artificial in­tel­li­gence is now pow­er­ing more than half of re­ported cy­ber­crime across Africa, al­low­ing crim­i­nals to launch faster, more con­vinc­ing and larger-scale at­tacks, ac­cord­ing to INTERPOLs African Cyberthreat Assessment Report 2026.

The re­port found that 55% of cy­ber­crime cases recorded across the con­ti­nent in­volve the use of AI, rais­ing con­cerns as Africa’s dig­i­tal econ­omy con­tin­ues to ex­pand.

With more than 1.1 bil­lion mo­bile sub­scribers in 2025, mil­lions of peo­ple are re­ly­ing on dig­i­tal ser­vices, cre­at­ing new op­por­tu­ni­ties for both in­no­va­tion and cy­ber­crim­i­nals.

Based on data from 36 African coun­tries, the 40-page as­sess­ment says cy­ber­crime has evolved into a highly or­gan­ised, cross-bor­der in­dus­try that is be­com­ing harder for au­thor­i­ties to de­tect and stop.

Online scams re­main Africa’s biggest cy­ber threat

According to the re­port, on­line scams re­mained the most com­mon form of cy­ber­crime in 2025. Criminals in­creas­ingly used ar­ti­fi­cial in­tel­li­gence along­side so­cial me­dia plat­forms and mo­bile money ser­vices to tar­get vic­tims.

INTERPOL said cy­ber­crime-re­lated fi­nan­cial losses have risen sharply over the past year, climb­ing from $192 mil­lion in 2024 to $484 mil­lion. Investigators at­tribute the in­crease to AI-powered fraud, stolen lo­gin cre­den­tials and so­phis­ti­cated so­cial en­gi­neer­ing at­tacks.

The re­port also found that 72% of sur­veyed coun­tries iden­ti­fied scam cen­tres op­er­at­ing within their bor­ders, with the high­est con­cen­tra­tion in West and Southern Africa.

Different re­gions face dif­fer­ent cy­ber risks

The re­port high­lights dis­tinct cy­ber­crime trends across the con­ti­nent.

In East Africa, mo­bile money fraud and ran­somware at­tacks tar­get­ing crit­i­cal in­fra­struc­ture are among the biggest threats.

West and Central Africa con­tinue to ex­pe­ri­ence high lev­els of busi­ness email com­pro­mise and ro­mance scams af­fect­ing both com­pa­nies and in­di­vid­u­als.

Meanwhile, Southern Africa’s ad­vanced dig­i­tal con­nec­tiv­ity has made the re­gion an at­trac­tive tar­get for in­ter­na­tional cy­ber­crim­i­nal net­works seek­ing to max­imise dis­rup­tion.

AI is mak­ing cy­ber­crime more con­vinc­ing

INTERPOL warned that ar­ti­fi­cial in­tel­li­gence is trans­form­ing the way cy­ber­crim­i­nals op­er­ate.

Deepfake tech­nol­ogy and AI-generated con­tent are in­creas­ingly be­ing used in dig­i­tal sex­tor­tion and on­line ha­rass­ment cam­paigns. One of INTERPOLs tech­nol­ogy part­ners, TrendAI, de­tected around 600,000 sex­tor­tion cases linked to these tac­tics.

The re­port also noted a sharp rise in Business Email Compromise (BEC) scams, where crim­i­nals use AI to pro­duce re­al­is­tic emails that im­i­tate trusted con­tacts.

Some Africa-based cy­ber­crim­i­nal groups have tar­geted busi­nesses and in­di­vid­u­als in Europe and North America, us­ing in­fra­struc­ture spread across sev­eral coun­tries to hide their ac­tiv­i­ties.

Another grow­ing con­cern is the use of syn­thetic iden­ti­ties. Rather than sim­ply steal­ing per­sonal in­for­ma­tion, cy­ber­crim­i­nals are com­bin­ing gen­uine data with fab­ri­cated de­tails to cre­ate en­tirely new dig­i­tal iden­ti­ties.

These fake pro­files have re­port­edly been used to open bank ac­counts, ob­tain mo­bile loans and reg­is­ter SIM cards while evad­ing some bio­met­ric ver­i­fi­ca­tion sys­tems.

Gaps in co­op­er­a­tion leave fi­nan­cial sys­tems ex­posed

INTERPOL said weak co­or­di­na­tion be­tween banks, tele­com com­pa­nies and law en­force­ment agen­cies con­tin­ues to ham­per ef­forts to com­bat cy­ber­crime.

The ab­sence of real-time in­for­ma­tion shar­ing cre­ates op­por­tu­ni­ties for crim­i­nals to move stolen funds quickly and ex­ploit weak­nesses across mul­ti­ple ju­ris­dic­tions be­fore au­thor­i­ties can re­spond.

The re­port also found that many African law en­force­ment agen­cies are still not ad­e­quately pre­pared to re­spond to AI-driven cy­ber threats, de­spite the rapid pace at which the tech­nol­ogy is be­ing adopted by crim­i­nal net­works.

Countries step up ef­forts against cy­ber­crime

Despite the grow­ing threat, the re­port points to progress across the con­ti­nent.

In 2025, 17 African coun­tries in­tro­duced or up­dated cy­ber­crime leg­is­la­tion. Senegal also launched an on­line re­port­ing plat­form de­signed to im­prove re­sponses to on­line of­fences in­volv­ing chil­dren.

INTERPOL said joint in­ter­na­tional op­er­a­tions have also de­liv­ered sig­nif­i­cant re­sults. Four ma­jor op­er­a­tions, Operation Serengeti 2.0, Operation Contender 3.0, Operation Sentinel and Operation Red Card 2.0, led to more than 1,500 ar­rests, the seizure of hun­dreds of elec­tronic de­vices and the re­cov­ery of over $100 mil­lion linked to cy­ber­crime.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.