10 interesting stories served every morning and every evening.

Discovery Loop — Continuous Exploration

www.discoveryloop.com

Continuous Exploration

Automating dis­cov­ery to ac­cel­er­ate sci­ence and en­gi­neer­ing for the world.

Scientific dis­cov­ery is bot­tle­necked.

The sci­en­tific method is one of the great­est tools hu­man­ity has ever de­vised, yet ex­e­cu­tion en­tails repet­i­tive ex­per­i­men­tal loops that are hard to scale with to­day’s man­ual ef­forts: you pro­pose an ex­per­i­ment, im­ple­ment and run it, ex­am­ine the re­sults, then it­er­ate to re­fine your ap­proach.

Historically, sci­en­tific progress has re­lied on these se­quen­tial hu­man it­er­a­tions. In many do­mains, this process re­mains in­cred­i­bly slow and la­bor-in­ten­sive.

01 — The Approach

Automating the ex­per­i­men­tal loop.

At Discovery Loop, we are build­ing sys­tems to au­to­mate these en­tire ex­per­i­men­tal loops. By uti­liz­ing fron­tier AI mod­els and large-scale com­pu­ta­tional in­fra­struc­ture, our sys­tems will be able to rapidly pro­pose, run, and learn from eval­u­a­tions.

This ap­proach al­lows for the par­al­lel ex­e­cu­tion of thou­sands of ex­per­i­ments, dras­ti­cally com­press­ing it­er­a­tion time and dri­ving up the quan­tity and qual­ity of sci­en­tific and en­gi­neer­ing out­put.

Start with Machine Learning

We will ini­tially fo­cus on au­tomat­ing the process of ma­chine learn­ing re­search and en­gi­neer­ing.

Act as Our Own First Customer

We will use these au­to­mated ML ca­pa­bil­i­ties to rapidly op­ti­mize our own tech­nol­ogy stack be­fore ex­pand­ing to other do­mains.

Grand Challenges

We be­lieve our ap­proach will be able to solve any learn­ing loop with mea­sur­able out­comes within the do­mains of sci­ence and en­gi­neer­ing. Ultimately, we are build­ing sys­tems ca­pa­ble of tak­ing on National Academy of Engineering (NAE) Grand Challenges—such as en­gi­neer­ing bet­ter med­i­cines, ad­vanc­ing health in­for­mat­ics, mak­ing so­lar en­ergy eco­nom­i­cal, pro­vid­ing ac­cess to clean wa­ter, se­cur­ing cy­ber­space, and en­gi­neer­ing the tools of sci­en­tific dis­cov­ery.

02 — Mission

Our mis­sion is straight­for­ward: we are build­ing AI so­lu­tions that can au­to­mat­i­cally solve im­por­tant prob­lems in ma­chine learn­ing, sci­ence, and en­gi­neer­ing. By ad­vanc­ing the pace at which we con­duct en­gi­neer­ing and sci­en­tific dis­cov­ery, we can bring the ben­e­fits of sci­ence and tech­nol­ogy to the world much faster. Ultimately, our goal is to build AI sys­tems that act as a deeply pos­i­tive, em­pow­er­ing force for hu­man­ity, de­liv­er­ing tech­nol­ogy so­lu­tions that im­prove peo­ple’s lives on a global scale.

04 — The Team

The brain trust.

Our found­ing team — Jeff Dean, Sanjay Ghemawat, Quoc Le, and Oriol Vinyals — has a shared his­tory of deep friend­ship and decades of close and im­pact­ful col­lab­o­ra­tion.

From left Oriol Vinyals  ·  Sanjay Ghemawat  ·  Jeff Dean  ·  Quoc Le

Collectively, we rep­re­sent three of the most-cited re­searchers in ar­ti­fi­cial in­tel­li­gence and two of the most-cited re­searchers in dis­trib­uted sys­tems.

Between us, we have pi­o­neered mas­sive scale com­put­ing and led the cre­ation of crit­i­cal in­fra­struc­ture, prod­ucts, and foun­da­tional AI ad­vances that the world re­lies on, in­clud­ing mul­ti­ple gen­er­a­tions of Google Search, Google Ads, Google News, Google Translate, Google File System, MapReduce, BigTable, Spanner, TensorFlow, Pathways, TPUs, AlphaChip, AlphaStar, AlphaCode, AlphaFold, Gemini, model dis­til­la­tion, mix­ture-of-ex­perts model ar­chi­tec­tures, word2vec, se­quence-to-se­quence mod­els, chain of thought rea­son­ing, neural ar­chi­tec­ture search, and mul­ti­ple gen­er­a­tions of Large Language Models (LLMs) among oth­ers.

Our rel­a­tive ad­van­tage is­n’t just our tech­ni­cal abil­ity; it is the un­prece­dented scale of the sys­tems we have pre­vi­ously built. We pos­sess true full-stack depth that spans chips, hard­ware in­fra­struc­ture, soft­ware in­fra­struc­ture, ML mod­els, and prod­ucts.

04 — What’s Next

Imagine a fu­ture where a hand­ful of peo­ple can con­duct sci­en­tific re­search and en­gi­neer­ing tasks much more rapidly, and with higher qual­ity, than mas­sive teams of sci­en­tists and en­gi­neers do to­day. By au­tomat­ing the loops of dis­cov­ery, the world will be able to make much more rapid ad­vances across count­less fields of sci­ence.

We are build­ing a lean, in-per­son team to ex­e­cute this trans­for­ma­tive vi­sion.

The next chapter of our AI momentum

blog.google

Editor’s note: Today, Google and Alphabet CEO Sundar Pichai shared some changes with Google DeepMind teams, in­clud­ing new roles for Demis Hassabis and Koray Kavukcuoglu. Below are the mes­sages Sundar and Demis sent to em­ploy­ees.

Message from Sundar Pichai

We’ve made ex­tra­or­di­nary progress to de­liver on our full AI stack. We’ve got amaz­ing tal­ent, world-class com­pute, and prod­ucts that bring AI to more peo­ple than any other com­pany. And you saw the in­cred­i­ble mo­men­tum at earn­ings across all our busi­nesses, in­clud­ing Search, YouTube, and Cloud. Our Gemini mod­els are in high de­mand among de­vel­op­ers and busi­nesses, and the Gemini app reached 950M+ monthly users. Meanwhile, our AI re­search con­tin­ues to drive field-defin­ing break­throughs (like last week’s Gemini Robotics ad­vances).

We have to ac­cel­er­ate all this work and stay fo­cused on the AI fron­tier. At the same time, there’s never been a more im­por­tant mo­ment to shape the fu­ture of AGI and sci­ence. Today Demis, Koray and I are shar­ing a few changes to our Google DeepMind teams that will en­able us to do both.

AGI and sci­ence: Demis has de­scribed us as stand­ing in the foothills of the sin­gu­lar­ity, and has been spend­ing a lot of his time en­gag­ing ex­ter­nally. He and I have been long dis­cussing a role that al­lows him to put his full at­ten­tion on ac­tively shap­ing the fu­ture of AGI. It’s work that is vi­tally im­por­tant to Alphabet and hu­man­ity, and I can’t imag­ine a bet­ter per­son than Demis to do it. So, mov­ing for­ward, Demis will be­come the Chair of GDM and Chief Scientist of Alphabet, while con­tin­u­ing to lead Isomorphic Labs. He’ll re­main closely con­nected to Koray, Josh, and our GDM teams, ad­vis­ing across mod­els and re­search. I’m so ex­cited for Demis — this is truly his life’s work and pur­pose. You can read Demis’s note to GDM be­low.

Google DeepMind: We are build­ing strong mo­men­tum: Flash is in high de­mand, our Cyber model is live, and Gemma mod­els have sur­passed 900M+ down­loads. We are com­mit­ted to be­ing at the fron­tier, and are su­per fo­cused on the ar­eas where we need to im­prove. I’m re­ally ex­cited for our up­com­ing model re­leases and the progress we’re see­ing. We have to con­tinue to move fast and with clear pur­pose here. Koray, the cur­rent Chief Technology Officer of GDM and our Chief AI Architect, will step up as SVP of Google DeepMind, re­port­ing to me. He will over­see Gemini model de­vel­op­ment, Frontier AI re­search, and the Gemini app and de­vel­oper teams. Koray has been at DeepMind since its early days, and over his 13 years there, he has started our deep learn­ing team and led the way on break­throughs like WaveNet and DQN. I look for­ward to see­ing him lead GDM into this next chap­ter.

Lastly, af­ter an in­cred­i­ble 27-year run, Jeff Dean is at a mo­ment where he wants to try some­thing new, and we’re ex­cited to sup­port him in that. Jeff and Google Senior Fellow Sanjay Ghemawat are launch­ing an in­de­pen­dent pub­lic ben­e­fit cor­po­ra­tion to ac­cel­er­ate dis­cov­er­ies in ML, sci­ence, and en­gi­neer­ing. Jeff and Sanjay helped to drive some of the most sig­nif­i­cant tech­nol­ogy tran­si­tions, from our early search in­fra­struc­ture to the neural net­works that helped cre­ate the mod­ern AI era. On a per­sonal note, it’s been a priv­i­lege to work along­side Jeff and Sanjay, and I wish them all the best! We’ll con­tinue to work with them as a found­ing in­vestor and Cloud part­ner, and col­lab­o­rate on a re­search frame­work for ML sys­tems and re­lated in­fra­struc­ture ad­vances.

We are at a dy­namic mo­ment with so much op­por­tu­nity ahead. With to­day’s changes we’re go­ing to keep dri­ving our mo­men­tum. Onwards!

-Sundar

Message from Demis Hassabis

Hi Team

We have ar­rived at a piv­otal mo­ment in hu­man his­tory. I’ve been work­ing to­wards AGI my whole life and now, like many of you, I feel it is close at hand. It’s crit­i­cal that we col­lec­tively get the next steps right to en­sure this all goes well for hu­man­ity and we usher in an in­cred­i­ble new age of dis­cov­ery and won­der.

With this back­drop, I’ve de­cided that now is the right time for me to hand over my day-to-day op­er­a­tional re­spon­si­bil­i­ties at GDM, so that I have the time and space to fo­cus on the big pic­ture and help in­flu­ence what is to come to the best of my abil­ity. I will be tak­ing on a new strate­gic role as Chair of GDM and Chief Scientist of Alphabet, and I’m ex­cited to an­nounce that Koray will be step­ping up to lead GDM as SVP of Google DeepMind, in ad­di­tion to his role as Chief AI Architect of Google.

Koray and I have been work­ing to­gether for over 13 years, since the early days of DeepMind. He is one of the world’s fore­most AI ex­perts and has been cham­pi­oning GDMs mis­sion from day one. I have to­tal con­fi­dence in Koray, Josh, and the rest of the GDM exec team as they con­tinue to spear­head the lat­est AI de­vel­op­ments across Google. The Gemini mod­els are in good hands with Koray and the leads, as they have been for a while, and I’m ex­cited about the great progress we’re mak­ing with our new mod­els in­clud­ing Gemini 4.

In my new role, I will con­tinue to work closely with Sundar on strate­gic and global AGI mat­ters, and to ad­vise Koray, Josh, and the GDM leads, from our awe­some new London Platform 37 of­fices. As part of this tran­si­tion, I’ll also be lean­ing into my role at Isomorphic, where we are mak­ing ex­tremely rapid and promis­ing progress, to ac­cel­er­ate our mis­sion there even faster. As you’ve heard me say many times, I’ve al­ways be­lieved the No.1 ap­pli­ca­tion of AI should be to im­prove hu­man health. It’s time for AI to prove its un­equiv­o­cal value to the world, and what bet­ter way to demon­strate that than to help fi­nally cure dis­eases like can­cer.

We’ve built a unique cul­ture at GDM that has served us very well. I want to thank each and every one of you for your bril­liance, ded­i­ca­tion, and ef­fort that make GDM the huge suc­cess it is to­day. We should all be ex­tremely proud of the amaz­ing things we’ve achieved so far. We’ve be­come the AI en­gine room of Google, with Gemini de­liv­er­ing help­ful ex­pe­ri­ences every­where in­clud­ing AI Mode and AI Overviews, the Gemini App rock­et­ing to over 950M monthly users, and our fun­da­men­tal and sci­en­tific re­search con­tin­ues to lead the world. I’m very ex­cited for our next chap­ter and the best is yet to come!

As a busi­ness we are in an in­cred­i­bly strong po­si­tion. We are the only com­pany that has the full stack and we’re world-class at every layer from in­fra­struc­ture to cloud to fron­tier mod­els to AI-first ap­pli­ca­tions. We have all the in­gre­di­ents to lead from here, and I firmly be­lieve we will.

Best

Demis

Cloudflare OS: an open platform for agents, apps, and work

blog.cloudflare.com

Every or­ga­ni­za­tion has a mis­sion, a rea­son for be­ing. Organizations pass that mis­sion — along with their ter­mi­nol­ogy, pro­ce­dures, sys­tems, stan­dards, and ways of work­ing — to their peo­ple. People, in turn, take this con­text to­gether with their own ex­pe­ri­ence and work to­wards the mis­sion.

Work can take many forms, from code, to doc­u­ments and slides, to re­la­tion­ships, to out­comes in the phys­i­cal world.

Some of these are straight­for­ward: code ei­ther runs or it does­n’t. Agents have been us­ing this feed­back loop to pro­duce code that works” for de­vel­op­ers over the last cou­ple of years. But what about the rest of us?

Bringing the same lever­age to the rest of the or­ga­ni­za­tion is a harder prob­lem. Agents need to un­der­stand the con­text of the com­pany and be able to reach the sys­tems peo­ple use to do their jobs. They need to turn that con­text and ac­cess into work that moves the or­ga­ni­za­tion to­wards its mis­sion.

That’s why we cre­ated Cloudflare OS. It gives every per­son an agent and work­space built around their com­pany: how it works, what it knows, and the sys­tems it re­lies on.

In May of this year, we gave every per­son at Cloudflare ac­cess to the first ver­sion of Cloudflare OS. Thousands of peo­ple across every func­tion, many of them out­side of en­gi­neer­ing, use it every day to cre­ate doc­u­ments and slides, au­to­mate re­peat­able tasks, and build small apps to vi­su­al­ize data and help them do their work.

Cloudflare OS also gave every­one a shared li­brary of con­text and skills built by teams at Cloudflare. It cap­tures our ter­mi­nol­ogy, pro­ce­dures, and best-known ways of do­ing re­cur­ring work as in­struc­tions an agent can fol­low. When one per­son fig­ures out a bet­ter way to do some­thing, every­one else can use it.

Today, we are open sourc­ing a new ver­sion of Cloudflare OS. Any or­ga­ni­za­tion can de­ploy it, con­nect it to in­ter­nal sys­tems, and make it their own.

What we learned from the first ver­sion

The Cloudflare OS we are open sourc­ing to­day is based on what we learned from run­ning the first ver­sion in­ter­nally, a jour­ney our CIO, Sam Rhea, cov­ers in his blog post.

The first ver­sion cen­tered on in­di­vid­u­als work­ing with agents through pri­vate work­spaces. Apps were sta­tic rather than live soft­ware con­nected to in­ter­nal sys­tems, and mostly de­ter­min­is­tic jobs still re­quired run­ning an agent skill again and con­sum­ing more model to­kens.

Collaboration ex­posed a more fun­da­men­tal chal­lenge. Access to an MCP server told us which tools an agent could call, but not which un­der­ly­ing re­sources the agent had ob­served. Once peo­ple be­gan shar­ing work­spaces, apps, and out­puts, we needed to en­sure that col­lab­o­ra­tion could not ex­pose in­for­ma­tion some­one was not per­mit­ted to see.

We re­built Cloudflare OS on a new foun­da­tion to solve these prob­lems. Security had to be part of the plat­form, not some­thing every per­son build­ing an app or us­ing an agent has to im­ple­ment cor­rectly.

The re­sult is a plat­form de­signed to be­long to the com­pany run­ning it. You can cus­tomize the in­ter­faces, con­nect your tools, and add the skills and con­text that cap­ture how your or­ga­ni­za­tion works.

Introducing Cloudflare OS

Cloudflare OS starts with a con­ver­sa­tion in your browser, like many other AI tools. What makes it dif­fer­ent is that each con­ver­sa­tion is grounded in the con­text and skills your or­ga­ni­za­tion has cu­rated. Give your work­space a goal, and it can draw on that knowl­edge and work with the tools and data your or­ga­ni­za­tion al­ready uses to achieve it.

Cloudflare OS com­bines three parts:

An agent work­space grounded in con­text and skills your com­pany cu­rates, with an iso­lated run­time where agents can write and run code.

A new se­cu­rity and gov­er­nance frame­work for safe ac­cess to in­ter­nal data and ser­vices.

A plat­form for per­sonal, mod­i­fi­able apps that peo­ple can build, share, and con­tinue chang­ing.

What be­gins as a con­ver­sa­tion can be­come a doc, an app, or a work­flow that con­tin­ues do­ing the work.

An agent work­space for every­one in your com­pany

Agent work­spaces were de­signed for every­one in your or­ga­ni­za­tion to use. You in­ter­act with them in your browser, so you don’t have to be a de­vel­oper or know how to use a ter­mi­nal.

A work­space com­bines agent ses­sions, per­sis­tent state, out­puts and files, re­source ac­cess, and an iso­lated run­time where the agent can write and run code.

They come loaded with the cu­rated con­text and skills your team or com­pany has col­lected. No more rein­vent­ing the wheel for every task — if some­one on your team has fig­ured out the best way to do some­thing, every­one ben­e­fits. People no longer have to ex­plain the same process, ter­mi­nol­ogy, and best prac­tices to a model every time they start a task.

A few things you can do:

Research and ask ques­tions

Ask a work­space to re­search a topic us­ing com­pany con­text and the re­sources you make avail­able to it. The agent can write code to search, fil­ter, join, and an­a­lyze in­for­ma­tion in­stead of pulling an en­tire dataset into the mod­el’s con­text win­dow.

Create docs, slides, and spread­sheets

A work­space can turn its re­search into a doc­u­ment, pre­sen­ta­tion, or spread­sheet that you can con­tinue edit­ing. These out­puts do not have to be sta­tic files. They can re­main con­nected to live data, be up­dated as their sources change, and still be ex­ported to fa­mil­iar for­mats or ser­vices such as Google Drive.

Create col­lab­o­ra­tive, con­nected apps for your team

When a doc­u­ment or spread­sheet is not enough, the agent can build an app with its own in­ter­face, logic, and state. The app can use con­nected com­pany re­sources and sup­port mul­ti­ple peo­ple work­ing to­gether.

Run de­ter­min­is­tic work­flows

Not every job needs a full agent ses­sion. Many are a known se­quence of steps with one or two places where judg­ment is use­ful. A work­space can turn those jobs into mostly de­ter­min­is­tic work­flows, us­ing code for the pre­dictable steps and a model only where it adds value. Workflows can run on de­mand, on a sched­ule, or when an event oc­curs in a con­nected sys­tem.

Cloudflare OS gives agents and apps gov­erned ac­cess to sys­tems of record through Gatekeepers (more on this in the se­cu­rity sec­tion be­low). It also sup­ports ex­ist­ing Model Context Protocol (MCP) servers your or­ga­ni­za­tion al­ready uses via MCP Server Portals.

A new se­cu­rity and gov­er­nance frame­work for safe ac­cess to in­ter­nal data and ser­vices

As peo­ple be­gin ex­per­i­ment­ing with AI at work, one of their first re­quests is of­ten for API keys to com­pany sys­tems. This makes sense: AI is­n’t much use at work if it does­n’t have ac­cess to the sys­tems peo­ple use to do their jobs.

But hand­ing over API keys to peo­ple and agents is dan­ger­ous and does not scale. Keys of­ten pro­vide broad, long-lived ac­cess that is dif­fi­cult to con­strain, share safely, and au­dit.

MCP gives agents a bet­ter way to use these sys­tems. An MCP server can hold the cre­den­tial and ex­pose a de­fined set of tools in­stead of hand­ing the key di­rectly to the agent. But con­trol­ling which tools an agent can call is only the first step. MCP alone does not tell us which un­der­ly­ing re­sources an agent has ob­served. The agent can com­bine in­for­ma­tion across sys­tems, send it some­where less re­stricted, or ex­pose it through apps and out­puts to peo­ple who may not be al­lowed to see the orig­i­nal re­sources. Authorization has to ac­count for where the data can go next.

Agents start with no ac­cess

Cloudflare Access con­trols who can en­ter Cloudflare OS. Inside, every agent and app starts with ac­cess to noth­ing. An agent can ask for ac­cess to a spe­cific re­source, which you can grant or deny. Generated code re­ceives that re­source as a typed bind­ing:

const is­sues = await env.PRO­JECT.lis­tIs­sues({ teamId: ENG, state: open”, });

env.PRO­JECT is a ca­pa­bil­ity rep­re­sent­ing per­mis­sion to use a spe­cific re­source un­der a spe­cific pol­icy. The cre­den­tial re­mains com­pletely iso­lated from the agent and any gen­er­ated code.

Server code runs in a Dynamic Worker with global out­bound net­work­ing dis­abled. Client code runs in a sand­boxed frame in the browser. Neither can reach the Internet ex­cept through ca­pa­bil­i­ties you ex­plic­itly pro­vide.

Gatekeepers gov­ern re­sources and ac­tions

A Gatekeeper is a ser­vice-spe­cific Worker that sits be­tween Cloudflare OS and an ex­ter­nal ser­vice. It un­der­stands the ser­vice’s API, its re­sources, and the op­er­a­tions that can be per­formed on them.

Giving an agent ac­cess to your en­tire GitHub ac­count is likely too broad. A Gatekeeper can give it ac­cess to a sin­gle repos­i­tory, al­low it to read is­sues but not source code, mask par­tic­u­lar fields, ap­ply rate lim­its, and re­quire ap­proval be­fore merg­ing a pull re­quest.

The agent and its apps see a small TypeScript API. The Gatekeeper han­dles OAuth, holds the cre­den­tial, en­forces pol­icy, records what was read, and me­di­ates any­thing with an ex­ter­nally vis­i­ble side ef­fect.

Policy fol­lows what the agent has seen

Controlling the ini­tial read is not enough. Take, for ex­am­ple, the case where an agent reads a sen­si­tive table in a data ware­house and uses it to pro­duce a live dash­board. Sharing the dash­board must not be­come a way to share the table with peo­ple who could not ac­cess it di­rectly.

Cloudflare OS records every re­source agents ob­serve. These ob­ser­va­tions re­main at­tached to the agent and its work. When an­other per­son tries to open the work­space, in­ter­act with the agent, or view what it pro­duced, Gatekeepers ver­ify that per­son’s ac­cess to the ob­served re­sources.

The same ob­ser­va­tion log is used to in­form poli­cies that de­ter­mine when agents can make ex­ter­nal re­quests. A read of sen­si­tive data can pre­vent the agent from writ­ing data to cer­tain sources, invit­ing new col­lab­o­ra­tors, hand­ing work to an­other agent, or mak­ing an out­bound re­quest.

People us­ing agents or build­ing apps do not have to worry about mak­ing these mis­takes. The plat­form can now be used to han­dle this.

A plat­form for build­ing and shar­ing per­sonal, mod­i­fi­able apps

Most pro­duc­tiv­ity suites give you a fixed set of ap­pli­ca­tions: doc­u­ments, spread­sheets, and pre­sen­ta­tions. In Cloudflare OS, each file” can be its own ap­pli­ca­tion, writ­ten by an agent for one per­son, one pro­ject, or one team.

These are not pro­to­types that you have to ex­port and de­ploy some­where else. Each one is a full-stack ap­pli­ca­tion with client code, server code, an API, and durable state. Apps are pri­vate by de­fault, but can be shared like doc­u­ments.

Every app is a Worker

When you ask your work­space to build an app, the agent writes two parts:

Client code that ren­ders the ap­p’s UI in the browser

Server code that stores state and im­ple­ments the ap­p’s be­hav­ior

The server is loaded on de­mand as a Dynamic Worker and in­stan­ti­ated as a Durable Object Facet (both are fea­tures we built for this pro­ject). The facet gives the app its own SQLite data­base, sep­a­rate from the Cloudflare OS run­time man­ag­ing it. Dynamic Workers use light­weight V8 iso­lates, so every app can have its own iso­lated run­time with­out need­ing a ded­i­cated server or con­tainer sit­ting around.

The browser client talks to the server us­ing Cap’n Web, Cloudflare’s open source ob­ject-ca­pa­bil­ity Remote Procedure Call (RPC) sys­tem. A server method can be called from the client like a nor­mal JavaScript func­tion:

const is­sues = await app.lis­tIs­sues({ sta­tus: done”, });

The spe­cial part is that the agent can also call the same method.

So if you can build a tool to do a job your­self, agents can use your tool to do the job when you’re not there.

Share the app, or share how it was built

When you build an app in Cloudflare OS, you have two ways to share them:

Sharing your app it­self lets other peo­ple col­lab­o­rate in real time us­ing the same state.

Sharing a blue­print of your app lets other peo­ple cre­ate their own copy of your app.

An app in­stan­ti­ated from a blue­print con­tains the orig­i­nal ap­p’s code. But it does not con­tain its SQLite data, con­ver­sa­tion his­tory, cre­den­tials, or con­nected re­sources. Each new app starts with in­de­pen­dent state and re­sources.

This means when you share apps with your team, they can mod­ify them them­selves with AI in­stead of fil­ing a fea­ture re­quest and as­sign­ing you.

Use any model, and con­trol what it costs

Cloudflare OS can be used with any model. Every in­fer­ence call runs through Cloudflare AI Gateway, giv­ing your or­ga­ni­za­tion one place to de­cide which mod­els are avail­able and which model should han­dle each job.

Not every task needs the most ex­pen­sive model. You may not want to run the most ex­pen­sive fron­tier model to sum­ma­rize your un­read emails every morn­ing. AI Gateway gives you the con­trol needed to make sure ex­pen­sive mod­els are only be­ing used for the hard­est work.

Every re­quest is at­trib­uted to the per­son, team, or work­space that made it. Administrators can see where in­fer­ence spend is go­ing, set bud­gets and rate lim­its, and de­cide what hap­pens when a limit is reached.

Open source, so you can make it yours

Cloudflare OS is avail­able to­day and is open source. Check out the cloud­flare-os GitHub repos­i­tory. You can de­ploy it into your own Cloudflare ac­count and use your own Access poli­cies, AI Gateway con­fig­u­ra­tion, data, and in­te­gra­tions.

Our in­ter­nal de­ploy­ment re­flects Cloudflare’s sys­tems, ter­mi­nol­ogy, poli­cies, and ways of work­ing. Yours should re­flect your or­ga­ni­za­tion.

Cloudflare OS is de­signed so you can cus­tomize the in­ter­face, add in­ter­nal Gatekeepers, and build or­ga­ni­za­tion-spe­cific fea­tures with­out chang­ing the core prod­uct.

We are re­leas­ing two repos­i­to­ries: the Cloudflare OS core and an ex­am­ple de­ploy­ment based on how we run it in­ter­nally at Cloudflare. The de­ploy­ment repos­i­tory con­sumes the core with­out patch­ing it, pro­vid­ing a place for con­fig­u­ra­tion, cus­tom UI, in­ter­nal in­te­gra­tions, an­a­lyt­ics, and de­ploy­ment pipelines.

Delivered together with our part­ners

The source code is only the start­ing point. The con­text, skills, work­flows, in­ter­nal sys­tems, and poli­cies are what make Cloudflare OS even more use­ful for your or­ga­ni­za­tion.

Cloudflare’s strate­gic part­ners, Presidio and Happy Cog, will work with you to cus­tomize Cloudflare OS around how your or­ga­ni­za­tion op­er­ates and roll it out across your work­force.

Partners can help you cu­rate shared skills and in­sti­tu­tional con­text, build cus­tom in­ter­faces, con­nect in­ter­nal sys­tems through Gatekeepers and MCP Server Portals, and con­fig­ure se­cu­rity, model, and cost con­trols.

You get your own branded Cloudflare OS, con­nected to your sys­tems, run­ning on Cloudflare, and shaped around how your peo­ple ac­tu­ally work.

Get started

Cloudflare OS is avail­able to­day on GitHub. You can ex­plore the source code, try the demo, or de­ploy it into your own Cloudflare ac­count in a few min­utes us­ing our starter repos­i­tory.

We’re just get­ting started. We’re work­ing on bring­ing Cloudflare OS to the Cloudflare dash­board as a fully man­aged prod­uct, adding con­tain­ers for de­vel­op­ment work­flows, and bring­ing work­spaces into Slack and other chat tools.

If you’re in­ter­ested in talk­ing with our team, we would love to chat. Use this form to reach out!

A Civilian Plane Crashed in New Mexico. Was the Military’s Tech to Blame?

www.wired.com

Jul 30, 2026 6:00 AM

Drone war­fare is mak­ing the skies more dan­ger­ous, even for air­planes far from the bat­tle­field.

ANIMATION: Gabriel Gabriel Garble

This past May, a twin-en­gine Beechcraft King Air mede­vac plane took off from Roswell, New Mexico, and headed west to the town of Ruidoso to pick up a pa­tient. It should­n’t have been a chal­leng­ing flight for the two pi­lots and two nurses aboard. The tem­per­a­ture was 69 de­grees; the sky was clear. The 60-mile jour­ney nor­mally takes a half hour, at most.

But once air­borne, the plane ran into trou­ble. At the White Sands Missile Range that night, US mil­i­tary per­son­nel were con­duct­ing a GPS jam­ming ex­er­cise that left the King Air pi­lots—and any­one else within hun­dreds of miles—un­able to use mod­ern nav­i­ga­tion sys­tems. Forced to re­vert to older tech­nol­ogy, ones that they rarely if ever use, the mede­vac pi­lots got dis­ori­ented and crashed into the side of a moun­tain. There were no sur­vivors.

The ac­ci­dent marked the first time that GPS jam­ming had con­tributed to the crash of a civil­ian plane in the United States. But it was just one of a string of re­cent dis­rup­tions across the world. The skies are more con­tested than ever, whether it’s civil­ian drones wan­der­ing out of the ap­proved zone or US agen­cies get­ting their sig­nals crossed, as hap­pened ear­lier this year when New Mexico and Texas scared the pub­lic by tem­porar­ily clos­ing their air­space. (It turned out that US Customs and Border Patrol were us­ing anti-drone lasers in that area.) The GPS jam­ming ex­er­cise that led to this lat­est crash is not a sin­gu­lar event. In the past year, the US mil­i­tary ap­peared to have sent out no­tices for at least 10 such ex­er­cises. As drone war­fare and elec­tronic war­fare ex­pand, air­lines are in­creas­ingly en­coun­ter­ing nav­i­ga­tion dis­rup­tions hun­dreds of miles be­yond the ac­tual con­flict zone,” says Eliran Almog, CEO of the cy­ber­se­cu­rity firm Cyviation.

It’s worth tak­ing a closer look at what hap­pened last May. While the Ruidoso crash was the first fa­tal ac­ci­dent in the US known to be linked to elec­tronic war­fare, there’s no rea­son to think it will be the last.

Even ab­sent GPS jam­ming, mede­vac is one of the most dan­ger­ous cat­e­gories of civil avi­a­tion. (Kreindler, a law firm spe­cial­iz­ing in air crash lit­i­ga­tion, says that mede­vac flights have an ac­ci­dent rate more sim­i­lar to com­bat fly­ing than to civil avi­a­tion.) Flights are of­ten or­ga­nized on short no­tice, and they fly into airstrips that might be un­fa­mil­iar to the flight crew, and be­cause hu­man lives are at stake, there is an in­cen­tive to fly when weather con­di­tions are mar­ginal.

Some of those fac­tors were at play on the night of May 13. At 11 pm, the crew was no­ti­fied that they had to fly to Ruidoso to pick up a pa­tient and bring them to Albuquerque. (That in­for­ma­tion comes from the pre­lim­i­nary re­port put out by the National Transport Safety Board.) The pi­lots were cap­tain Keelan Clark, aged 30, and first of­fi­cer Ali Kawsara, aged 23. Clark had got­ten his com­mer­cial pi­lot’s li­cense just a year and a half be­fore; he’d been pro­moted from first of­fi­cer to cap­tain the pre­vi­ous month. Kawsara had just two months on the job. He’d only worked cargo jobs be­fore this one.

Both men had demon­strated pro­fi­ciency fly­ing in low-vis­i­bil­ity con­di­tions us­ing what’s called in­stru­ment flight rules, or IFR. There are two ba­sic ways to nav­i­gate in bad weather. Modern cock­pits are equipped with GPS-enabled equip­ment that shows where the plane is on a com­puter screen and por­trays a ma­genta-col­ored line that shows pi­lots where they need to go. This is called RNAV fly­ing; an RNAV ap­proach” brings planes all the way to the thresh­old of a run­way for land­ing in low-vis­i­bil­ity con­di­tions.

This method of fly­ing is much eas­ier than the pre­vi­ous it­er­a­tion. Before GPS be­came widely avail­able in the 2000s, air­lin­ers nav­i­gated us­ing a com­bi­na­tion of mag­netic com­passes, ground-based ra­dio bea­cons, and in­er­tial sys­tems de­rived from old-fash­ioned gy­ro­scopes. Flying by ra­dio bea­cons re­quires pi­lots to form a 3D men­tal map of their lo­ca­tion rel­a­tive to the bea­cons. They have to prac­tice un­til they be­come so ef­fi­cient that they can stay calm un­der pres­sure, lest they panic, lose their sit­u­a­tional aware­ness, and spi­ral out of con­trol. Just ask Kennedy,” says Kenneth Krentsa, a re­tired air­line pi­lot, re­fer­ring to JFK Jr.’s 1999 night­time crash.

Clark and Kawsara took off at eight min­utes to mid­night and ini­tially headed due west, to­ward Ruidoso. The weather was clear, but be­cause the night was nearly moon­less and the area is rural, the only vi­sual ref­er­ences avail­able were the lights of scat­tered set­tle­ments. It’s a black hole out there,” says Juan Browne, an air­line pi­lot who hosts a crash-in­ves­ti­ga­tion pod­cast.

Unable to ori­ent them­selves with­out vi­sual cues, the pi­lots called up Albuquerque Air Route Traffic Control Center—Albuquerque Center, for short—and re­quested per­mis­sion to fly in­stru­ments-only to Ruidoso. The re­quest was ap­proved.

Under nor­mal cir­cum­stances, the flight that fol­lowed would have been un­event­ful. The pi­lots would have fol­lowed the ma­genta line, and the GPS nav­i­ga­tion equip­ment would have lined them up for a smooth land­ing.

But 100 miles to the west, an Air Force Unit called the 746th Test Squadron of the 704th Test Group was hold­ing its an­nual NAVFEST event at the White Sands Missile Range. The event draws to­gether elec­tronic war­fare units from across the armed ser­vices for two weeks of ex­er­cises, in which units test dif­fer­ent tech­nolo­gies for dis­rupt­ing GPS and deal­ing with ad­ver­saries’ dis­rup­tion.

Courtesy of AirNavRadar.com

The event is held at White Sands be­cause it’s among the most re­mote and sparsely set­tled ar­eas of the con­ti­nen­tal United States. (Not co­in­ci­den­tally, the first atomic bomb was det­o­nated there.) But in the run-up to NAVFEST, the Federal Aviation Administration warned air­craft op­er­a­tors that GPS could be af­fected up to 400 miles away be­tween May 12 and May 18.

At mid­night on May 14, eight min­utes af­ter Clark and Kawsara took off from Roswell, they told Albuquerque Center that they’d lost their GPS. Unable to nav­i­gate on their own, they asked that the con­troller give them a head­ing—a mag­netic di­rec­tion to fly in. The con­troller gave them a head­ing to fly west, and then, a minute later, to turn north.

The pi­lots said that they wanted to fly an RNAV ap­proach to Ruidoso. This be­ing ruled out while GPS is jammed, the King Air pi­lots changed their re­quest and asked to use an al­ter­nate form of land­ing sys­tem called Instrument Landing System, or ILS, that does­n’t re­quire GPS re­cep­tion. At 12:05 am, the con­troller as­sured the King Air that they would pro­vide them with vec­tors to guide them in a cou­ple of min­utes.” In the mean­time, they kept fly­ing north.

In ret­ro­spect, tragedy might have been avoided if Albuquerque air traf­fic con­trol had been able to pay closer at­ten­tion to the young pi­lots in the King Air. But tonight they were busy. Three other air­craft also re­ported that they’d lost their GPS and needed help. One was strug­gling to get a bear­ing on a ra­dio bea­con.

While ATC helped other planes, the King Air con­tin­ued north. By 12:08 am, they had over­shot the land­ing pat­tern by 10 miles.

At this point the pi­lots had three op­tions. They could stick to the cur­rent plan and wait for the busy con­troller to give them the next vec­tor to­ward the land­ing. Or, now that GPS was work­ing again, they could ask to switch to the RNAV ap­proach and fly it them­selves. Or they could ditch the in­stru­ment ap­proach al­to­gether and fly what’s called a vi­sual ap­proach. You see a run­way, and you fly to it.

As the King Air flew north, they were high enough to see the lights of Ruidoso’s air­port 31 miles to the south­west. To the pi­lots in the cock­pit of the King Air, a vi­sual ap­proach must have seemed a tan­ta­liz­ing prospect. Why hang around wait­ing for ATC to give them vec­tors, why go through the men­tal ac­ro­bat­ics of try­ing to fig­ure out where they were rel­a­tive to the ILS bea­con? All they had to do was fly to­ward the lights that they could clearly see through their wind­shield.

The King Air called Albuquerque Center and asked to go vi­sual.” The re­quest was granted.

As they turned and de­scended to­ward Ruidoso’s lights, what the pi­lots could­n’t see was the 10,000-foot-high mass of the Capitan Mountains ly­ing across their path. As they drew closer, the dark mass of rock ap­peared to rise up, swip­ing away the lights of the val­ley. This could have cre­ated confusion and a loss of sit­u­a­tional aware­ness,” Browne says. When those lights go out, man, you know you are in big trou­ble.”

The pi­lots slowed their de­scent, even climb­ing a lit­tle, but it was­n’t enough. They kept fly­ing straight to­ward the moun­tain. When you are in that state of mind, you climb as high as you can,” Krentsa says. And you cir­cle. You stay in one place un­til you fig­ure out where you are. You don’t just keep press­ing for­ward, blindly.”

But that’s what the King Air pi­lots did. They flew straight into the ris­ing slope and hit it at full speed. All four oc­cu­pants died in­stantly.

The na­ture of war­fare is chang­ing pro­foundly, and quickly, as drones be­come cheaper, more nu­mer­ous, and more deadly. Hard to spot, and hard to shoot down, they pro­vide an ef­fec­tive way for smaller, less re­sourced na­tions to level the play­ing field against more pow­er­ful ad­ver­saries. Ukraine, nearly over­whelmed by Russia’s con­ven­tional war­fare might at the be­gin­ning of 2022, has rapidly de­vel­oped its drone force to gain what ap­pears to be an up­per hand in the con­flict. And while the US achieved to­tal air su­pe­ri­or­ity over Iran af­ter at­tack­ing the coun­try this February, it has found it­self help­less to stop Iran from us­ing drones and mis­siles to ef­fec­tively shut down traf­fic through the Strait of Hormuz.

To fight back, de­fend­ers can try to ex­ploit a drone’s nav­i­ga­tion. A cheap and sim­ple way for drones to reach their tar­gets is by GPS, which uses ra­dio sig­nals re­ceived from a con­stel­la­tion of satel­lites to cal­cu­late a po­si­tion. When those sig­nals are blocked, an en­e­my’s drones can be ren­dered blind. But the en­emy, too, can take coun­ter­mea­sures. Drone and anti-drone tech­nolo­gies find them­selves in an end­less cat-and-mouse bat­tle, each con­tin­u­ously try­ing to outdo the other. Exercises like NAVFEST of­fer a way for the US mil­i­tary to stay on top of the game.

Civilian GPS has be­come col­lat­eral dam­age, and air travel most of all. Since 2023, planes fly­ing over large swaths of the Middle East, the Black Sea, and the Baltic Sea re­gions have en­dured waves of GPS jam­ming. Airlines have learned to adapt, but a price is still be­ing paid. GPS was a ma­jor boost for air­line safety, and while re­mov­ing it may not in­stantly cause planes to fall from the sky, it re­moves a layer of pro­tec­tion from pas­sen­gers and crew. Add in other stres­sors—a dark night, an in­ex­pe­ri­enced crew, a lack of pro­fi­ciency in the backup tech­nol­ogy—and the sum to­tal is enough to yield dis­as­ter.

There is no ques­tion that in­creased lev­els of GPS jam­ming and spoof­ing around the world pose a safety risk for com­mer­cial avi­a­tion. When alarms go off rou­tinely in the cock­pit, and when pi­lots learn to dis­re­gard key read­ings from their in­stru­ments be­cause the read­ings can’t be trusted, we’re a long way from nor­mal op­er­a­tion,” says Todd Humphreys, a pro­fes­sor of aero­space en­gi­neer­ing at the University of Texas at Austin who has been a lead­ing re­searcher into GPS dis­rup­tion. Air travel is still very safe, but it may be stuck for the next five years or more in a mild-and-in­creas­ing risk sit­u­a­tion as we con­front ever more GPS in­ter­fer­ence within the very-slow-to-adapt avi­a­tion in­dus­try.”

For a few years, US avi­a­tion was spared the dis­rup­tions of anti-drone elec­tronic war­fare. Then it started to hap­pen here, too. In March 2025, air­lin­ers fly­ing into Ronald Reagan National Airport in Washington, DC, re­ceived spu­ri­ous alarms from a col­li­sion-avoid­ance sys­tem, and sev­eral had to abort their land­ings. It later turned out that the Secret Service was test­ing elec­tronic war­fare equip­ment at the vice pres­i­den­t’s res­i­dence. This year, two sep­a­rate in­ci­dents in West Texas in­volv­ing US Army and CBP drone op­er­a­tions led to air­space clo­sures and the dis­rup­tion of com­mer­cial flights.

The avi­a­tion in­dus­try has been slow to grap­ple with the pro­lif­er­a­tion of counter-drone mea­sures and their po­ten­tial ef­fects on flight safety. Airlines and other com­mer­cial op­er­a­tors are still heav­ily re­liant on GPS for nav­i­ga­tion, and other cru­cial tech­nolo­gies, like col­li­sion avoid­ance, are also vul­ner­a­ble. The harder prob­lem with drones is­n’t de­feat­ing them. It’s do­ing it with­out cre­at­ing a sys­tem that neg­a­tively im­pacts civil avi­a­tion,” says Kris Brost, gen­eral man­ager of Robin Radar Systems, a drone de­fense com­pany. Counter-drone tech­nol­ogy has to be sur­gi­cal, not a sledge­ham­mer, be­cause the air­space we’re try­ing to pro­tect is the same air­space the econ­omy runs on.”

Living in a world with drones of both the friendly and un­friendly va­ri­ety is go­ing to take a lot of ad­just­ing. Historically, ma­jor changes in avi­a­tion take place only af­ter crashes that kill a large num­ber of peo­ple. But a suf­fi­ciently mo­ti­vat­ing cat­a­stro­phe may not be far off.

On July 7, a 737 freighter, op­er­ated by a tiny Pakistani cargo air­line called K2 Airways, took off in the late af­ter­noon from Sharjah in the United Arab Emirates and flew east to­ward Karachi with a five-per­son crew. Its route took it just south of the Strait of Hormuz, an area that had been ex­pe­ri­enc­ing in­tense GPS jam­ming due to the US-Iran con­flict. Later, af­ter night­fall, the flight crew called Karachi air traf­fic con­trol and re­ported a navigational sys­tem is­sue,” ac­cord­ing to the Pakistan Civil Aviation Authority. In the three min­utes that fol­lowed, the plane dove 5,000 feet, climbed 6,000 feet, and then plunged 36,000 feet into the ocean in a near-ver­ti­cal dive, killing every­one aboard. It’s too early to know what caused the crash. But it won’t be any sur­prise if the elec­tronic war­fare made an­other pi­lot fly into dark­ness.

Update: 8/5/2026, 8:05 PM EDT: WIRED has up­dated the ar­ti­cle to cite NTSBs pre­lim­i­nary re­port.

Let us know what you think about this ar­ti­cle. Submit a let­ter to the ed­i­tor at [email protected].

Jeff Wise is a New York–based sci­ence jour­nal­ist spe­cial­iz­ing in avi­a­tion, tech­nol­ogy, and psy­chol­ogy who con­tributes fre­quently to New York magazine. A co­host of the pod­cast Find­ing MH370, he was widely seen in the Netflix doc­u­men­tary se­ries MH370: The Plane That Disappeared and ex­ec­u­tive-pro­duced the Showtime doc­u­men­tary fea­ture Gringo: The Dangerous Life of John McAfee. He is the … Read More

DeltaDB

zed.dev

Software is made be­tween com­mits

Be among the first to try DeltaDB, a ver­sion con­trol sys­tem that records the work as it un­folds and keeps every change con­nected to the con­ver­sa­tion that shaped it.

Rewind to any edit

DeltaDB cap­tures every op­er­a­tion in be­tween com­mits and gives each one a sta­ble iden­tity, so you can point to the code at any mo­ment in its evo­lu­tion.

Trace code to con­ver­sa­tion

Every change is linked to the agent con­ver­sa­tion that pro­duced it. From any line of code, find the con­ver­sa­tion. From any mes­sage, jump to the code it touched.

Branch at any mo­ment

DeltaDB vir­tu­al­izes the work­tree, so spin­ning up a new agent branch is ef­fec­tively free. Any point in his­tory is a valid branch point, in­clud­ing mid-run.

Share the thread, not the PR

A team­mate can join while the work is still hap­pen­ing, talk to the agent that did the work, and an­no­tate as they go, with­out wait­ing for you to com­mit and push first.

Just a moment...

www.axios.com

I'm switching my phone from Android to Linux.

runarcn.no

02 Aug, 2026

For the past months or even years, I’ve be­come more and more dis­sat­is­fied with the path Google has taken with the Android Open Source Project (AOSP). Be it the de­pen­dency and in­sane track­ing of Google Play Services, lock­ing down of de­vice trees to hin­der cus­tom ROM de­vel­op­ment, all the AI stuff be­ing put into the sys­tem, or most re­cently (and per­haps worst), re­mov­ing the abil­ity of in­stalling apps per your own wish­ing - it’s be­come sort of like a death by a thou­sand pa­per­cuts” sit­u­a­tion. While the AOSP is­n’t ex­actly dead (yet), it kind of feels like it’s just a ques­tion of time. Therefore, I’ve de­cided to jump ship. I’m in­stalling linux on my phone.

The mo­bile linux space is­n’t ex­actly look­ing good right now, but it holds some promise. Ubuntu Touch man­aged to sur­vive be­ing aban­doned by Canonical, and post­mar­ket is still see­ing good de­vel­op­ment even if not hav­ing great hard­ware ada­p­a­tion at the mo­ment. SailfishOS (even if it has pro­pri­etary parts) has also made great progress with strong sup­port for an­droid apps on of­fi­cial de­vices and what seems to be a real good launch of their Jolla 2.

Personally I’m lucky enough to own a Fairphone 4 that just so hap­pens to sup­port all of these. After a bit back and forth I’ve ul­ti­mately ended up with SailfishOS. Overall it’s a good sys­tem - I like the ges­ture based nav­i­ga­tion and its ap­pli­ca­tion frame­work is beau­ti­ful. It’s also cool how it’s very linux. If I have an is­sue or some­thing I want to tin­ker with I can just ssh over from my PC ei­ther wire­lessly or via USB. It’s not with­out it’s is­sues though. For some rea­son it ships hor­ri­bly out­dated ver­sions of python and glibc mak­ing some things more dif­fi­cult than needed, and way­droid and GPS is bro­ken on the fair­phone port - worth not­ing an un­of­fi­cial port. Many of the com­mu­nity-built apps are also ei­ther in-part or com­pletely slop-coded, such as a what­sapp client that I (sadly) am de­pen­dant on.

Ubuntu Touch is also an op­tion, but is­n’t with­out it’s own bag of is­sues. Waydroid runs there which is a huge plus, but no­ti­fi­ca­tions and clip­board does­n’t sync across mak­ing it re­ally hard to use ie. Bitwarden (Sailfish has an un­of­fi­cial port which is al­most flaw­less). The de­fault and na­tive apps are pretty lack­lus­ter com­pared to both Android and Sailfish. Among other things I was un­able to fig­ure out how to block phone num­bers, a highly needed fea­ture for any­one that has VIVO as their cel­lu­lar provider. I’m also not a big fan of how the UI works with app nav­i­ga­tion and with the top bar”/“​drop down menu”. If you want to know more about it, The Linux Experiments has some good videos on both youtube and peer­tube.

Sadly, I won’t be able to aban­don an­droid com­pletely yet. The no way­droid means that I can’t ac­cess some apps that I need be it for on­line ver­i­fi­ca­tion for log­ging onto bank and gov­ern­ment ser­vices in Norway or apps I need for my own se­cu­rity here in Brazil such as Uber. Luckily I’ve had a Galaxy A17 ly­ing around as a backup phone for half a year now which I can carry around for these ex­act ser­vices. I just open up a wifi hotspot from my FP4, do the ex­act tasks, and close it off again. Other than that, I will try to avoid us­ing it as much as phys­i­cally pos­si­ble.

As this pro­gresses I will take note of what works and what does­n’t be­fore mak­ing ei­ther a proper writeup or video in a while. It’s not gonna be easy, but hope­fully it will be worth it.

Oh, and slight spoil­ers for what my ex­pe­ri­ence with Sailfish is af­ter some us­age, but I might go ahead and buy a Jolla Phone 2 when I re­turn to Norway.

Reply via email

#free soft­ware

#tech

How Castform + Neon Beats Frontier Models on Price and Efficiency

neon.com

Most teams’ best train­ing data is just sit­ting in their data­bases. The prob­lem is that turn­ing raw data into some­thing us­able is hard, and let­ting agents read, search, and mu­tate data cheaply at scale re­quires ad­vanced in­fra. Pointing Castform at Neon skips both.”Ying Hang Seah, co­founder, Castform

Most teams’ best train­ing data is just sit­ting in their data­bases. The prob­lem is that turn­ing raw data into some­thing us­able is hard, and let­ting agents read, search, and mu­tate data cheaply at scale re­quires ad­vanced in­fra. Pointing Castform at Neon skips both.”

A good agent” needs to be strong in 2 ar­eas:

Context: can we pro­vide the tools to find the right data?

Model: can the model de­cide what to search for?

Neon (Lakebase Postgres) and their new Search ex­ten­sions solve the first; Castform solves the sec­ond.

In ~2022, the in­dus­try was go­ing all in on em­bed­ding search. Every data­base provider added one, and pgvec­tor was Neon’s most down­loaded ex­ten­sion. To pro­vide con­text to LLMs, en­gi­neers hand­crafted RAG pipelines, which in essence, is some form of em­bed­ding sim­i­lar­ity search.

In ~2025, agents started to gain more trac­tion. Developers started cre­at­ing multi-hop search work­flows, de­com­pos­ing big prob­lems into smaller ones. Retrieval has shifted from the one-shot search sys­tems to agen­tic re­trieval. Instead of is­su­ing a sin­gle query, mod­els plan and search mul­ti­ple times in a loop. Every loop it­er­a­tion meant an­other call to the fron­tier model, in­creas­ing the over­all cost and la­tency per user re­quest.

Concretely, a typ­i­cal multi-turn search re­quest with gpt-5.6-sol takes >10s and costs ~$0.03 end-to-end, mak­ing it pro­hib­i­tively slow and ex­pen­sive.

Meanwhile, small open-weights mod­els are 100x cheaper. But, out of the box, their ca­pa­bil­i­ties lag be­hind closed api mod­els. RL post-train­ing helps bridge this gap. On spe­cific tasks like search, post-trained open-source mod­els can match & beat fron­tier mod­els while cost­ing or­ders of mag­ni­tude less per re­quest.

That is why we built Castform: to en­able de­vel­op­ers to RL post-train mod­els with­out hav­ing to deal with ma­chine learn­ing & gpu in­ter­nals. The goal’s to make post-train­ing as ap­proach­able as prompt en­gi­neer­ing.

Castform’s pipeline runs against Neon via Lakebase Search:

To per­form RL post-train­ing ef­fec­tively, you need a task (e.g. an­swer a user’s ques­tion), the en­vi­ron­ment for the agent to run in (e.g. a search tool for your cor­pus) and a re­ward func­tion (e.g. is the an­swer cor­rect?).

With all 3 pieces in place, the RL post-train­ing is a loop of trial and er­ror: the model at­tempts the task given the tools, the re­ward func­tion scores the at­tempt, and the feed­back sig­nal guides the model on how to hill-climb its way to op­ti­mal per­for­mance.

Yet, most com­pa­nies do not have a clean dataset of tasks and re­ward func­tions ready for post-train­ing.

Enterprises do have a large set of pro­pri­etary data:

in­ter­nal doc­u­men­ta­tion

prod­uct records

sup­port ar­ti­cles

cus­tomer in­ter­ac­tions

wikis

op­er­a­tional data­bases

This data con­tains the knowl­edge an agent needs, but turn­ing it into an ef­fec­tive train­ing dataset nor­mally re­quires sub­stan­tial data en­gi­neer­ing and man­ual la­bel­ing.

That leads many teams to dis­miss post-train­ing for one of two rea­sons:

We don’t have the train­ing data.”

Fine-tuning is too dif­fi­cult and re­quires in­fra­struc­ture we don’t have.”

Castform ad­dresses both. It turns an ex­ist­ing cor­pus into train­ing tasks, then man­ages the RL loop needed to teach an open-source model how to use that data ef­fec­tively.

With Castform, you can turn your com­pany knowl­edge base into a model:

Document (from your data): Trains booked through Navan will be paid by GitLab travel card. Train rides must be stan­dard cabin class with 14 day book­ing lead time

Ground truth (inferred from your data): Train rides must be stan­dard cabin class with a 14 day book­ing lead time.

Question (synthetically gen­er­ated): When book­ing a rail trip in Navan, what are the rules for how early I need to re­serve it and which seat­ing level I’m ex­pected to choose?

With the gen­er­ated ques­tion-an­swer dataset, Castform lets you scaf­fold the train­ing run by spec­i­fy­ing the tools the agent has ac­cess to and a re­ward func­tion.

The re­ward func­tion spec­i­fies what you want your model to get good at. In our case, we want it to re­trieve the cor­rect chunks, cite the right sources along with pro­vid­ing the right fi­nal an­swer.

def run_­tool(tool, tool_args): ”″Single tool: hy­brid search over Lakebase.“”″ if tool == search”: query = tool_args[“query”] bm25 = neon.lake­base_­text(query, k) vec­tor = neon.lake­base_vec­tor(query, k) re­turn rrf_merge(bm25, vec­tor, k)

def re­ward(trace, ground_truth): ”″Grade a trace against the ground-truth an­swer.“”″ an­swer = parse_­trace(trace) re­trieval = … # did it re­trieve the right source ci­ta­tion = … # did it cite the right chunk cor­rect­ness = … # did it land on the right an­swer re­turn re­trieval + ci­ta­tion + cor­rect­ness

See a com­pre­hen­sive code ex­am­ple here.

Castform gives you full ob­serv­abil­ity into your RL run. You can mon­i­tor your re­ward climb with each step, but more im­por­tantly you can drop into in­di­vid­ual tasks/​prompts to watch how the model per­forms qual­i­ta­tively, al­low­ing you to de­bug prob­lems such as bro­ken tools or re­ward hack­ing.

For more de­tails on how to mon­i­tor your train­ing runs, you can check out the Castform blog here. You can also check out our ex­am­ple train­ing run here.

During train­ing, the agent re­peat­edly calls Lakebase Search un­til it has enough con­text to an­swer. Across thou­sands of par­al­lel roll­outs, each po­ten­tially mak­ing dozens of calls, this cre­ates a highly bursty work­load.

Neon’s dy­namic com­pute scal­ing ab­sorbs these peaks with­out re­quir­ing Castform to pro­vi­sion for max­i­mum ca­pac­ity around the clock. Training runs get low-la­tency search when de­mand spikes, while com­pute scales down dur­ing idle pe­ri­ods.

This in­fra­struc­ture be­comes even more valu­able as agents move be­yond search and be­gin mod­i­fy­ing data. Training state­ful agents re­quires iso­lated en­vi­ron­ments that can be cre­ated and re­set cheaply, pre­vent­ing one roll­out’s ac­tions from af­fect­ing an­other or touch­ing pro­duc­tion.

Neon branch­ing can give each roll­out an iso­lated data­base state, while time-travel queries make it pos­si­ble to re­con­struct and in­spect the state an agent en­coun­tered. Combined with au­toscal­ing and scale-to-zero, this cre­ates a path to­ward train­ing thou­sands of state­ful agent roll­outs with­out main­tain­ing thou­sands of con­tin­u­ously run­ning en­vi­ron­ments.

Castform makes it easy for any de­vel­oper to post-train open-source mod­els to be cheaper, faster, bet­ter than the fron­tier. Post-train your first model to­day at cast­form.com.

Cops Used Flock to Track a Man Across State Lines to Create Pretext to Search His Car for Weed

www.404media.co

Police in Wisconsin used Flock to de­ter­mine that a man travels to Michigan fre­quently,” where mar­i­juana is le­gal, then back to Wisconsin, where it is il­le­gal. They then used his travel across state lines as tracked by Flock as part of the prob­a­ble cause jus­ti­fi­ca­tion to search his car for weed; he was even­tu­ally ar­rested on mar­i­juana pos­ses­sion charges, ac­cord­ing to court records re­viewed by 404 Media.

The searches came to light in a Wisconsin crim­i­nal com­plaint against Edward Abrams-Phillips, who was wanted for bail jump­ing on do­mes­tic vi­o­lence charges. But the crim­i­nal com­plaint makes clear that be­yond the bail jump­ing and do­mes­tic vi­o­lence charges, po­lice specif­i­cally stud­ied Abrams-Phillips’ in­ter­state travel to cre­ate the pre­text for search­ing his car for mar­i­juana. The bail jump­ing charge was dis­missed; Abrams-Phillips was found guilty only of weed pos­ses­sion in the case, ac­cord­ing to the court records.

The com­plaint ex­plains that Abrams-Phillips was tracked via Flock’s net­work over the course of the day to de­ter­mine that he drove from Wisconsin to Michigan, a known source state for mar­i­juana as it is le­gal there,” the com­plaint states, adding that pre­vi­ous Flock hits in­di­cated that he travels to Michigan fre­quently.” Police note that, us­ing Flock, they were able to track Abrams-Phillips dri­ving from Wisconsin to Michigan, then back to Wisconsin over the course of sev­eral hours, where he was pulled over and ar­rested. The Flock searches and ar­rests hap­pened in April 2025.

The ve­hi­cle was ob­served hit­ting flock on sev­eral oc­ca­sions to in­clude 41 north­bound from Brown Rd, 41NB and County Line in Marinette [Wisconsin], and 41 NB on Bridge St. go­ing into Michigan. Based on prior flock hits, the ve­hi­cle trav­els to Michigan fre­quently which is a known source State for Marijuana as it is le­gal there,” the charg­ing doc­u­ment notes. Around 3:56 p.m., the ve­hi­cle was seen on Flock head­ing south­bound on in­ter­state 41 to­wards Green Bay [Wisconsin]. Deputies made a co­or­di­nated ef­fort to in­ter­cept the ve­hi­cle on 41 from Brown Rd. Deputy Kowalski ini­ti­ated a traf­fic stop on the ve­hi­cle as the dri­ver matched the de­scrip­tion of Edward.”

This post is for paid mem­bers only

Become a paid mem­ber for un­lim­ited ad-free ac­cess to ar­ti­cles, bonus pod­cast con­tent, and more.

Subscribe

Sign up for free ac­cess to this post

Free mem­bers get ac­cess to posts like this one along with an email round-up of our week’s sto­ries.

Subscribe

Already have an ac­count? Sign in

nytimes.com

www.nytimes.com

Please en­able JS and dis­able any ad blocker

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.