10 interesting stories served every morning and every evening.

Gemini 3.7 Flash

ai.google.dev

Gemini 3.7 Flash is the next it­er­a­tion in the Gemini 3 se­ries of highly-ca­pa­ble, na­tively mul­ti­modal, rea­son­ing mod­els.

Documentation

Visit the Latest model page for full cov­er­age of fea­tures and ca­pa­bil­i­ties.

gem­ini-3.7-flash

Inputs

Text, Image, Video, Audio, and PDF

Output

Text

Input to­ken limit

1,048,576

Output to­ken limit

65,536

Audio gen­er­a­tion

Not sup­ported

Caching

Supported

Code ex­e­cu­tion

Supported

Computer use

Supported (Preview)

File search

Supported

Function call­ing

Supported

Grounding with Google Maps

Supported

Image gen­er­a­tion

Not sup­ported

Live API

Not sup­ported

Search ground­ing

Supported

Structured out­puts

Supported

Thinking

Supported (low, medium, high)

Note: min­i­mal is not sup­ported and re­turns an er­ror.

URL con­text

Supported

Batch API

Supported

Flex in­fer­ence

Supported

Priority in­fer­ence

Supported

Stable: gem­ini-3.7-flash

Except as oth­er­wise noted, the con­tent of this page is li­censed un­der the Creative Commons Attribution 4.0 License, and code sam­ples are li­censed un­der the Apache 2.0 License. For de­tails, see the Google Developers Site Policies. Java is a reg­is­tered trade­mark of Oracle and/​or its af­fil­i­ates.

Last up­dated 2026 – 08-13 UTC.

Introducing Gemini 3.7 Flash

blog.google

Aug 13, 2026

|

Our most in­tel­li­gent work­horse model yet for cod­ing and agents.

Your browser does not sup­port the au­dio el­e­ment.

Listen to ar­ti­cle

[[duration]] min­utes

This con­tent is gen­er­ated by Google AI. Generative AI is ex­per­i­men­tal

Today, we’re build­ing on the progress of our widely used Flash se­ries by in­tro­duc­ing Gemini 3.7 Flash, our most in­tel­li­gent work­horse model yet for cod­ing and agents.

This re­lease comes just three weeks af­ter Gemini 3.6 Flash, and is a di­rect re­sult of de­vel­oper feed­back and al­go­rith­mic in­no­va­tions that we look for­ward to bring­ing to fu­ture mod­els. 3.7 Flash de­liv­ers sub­stan­tial im­prove­ments across soft­ware en­gi­neer­ing, knowl­edge work, and web de­vel­op­ment work­flows — with an in­tro­duc­tory price of half the orig­i­nal 3.6 Flash cost per mil­lion to­kens.

Better in­tel­li­gence for com­plex work­flows

3.7 Flash shows strong gains over 3.6 Flash in cod­ing tasks like de­bug­ging and is­sue res­o­lu­tion. It also achieves higher first-pass code ac­cu­racy and has im­proved per­for­mance in gen­er­at­ing pro­duc­tion-ready code as seen in FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%).

In web de­vel­op­ment, 3.7 Flash gen­er­ates more func­tional lay­outs and fea­ture-com­plete apps in fewer prompts. For UI gen­er­a­tion, the model shows high de­sign ad­her­ence and par­ity based on a ref­er­ence in­put, whether it’s a screen­shot, an im­age, or a full de­sign sys­tem. It out­per­forms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowl­edge-dense fields like fi­nance, law, and bio­sciences, 3.7 Flash de­liv­ers im­proved rea­son­ing and ac­cu­racy. It sig­nif­i­cantly out­per­forms 3.6 Flash on the GDP.pdf bench­mark (34.0% vs 22.0%), an eval for test­ing a mod­el’s abil­ity to process com­plex doc­u­ments. It also sur­passes 3.6 Flash in AutomationBench, demon­strat­ing it can more ef­fec­tively com­plete real-world busi­ness work­flows (30.4% vs 17.0%).

Better de­vel­oper ex­pe­ri­ence and price

Gemini 3.7 Flash de­liv­ers a no­tice­ably im­proved de­vel­oper ex­pe­ri­ence over 3.6 Flash. It bet­ter adapts to road­blocks, clar­i­fies in­tent when needed, and fol­lows in­struc­tions with greater fi­delity. It thinks more dili­gently, putting in more ef­fort into multi-step plan­ning and tool calls. A more dis­ci­plined ex­e­cu­tion means less man­ual over­sight and fewer re­tries across en­gi­neer­ing work­flows.

3.7 Flash is avail­able through the end of the year at an in­tro­duc­tory price

1

of $0.75/1M in­put to­kens and $3.75/1M out­put to­kens. This price com­bined with the en­hanced model per­for­mance en­ables de­vel­op­ers and cus­tomers to scale pro­duc­tion-ready agents cost ef­fec­tively.

Early cus­tomer feed­back is high­light­ing 3.7 Flash’s per­for­mance and pre­ci­sion, achiev­ing re­sults that are sig­nif­i­cantly bet­ter than 3.6 Flash at a low cost.

Improving Gemini Spark with 3.7 Flash

Gemini Spark, avail­able to Google AI Pro and Ultra sub­scribers in over 160 coun­tries, will be us­ing Gemini 3.7 Flash start­ing to­day. We launched Spark at I/O as your per­sonal AI agent that runs 24/7, tak­ing ac­tion on your be­half while un­der your di­rec­tion. This model up­date makes Spark more ef­fi­cient for knowl­edge work with im­proved tool use for Google Workspace apps, de­liv­er­ing im­proved ac­cu­racy and out­put qual­ity for com­plex, multi-skill work­flows.

With 3.7 Flash, Gemini Spark can turn ideas into ac­tion more ef­fi­ciently by con­sol­i­dat­ing files, draft­ing emails, and up­dat­ing sta­tus doc­u­ments.

Built with safety in mind

We con­tin­u­ally work to im­prove the cov­er­age and ro­bust­ness of Frontier Safety safe­guards. Gemini 3.7 Flash is ship­ping with up­dated safe­guards against mis­use in the do­mains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cy­ber of­fense, while en­abling ben­e­fi­cial use cases, in ac­cor­dance with our ap­proach to biore­silience and our cy­ber pro­gram.

For more in­for­ma­tion, see the 3.7 Flash model card.

Try it to­day

Developers: Explore agent-first work­flows in Google Antigravity or start build­ing to­day in the Gemini API via Google AI Studio and Android Studio. Get started with our de­vel­oper guide.

Enterprises: Access 3.7 Flash in Gemini Enterprise Agent Platform and the Gemini Enterprise app.

Individuals: Available via Spark, your 24/7 per­sonal agent in the Gemini app for Google AI Pro and Ultra sub­scribers in sup­ported coun­tries.

Detailed bench­marks

Get the lat­est news from Google in your in­box

Sign up for our newslet­ters with prod­uct up­dates, event in­for­ma­tion, spe­cial of­fers, and more.

Your in­for­ma­tion will be used in ac­cor­dance with Google’s pri­vacy pol­icy. You may opt out at any time.

z.ai

DeepSeek Harness developer preview: Everything is a plugin

deepseek.com

Agent = Model + Harness

Harnesskeeps agents work­ing in real-world en­vi­ron­ments

The model is the soul of an agent.

A har­ness lets an agent un­der­stand its en­vi­ron­ment, use tools, and keep work­ing in real-world set­tings.

Cordis ker­nel

The Cordis ker­nel man­ages plu­gin mount­ing, un­mount­ing, and de­pen­den­cies. Agent ca­pa­bil­i­ties live in the plu­g­ins.

Capabilities as plu­g­ins

Plugins pro­vide every agent ca­pa­bil­ity, in­clud­ing mod­els, tools, skills, ses­sions, sand­boxes, stor­age, loops, sched­ul­ing, and the UI. Cordis ser­vices and events let the plu­g­ins work to­gether.

Compose with con­fig­u­ra­tion

Developers can se­lect, swap, or ex­tend any ca­pa­bil­ity in con­fig­u­ra­tion with­out chang­ing the DeepSeek Harness source code.

Customize your DeepSeek Harness

Get started

Try it now or in­stall from source

Quick start

Install Node.js, then launch the Web UI with npx.

$ npx @deepseek-ai/dsh web

Install from source

Clone the full source and fol­low the setup in­struc­tions in the repos­i­tory.

$ git clone https://​github.com/​deepseek-ai/​deepseek-har­ness

Accelerating GPT-5.6 Sol Ultrafast with OpenAI

www.cerebras.ai

Today, Cerebras and OpenAI are shar­ing an early look at Ultrafast Mode, a new ser­vice tier launch­ing first in the OpenAI API and pow­ered by Cerebras. Ultrafast is avail­able ini­tially to a se­lect group of cus­tomers, with ac­cess ex­pand­ing over time. Cerebras pow­ers GPT-5.6 Sol on Ultrafast mode, de­liv­er­ing up to 750 out­put to­kens per sec­ond and with­out any qual­ity com­pro­mise, al­low­ing Sol Ultrafast to ac­cel­er­ate your most time-sen­si­tive, mis­sion-crit­i­cal work.

Frontier Intelligence at Unprecedented Speed

AI builders have al­ways needed to choose be­tween speed and in­tel­li­gence. As mod­els scale up in size and in­tel­li­gence, they in­cur higher com­pu­ta­tional and data move­ment costs, slow­ing down re­sponse times. Users of­ten need to wait for high-qual­ity re­sults or ac­cept in­fe­rior re­sults within a shorter time­frame.

GPT-5.6 Sol Ultrafast re­solves this trade­off, bring­ing fron­tier in­tel­li­gence to prod­ucts and work­flows where every sec­ond mat­ters. Compared with out­put speeds re­ported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

At Cerebras, we put Ultrafast to the test by run­ning it head-to-head with pop­u­lar mod­els on Humanity’s Last Exam. HLE is a chal­leng­ing model bench­mark that con­sists of 2,500 ques­tions typ­i­cally an­swer­able only by those hold­ing PhDs in fields such as chem­istry, eco­nom­ics, and lit­er­a­ture.

In our eval­u­a­tions, GPT-5.6 Sol on Ultrafast mode an­swered all 2,500 HLE ques­tions in 11 hours and 11 min­utes. Claude Fable 5 needed 78 hours and 27 min­utes, more than three days of con­tin­u­ous com­pute, to ar­rive at the same con­clu­sions. In other words, Ultrafast worked through the fron­tier of hu­man knowl­edge in a sin­gle work­ing day, achiev­ing com­pa­ra­ble ac­cu­racy nearly faster.

Benchmarking was per­formed by Cerebras us­ing GPT 5.6 Sol Ultrafast with Codex on xhigh rea­son­ing on July 10 and Claude Fable 5 with Claude Code on xhigh rea­son­ing on July 13 – 15.

As model ca­pa­bil­i­ties con­tinue to ad­vance, the range of ap­pli­ca­tions for fast in­fer­ence ex­pands. GPT-5.6 Sol is OpenAI’s best model yet for le­gal briefs, fi­nan­cial mod­els, and en­gi­neer­ing re­ports. On GDP-Val, a bench­mark for eco­nom­i­cally valu­able knowl­edge work tasks, Ultrafast de­liv­ered a 5.6x end-to-end speedup with no qual­ity degra­da­tion, show­ing how faster in­fer­ence can ac­cel­er­ate eco­nom­i­cally valu­able work.

Benchmarking was per­formed by Cerebras on July 31 2026 us­ing GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium rea­son­ing within Codex.

High-Speed Intelligence Powers High-Stakes Work

Faster in­tel­li­gence changes what’s pos­si­ble for in­di­vid­u­als and or­ga­ni­za­tions. With Ultrafast, you can now put agents on the crit­i­cal path of prob­lems where every sec­ond counts.

With GPT-5.6 Sol Ultrafast, Cerebras en­ables AI that keeps up with how you think, code, and col­lab­o­rate. We’re ex­cited to see how work­flows and ap­pli­ca­tions are trans­formed by Ultrafast in­fer­ence.”

Rohan Varma

Product at OpenAI

Ultrafast is a per­sis­tent edge for or­ga­ni­za­tions us­ing fron­tier AI to quickly re­spond to in­com­ing in­for­ma­tion. Companies op­er­at­ing web ser­vices can lever­age Ultrafast to root-cause and ad­dress pro­duc­tion out­ages, pre­serv­ing cus­tomer trust, pre­vent­ing lost rev­enue, and sav­ing down­time min­utes against their SLAs. And in ad­ver­sar­ial, high stakes cy­ber­at­tacks, Ultrafast is an in­valu­able tool for se­cu­rity teams who must quickly de­tect and re­spond to bad ac­tors to con­tain cat­a­strophic losses.

More broadly, Ultrafast en­ables en­tirely new modes of work­ing with agents, it de­liv­ers real-time in­sights and up­dates, so you don’t have to con­text-switch across mul­ti­ple par­al­lel ses­sions to get the most out of your agents.

Whereas for­merly I might have to wait a cou­ple min­utes for a task to fin­ish, it now fin­ishes for me be­fore I even have the op­por­tu­nity to con­text-switch. It makes me way more pro­duc­tive.”

Jeffrey Wang

OpenAI Researcher

With Ultrafast, re­searchers and en­gi­neers can re­serve their at­ten­tion for go­ing deep on se­lect prob­lems that mat­ter most, while con­tin­u­ing to use Standard pro­cess­ing for par­al­leliz­ing com­mod­ity tasks. Cerebras is ex­cited to power the next wave of AI in­no­va­tion, rais­ing the ceil­ing for what in­di­vid­u­als and or­ga­ni­za­tions can ac­com­plish with re­spon­sive AI.

Breakneck Speed is Enabled by Breakthrough Innovation

GPT-5.6 Sol on Ultrafast mode is pow­ered by Cerebras’ rev­o­lu­tion­ary Wafer-Scale Engine ar­chi­tec­ture, pur­pose-built for fron­tier AI work­loads. Fast fron­tier in­fer­ence is a data move­ment prob­lem: on GPUs, in­fer­ence on large mod­els is bot­tle­necked by mem­ory band­width, as model weights must be re­peat­edly trans­ferred be­tween on-chip mem­ory and off-chip stor­age to gen­er­ate suc­ces­sive to­kens within a model re­sponse.

Cerebras takes a con­trar­ian ap­proach to elim­i­nat­ing this in­ef­fi­cient data move­ment: we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and to­kens flow un­in­ter­rupted through model lay­ers pipelined across wafers. This tech­ni­cal ap­proach scales smoothly with model size, paving the way for a con­tin­ued speed ad­van­tage on fu­ture fron­tier mod­els.

Ultrafast: Now in Limited Preview

GPT-5.6 Sol on Ultrafast mode is avail­able in a lim­ited pre­view to­day to a se­lect group of cus­tomers. Access will ex­pand as ca­pac­ity grows. Sign up for up­dates.

Understanding is the new bottleneck

www.geoffreylitt.com

Hot take: I think it’s still im­por­tant to un­der­stand the code that our agents write!

In this talk I’ll ex­plain why that’s the case, and show some ideas for how to ef­fi­ciently un­der­stand code. Alright, let’s dive in.

Agents are writ­ing more and more code for us, and we all know it’s get­ting harder to keep up.

But the good news is: there are many ways to un­der­stand code! Reading diffs line by line is not the only way.

Most of this talk will be about tech­niques I have found help­ful to un­der­stand sys­tems my agents are build­ing:

Code ex­plainer docs

Quizzes to check my un­der­stand­ing

Micro-worlds that I can play with to un­der­stand the sys­tem

But first we have to ask a more ba­sic ques­tion…

Why un­der­stand?

Why? Why un­der­stand?

Aren’t we sup­posed to be tak­ing our­selves out of the loop now, and let­ting the agents loop them­selves? As the agents get smarter, does­n’t it be­come less im­por­tant for us to be in the de­tails?

I think many peo­ple — even those who are pro-un­der­stand­ing — have a slightly in­cor­rect an­swer to this ques­tion!

One pos­si­ble an­swer: we un­der­stand to ver­ify. We check the agen­t’s work, we see if it’s cor­rect.

Correct can mean many things: does it match the spec, is it well ar­chi­tected… but it’s fun­da­men­tally a thumbs-up / thumbs-down ques­tion.

Here’s the thing: the agents are get­ting bet­ter and bet­ter at ver­i­fy­ing their own work. And this is good! I like it when my agent does­n’t make mis­takes.

But hmm. Where does that leave us hu­mans?

That’s where an­other an­swer comes in: we can un­der­stand to par­tic­i­pate.

You can learn what the agent is do­ing to make sure you can be an ac­tive par­tic­i­pant in the cre­ative process. Here’s why this mat­ters…

It’s never just one loop! A pro­ject is many, many loops with the agent.

And the un­der­stand­ing you have of the sys­tem is part of your abil­ity to come up with the next idea to evolve it.

You need a rich set of con­cepts in your mind to think cre­atively and flu­ently about how to move some­thing for­ward. If you’re lack­ing that flu­ency, your abil­ity to par­tic­i­pate in the pro­ject is mean­ing­fully lim­ited.

By the way, this re­lates closely to the idea of cog­ni­tive debt, pop­u­lar­ized by Margaret Storey and Simon Willison.

It’s like tech debt: you can get away with not un­der­stand­ing what’s go­ing on in the short term, but it’ll bite you even­tu­ally.

OK, so fine, un­der­stand­ing mat­ters.

But this raises the next ques­tion: how? How do we build this hu­man un­der­stand­ing when we’re work­ing with AI and mov­ing fast?

Well, turns out this is not the first time any­one has ever thought about how to com­mu­ni­cate un­der­stand­ing. I think we can look to ed­u­ca­tion as an in­spi­ra­tion. Can we steal the best ideas ever in­vented for ed­u­ca­tion and ap­ply them to this prob­lem?

Technique 1: Explanations

Today I want to share three tech­niques that show how we can at­tempt this.

First: ex­pla­na­tions. What makes a good ex­pla­na­tion?

Whenever an agent fin­ishes some work, it’s an op­por­tu­nity for an ex­pla­na­tion — an ar­ti­fact.

Most naively, we can read a code diff: the raw ma­te­r­ial that changed.

But what if we ask:

What would the best ex­pla­na­tion be? If you had a team — hu­man or AI — that re­ally sweat the de­tails of ex­plain­ing some­thing well to you, how would that feel?

Here’s one an­swer. I made a skill called /explain-diff, which I use every day and many cowork­ers have found valu­able.

It out­puts thought­fully struc­tured code ex­plain­ers as HTML, mark­down, or Notion docs. Notion is a good place for col­lab­o­rat­ing on and dis­cussing these ex­plain­ers as a team. (Disclaimer: I work at Notion so I’m bi­ased.)

Let’s see what’s in one of these ex­plain­ers, us­ing an ex­am­ple of edit­ing the per­spec­tive of a video game.

First prin­ci­ple: teach me back­ground info!

Before we even get to what changed, help me un­der­stand what was al­ready there. In this case, teach me about the game en­gine.

Second prin­ci­ple: in­tu­ition be­fore de­tails.

Before any code, it states the goal — make the gar­den feel three-di­men­sional with 2D draw­ing tricks” — and ex­plains re­lated con­cepts, like what iso­met­ric pro­jec­tion is.

All of this builds my in­tu­ition for the essence of the change. It’s catch­ing me up as the hu­man so I can be an equal par­tic­i­pant in un­der­stand­ing.

You can also build in­tu­ition with in­ter­ac­tive fig­ures.

Here I’m un­der­stand­ing the iso­met­ric per­spec­tive by drag­ging rocks around the gar­den and watch­ing their co­or­di­nates move.

(This is us­ing a new fea­ture Notion just shipped: you can now em­bed in­ter­ac­tive HTML in­side pages.)

We fi­nally get to the code. But a typ­i­cal diff is a pile of files edited in al­pha­bet­i­cal or­der with no ex­pla­na­tion.

A literate diff” as I call it is struc­tured as prose — walk­ing through the changes in a sen­si­ble or­der, with sur­round­ing ex­pla­na­tion and em­bed­ded code snip­pets. Faster to re­view than a raw diff.

The end re­sult of all of this is a nice ex­plainer packet. I still read the code diff but I al­ways read this first.

Sometimes I’ll print these out and take them to the café — less dis­tract­ing.

It’s beau­ti­fully ironic: AI turns an in­ter­ac­tive ac­tiv­ity into a sta­tic pa­per re­port I can fo­cus on deeply :)

I do some­thing sim­i­lar with my code ex­plain­ers now. At the bot­tom of an ex­plainer there’s an in­ter­ac­tive quiz — five ques­tions about the change — and I try to an­swer them.

My rule: I won’t send code to oth­ers un­til I can pass the quiz, and I do the same when re­view­ing oth­ers’ code.

A quiz is a speed reg­u­la­tor. Working with AI, it’s easy for the loop to run faster than the speed of hu­man un­der­stand­ing.

The quiz is a coun­ter­bal­anc­ing force: I me­chan­i­cally ask do I ac­tu­ally un­der­stand?” so that I can re­main a full cre­ative par­tic­i­pant.

Technique 2: Micro-worlds

Next idea: mi­cro-worlds. This one’s in­spired by the vi­sion­ary ed­u­ca­tor Seymour Papert.

Papert had this beau­ti­ful idea he called liv­ing in Mathland: if you want to learn math, live in Mathland — just like if you want to learn French, you go live in France. Could we build an en­vi­ron­ment where chil­dren learn math nat­u­rally, as a con­se­quence of their cu­rios­ity?

So how do we ap­ply that to code? Can we make worlds you in­habit and nat­u­rally in­tuit how the sys­tem works and how it’s chang­ing?

Last year I was cod­ing a Prolog in­ter­preter and strug­gling to in­tuit what was hap­pen­ing in­side.

I worked with an agent to build this de­bug­ger, which let me step through the ex­e­cu­tion of my logic lan­guage — scrub through time, see what’s on the stack and which rules are eval­u­ated at each step. I could even leave com­ments for my­self (“nice, we cor­rectly ap­plied that rule”).

There’s a big dif­fer­ence be­tween mak­ing a tool for me to de­bug and let­ting the agent de­bug — do­ing it my­self is how I de­velop un­der­stand­ing along the way.

Another ex­am­ple. I was mi­grat­ing my per­sonal web­site from one frame­work to an­other, and Claude wrote a script that did it. But it was very hard to re­view: I was­n’t fa­mil­iar with the new frame­work, and all I could say was I guess that looks about right.”

So I asked Claude to make me a video game — a com­mand cen­ter where I do the port my­self, step by step, watch­ing the vis­i­ble ef­fects and the file tree evolve. It pro­duced a UI where I click but­tons to run the port step by step, with my old site and new site run­ning side by side.

In this com­mand cen­ter I watched the new site come to life in­cre­men­tally. That left me with a sim­i­lar un­der­stand­ing to do­ing it by hand — but much faster, be­cause the whole ex­pe­ri­ence was laid out for me.

The point here is that agents can write bits of code that help us hu­mans un­der­stand other code.

This is a big deal!

Technique 3: Shared spaces

Alright, last tech­nique: shared spaces. So far this has all been about un­der­stand­ing solo… but when you’re work­ing on a team, you need to un­der­stand to­gether.

When you and some­one else hold the same men­tal model, you can com­mu­ni­cate ef­fi­ciently. You have a shared vo­cab­u­lary that evokes the same im­ages, so you can jam and riff and have cre­ative con­ver­sa­tions. Without those shared struc­tures, those con­ver­sa­tions are much harder.

I’m re­ally ex­cited about cre­at­ing shared en­vi­ron­ments where teams build that un­der­stand­ing to­gether. It’s kinda what Notion is all about too.

Recently in Notion we’ve been ship­ping tons of new fea­tures for hu­mans and agents to work to­gether, so your whole team de­vel­ops a shared un­der­stand­ing in­stead of each work­ing in a silo.

One tiny ex­am­ple: you can now run Claude and Cursor agents in Notion. I do a lot of my cod­ing that way now.

And when those agents make a tech­ni­cal plan in Notion, it’s in a col­lab­o­ra­tive page by de­fault, so I can com­ment on it with my team and dis­cuss im­me­di­ately. Thinking to­gether, not alone!

The point was al­ways to aug­ment

Alright, let’s wrap up. Today we’ve cov­ered some tech­niques that were about un­der­stand­ing code… but ac­tu­ally I think this is a much big­ger is­sue.

It’s still im­por­tant for hu­mans to un­der­stand how things work in gen­eral! Not just to ver­ify, but to par­tic­i­pate.

And sur­prise sur­prise, this is not a new idea. It harkens back to the very ori­gins of our field of com­put­ing…

50 years ago Alan Kay en­vi­sioned that com­put­ers could be a new medium, bet­ter than the book, for teach­ing peo­ple — es­pe­cially kids — how to think about the world.

In this pic­ture, it might look like these kids are watch­ing YouTube on an iPad, but they’re not. They’re play­ing an in­ter­ac­tive game and edit­ing the code as they play it to get a bet­ter un­der­stand­ing of physics. This was 50 years ago!!

And now hope­fully you un­der­stand this meme.

The point was al­ways to aug­ment, not just au­to­mate.

It’s beau­ti­ful that AI now makes cre­at­ing sim­u­la­tions so ac­ces­si­ble… Having AI teach us is one of the great­est pos­si­bil­i­ties com­put­ing has ever opened up.

This makes me very op­ti­mistic about the fu­ture!

If we build the right tools, we can now un­der­stand the world bet­ter than we ever could be­fore. We don’t have to merely take our­selves out of the loop, we can get deeper in the loop too. It’s up to us.

FIN

Related reads

If you en­joyed this talk, you might like these other posts I’ve writ­ten about hu­man-AI col­lab­o­ra­tion:

Enough AI copi­lots! We need AI HUDs — anyone se­ri­ous about de­sign­ing for AI should con­sider non-copi­lot form fac­tors that more di­rectly ex­tend the hu­man mind…”

AI-generated tools can make pro­gram­ming more fun — Instead, I used AI to build a cus­tom de­bug­ger UI… which made it more fun for me to do the cod­ing my­self…”

Code like a sur­geon — identify and del­e­gate the sec­ondary grunt work tasks, so you can fo­cus on the main thing that mat­ters.”

OCR 4.1 - Mistral AI

docs.mistral.ai

July 16, 2026

Public Previewv4.1

Our lat­est OCR ser­vice pow­er­ing our Document AI stack, with na­tive para­graph-level bound­ing box ex­trac­tion, struc­tural block la­bels, and block-level con­fi­dence scores.

Every Fucking Website

lxe.github.io

In case you’re not aware, there’s COVID-19 hap­pen­ing! Here’s some stuff we wrote that you won’t read.

You ob­vi­ously know what cook­ies are. If we don’t put this here, de­lighted lawyers from EU and CA will sue us. Not only this is very ex­pen­sive, there are no browser set­tings to re­move this, since every site does this dif­fer­ently. You voted for this! Also you have to click I agree”

Nine PBS sues Iron Mountain over blocked access to archival data

current.org

Nine PBS in St. Louis filed a law­suit against in­for­ma­tion man­age­ment cor­po­ra­tion Iron Mountain Data Centers July 28, seek­ing to re­cover over 50 ter­abytes of archival ma­te­ri­als stored in one of the com­pa­ny’s Denver-based data cen­ters.

The law­suit filed in Denver District Court al­leges that the sta­tion’s cloud-stor­age ven­dor, Open Source Storage, abruptly cut off ac­cess to Nine PBS data ear­lier this year with­out warn­ing. It states OSS, which had a sep­a­rate re­la­tion­ship with Iron Mountain to pro­vide data stor­age, went defunct,” leav­ing Nine PBS archives in a data cen­ter op­er­ated by Iron Mountain.

Iron Mountain has re­fused to re­turn the ma­te­ri­als to the sta­tion be­cause its client, OSS, tech­ni­cally owned the phys­i­cal ser­vices hous­ing the data” within Iron Mountain, ac­cord­ing to the com­plaint.

The sta­tion re­quested tem­po­rary and pre­lim­i­nary re­lief that would pre­vent Iron Mountain from delet­ing, mod­i­fy­ing or over­writ­ing its ma­te­ri­als in the suit. A dis­trict judge granted the mo­tion and set a hear­ing for Wednesday.

In a state­ment to Current, Nine PBS VP and CCO Leah Freeman con­firmed the sta­tion’s law­suit against Iron Mountain and its ded­i­ca­tion to re­triev­ing the archival ma­te­ri­als and pro­gram­ming, which she says span over 70 years of our or­ga­ni­za­tion’s his­tory.”

We are com­mit­ted to en­sur­ing we can re­cover and re­store full ac­cess to this valu­able con­tent, which Nine PBS right­fully owns, as it holds sig­nif­i­cant his­tor­i­cal im­por­tance for St. Louis.”

According to the law­suit, the blocked ma­te­ri­als in­clude his­tor­i­cal items such as Nine PBS cov­er­age on the his­tory of East St. Louis, the COVID-19 pan­demic and the Great Flood of 1993, which rav­aged along the Mississippi and Missouri rivers.

No choice but to file’

Nine PBS en­tered a re­la­tion­ship with a com­pany de­scribed in the com­plaint as OSS pre­de­ces­sor” in 2019. This uniden­ti­fied ven­dor pro­vided hardware, soft­ware and cloud-stor­age ser­vices” for stor­ing the pub­lic broad­cast­er’s archival ma­te­ri­als and other data.

Nine PBS re­newed its con­tracts with the data ser­vices ven­dor and sub­se­quently OSS an­nu­ally, the com­plaint states.

When Nine PBS at­tempted to sched­ule a meet­ing with OSS in February to dis­cuss re­new­ing for 2026, OSS did­n’t re­spond or in­di­cate any in­ten­tion not to re­new the agree­ment.” The agree­ment was set to ex­pire on March 6.

The con­tract pro­vided 30 days for Nine PBS to re­trieve its data from OSS stor­age upon ter­mi­na­tion of ser­vices.” But on March 6, OSS cut off the sta­tion’s ac­cess with­out warn­ing, ac­cord­ing to the com­plaint. When Nine PBS at­tempted to con­tact OSS to sort out the prob­lem, it dis­cov­ered that OSS web­site was de­funct and the com­pany had delin­quency sta­tus with the Colorado Secretary of State.

To en­sure the sta­tion’s data re­mained se­cure, Nine PBS in­ves­ti­gated fur­ther and dis­cov­ered that OSS had a re­la­tion­ship with Iron Mountain, ac­cord­ing to the com­plaint.

Nine PBS sent a de­mand let­ter March 13, de­mand­ing that Iron Mountain pre­serve and re­turn its data and of­fer­ing to pay any rea­son­able costs as­so­ci­ated with its de­mand.” Iron Mountain nei­ther con­firmed nor de­nied the data was in its pos­ses­sion, the law­suit states.

Nine PBS filed a law­suit against OSS and its purported” pres­i­dent Charles Wells in the St. Louis Circuit Court April 16, the com­plaint states. The broad­caster later paused the lit­i­ga­tion af­ter James Tramel, a managing part­ner of the group that of­fi­cially ac­quired” OSS as­sets, con­firmed that Nine PBS data was se­cure within Iron Mountain’s Denver data cen­ter.

After com­mu­ni­cat­ing with Nine PBS for about a month, Tramel stopped re­spond­ing. Weeks later, an au­to­matic re­ply email from his ac­count stated that he was no longer af­fil­i­ated with OSS. Tramel re­vealed in a sub­se­quent phone call that he had been de­frauded” into pur­chas­ing OSS, ac­cord­ing to the com­plaint. At this point, the com­pa­ny’s pre­vi­ous own­ers, in­clud­ing Wells, Ben Nicholson, and Justine Ririe, re­sumed con­trol of OSS op­er­a­tions.

After Nine PBS at­tempted to con­tact OSS lead­er­ship with­out suc­cess, the sta­tion re­turned to the St. Louis Circuit Court and ob­tained a de­fault judg­ment against OSS. The court’s rul­ing stated that the sta­tion both owned its data and had an immediate right to pos­sess the data,” ac­cord­ing to the com­plaint. The judg­ment also or­dered OSS to re­turn the data to Nine PBS and/or fa­cil­i­tate its trans­fer to a new ven­dor.”

An at­tor­ney rep­re­sent­ing Nine PBS con­tacted Iron Mountain about re­turn­ing the data, and noted that litigation was im­mi­nent in Colorado,” and the com­pany ad­mit­ted that it pos­sessed the data, the com­plaint states. Iron Mountain ini­tially in­di­cated that it wanted to avoid lit­i­ga­tion and com­ply with Nine PBS re­quest, but later re­fused to do so, cit­ing OSS own­er­ship of the in­fra­struc­ture that houses the data.

Nine PBS com­plaint states that, fol­low­ing mul­ti­ple un­suc­cess­ful at­tempts to re­trieve its dig­i­tal prop­erty, it had no choice but to file this law­suit to force Iron Mountain to pro­tect and ul­ti­mately pro­vide” the data.

Iron Mountain did not re­spond to a re­quest for com­ment.

Why does Opus 5 feel worse to work with?

mun-logadan.github.io

In my opin­ion and that of the col­leagues I’ve spo­ken with, work­ing with Opus 5 feels like a down­grade com­pared to Opus 4.7, Opus 4.8, and Fable.

I’m not claim­ing a step back­wards in ca­pa­bil­i­ties — it is a more ca­pa­ble model than Opus 4.7 and Opus 4.8 and even ri­vals Fable in bench­marks, yet these other mod­els feel bet­ter to work with. I be­lieve this is be­cause they:

stop and ask ques­tions if my in­tent was un­clear,

don’t make as­sump­tions with­out check­ing,

and don’t rein­ter­pret or up­date my plans with­out ask­ing.

Because of this, they don’t re­quire the care­ful babysit­ting that Opus 5 does.

Baseless spec­u­la­tion

I sus­pect this is the re­sult of two com­pound­ing forces at Anthropic, and in cur­rent fron­tier labs in gen­eral.

First, the de­sire to cre­ate a self-im­prov­ing AI that is ca­pa­ble of re­cur­sively boot­strap­ping it­self to AGI/ASI.

Second, the pres­sure to score highly on bench­marks. Although it’s an open se­cret that many bench­mark tasks are ill-de­fined, un­fair, hack­able, or oth­er­wise bro­ken, a good bench­mark task is self-con­tained. It can be solved. It does­n’t re­quire hints, read­ing the task cre­ator’s mind, or out­side in­for­ma­tion to pass.

That does­n’t mean a good task can only have one cor­rect an­swer, just that it should score all un­am­bigu­ously cor­rect an­swers equally.

Selecting for mod­els that do well on bench­marks (and in­deed train­ing for them or on RLVR tasks in gen­eral) in­her­ently se­lects for mod­els that make bold, usu­ally-cor­rect as­sump­tions in the face of am­bi­gu­ity. It pe­nal­izes mod­els with a ten­dency to stop and ask for clar­i­fi­ca­tion or di­rec­tion.

Unfortunately, that’s ex­actly what most of us want from a cod­ing agent.

Try as you might, it’s nearly im­pos­si­ble to get the en­tirety of the con­text, in­ten­tions, busi­ness im­pli­ca­tions, bud­get con­straints, and what-have-you writ­ten down and ac­ces­si­ble to a cod­ing agent. There will in­vari­ably be am­bi­gu­ity and choices to be made, and it is nice to know that an agent will stop and ask when needed.

Real life just is­n’t a bench­mark. There is­n’t a guar­an­teed right an­swer to every ques­tion, nor even a set of right an­swers, and with real-life con­se­quences on the line, I do not want an agent tak­ing its best guess!

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.