10 interesting stories served every morning and every evening.

openai.com

Just a moment...

ads.openai.com

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

fireworks.ai

K3 is a fron­tier qual­ity open model at a frac­tion of the cost. Even big­ger is that it com­ple­ments Fable pre­dictably, which makes it pos­si­ble to get the high­est qual­ity in­tel­li­gence by rout­ing tasks.

🧭 tl;dr: We ran Kimi K3 (open) against Fable 5 (closed) on ~1,000 agen­tic tasks find­ing:

We achieved 93% ac­cu­racy with rout­ing be­tween K3 and Fable.

Results were up to ~50X more cost ef­fec­tive than Fable alone on long agen­tic loops, and con­sis­tently lower cost across every use case.

How We Measured

We av­er­aged bench­marks, each aimed at a dif­fer­ent kind of work, and ran K3 and Fable 5 through the same har­ness. About 1,030 tasks in all, in real agent loops.

One quick de­f­i­n­i­tion be­fore we get into the re­sults. Oracle rout­ing is a method for mea­sur­ing the best the­o­ret­i­cal per­for­mance by run­ning the task through each model and then pick­ing the cheap­est cor­rect op­tion (the cost/​per­for­mance ceil­ing). In a prac­ti­cal router, you don’t get to run your task against mul­ti­ple mod­els. The router makes a pre­dic­tion of which model has the best cost and qual­ity trade off, but ul­ti­mately it’s a guess.

In this study, or­a­cle rout­ing demon­strated K3 is se­lected for 72 – 96% of tasks. This sug­gests a near-per­fect router might be achiev­able, by learn­ing the dif­fer­ence be­tween day-to-day tasks and the true long tail of fron­tier work. It will re­quire an or­der of mag­ni­tude more rout­ing data, and real world per­for­mance to say de­fin­i­tively.

K3 is a good model.

From a 10,000 foot view, it can be easy to look at both mod­els and call the head-to-head a tie. For ex­am­ple, if you look at SWE, the head­line bench­mark, K3 gets 92.4%, Fable 92.6%. Across the five types of tasks we bench­marked on, the two mod­els tend to stay within a few points of each other, with Fable pulling slightly ahead on its cod­ing-lan­guage breadth (Multi-lang).

It’s easy to stop there and say they’re roughly even”. The news is that they have dis­cretely bet­ter per­for­mance across dif­fer­ent task types.

Two Models is Better than One

If you take a peek in­side a sin­gle bench­mark, there’s more to see than just a top-line ac­cu­racy num­ber. Take SWE, where the two are dead even over­all. If you split SWE by prob­lem do­main you can see where each model shines. K3 is sharpest on sym­bolic math and dev tool­ing; Fable wins on web & data vi­su­al­iza­tion work. The same pat­tern runs through the multi-lan­guage set, where Fable’s breadth car­ries Java, Python and C++, while K3 draws even on JavaScript and Rust.

For long-hori­zon work at a ter­mi­nal, dri­ving a shell and prod­ding at sys­tems across dozens of turns, K3 showed its true col­ors. It cleared a batch of tasks Fable never cracked: a 7z hash, FEAL crypt­analy­sis, leaked se­crets, a live vul­ner­a­bil­ity, run­away async jobs.

K3 can be up to 50x lower cost on Fireworks. 🫳🎤

While qual­ity is a near-tie at a high level, price is­n’t close.

So where’s this huge price gap com­ing from? to­ken pric­ing, prompt caching, and ef­fort-per-task. On SWE for ex­am­ple, K3 works much harder than Fable: roughly 55 turns and 1.3M to­kens a task ver­sus 21 turns and 130K. On the long ter­mi­nal tasks it’s the other way around: Fable is the one that spi­rals, run­ning up 64 turns and 1.5M to­kens (sometimes straight into a time­out).

Prompt caching does most of the work of turn­ing that ef­fort into K3′s price ad­van­tage: even when K3 reads ten times the to­kens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a trade­off. Tasks with ex­tra turns gen­er­ally mean more wall-clock time per run i.e. slower runs. If you need an an­swer in two sec­onds, that mat­ters; if you’re run­ning agents in the back­ground at scale, a bill that’s a frac­tion of the size mat­ters a lot more.

Don’t pick a model. Route.

If you send every task to who­ever han­dles it best, you don’t land some­where be­tween the two mod­els, you land above both.

Per-task rout­ing al­ways out per­forms any sin­gle model run:

The or­a­cle router choose K3, 72 – 96% of task traf­fic. By ar­chi­tect­ing a router this way, you end up with over­all qual­ity above ei­ther model alone at a cost close to just us­ing just the cost-op­ti­mized one.

K3 is cost op­ti­mized on all work types

Put both qual­ity and cost on one plot. K3 in blue lands to the left (the more cost-ef­fec­tive side) of Fable in red in all five task-fam­i­lies. Accuracy trades back and forth: Fable pulls ahead on multi-lan­guage, K3 on ter­mi­nal and le­gal, the rest roughly level.

Single Models Are Wasteful and No Longer SoTA

Kimi K3 + Fable routed to­gether un­locks their best qual­i­ties at the best price.

The sin­gle model provider, to­ken maxxing days, are com­ing to an end. The task-level data says these mod­els are spe­cial­ists at very dif­fer­ent prices. The best AI no longer comes out of a sin­gle lab, it’s a mix­ture of mod­els.

What this means in prac­tice:

Open as the de­fault. A 50x lower cost open model like K3 should be your base case, since the or­a­cle sends it most of the traf­fic any­way.

The router is your moat. A router must be tai­lored to your work­load and learn­ing that task/​model split con­tin­u­ously is the best chance you’ll have at stay­ing ahead.

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

blog.google

Jul 21, 2026

|

Our newest Gemini mod­els de­liver the ef­fi­ciency, la­tency, and re­li­a­bil­ity to build AI agents at scale.

Developers and cus­tomers build­ing pro­duc­tion AI agents need higher to­ken ef­fi­ciency, lower la­tency, and more re­li­able per­for­mance. Our Flash se­ries of mod­els is built to meet the sweet spot of ef­fi­ciency and qual­ity to en­able scal­ing agen­tic work­flows. Building on Gemini 3.5 Flash, we’re in­tro­duc­ing new Gemini mod­els:

3.6 Flash: Our work­horse model that de­liv­ers bet­ter cod­ing, knowl­edge work, and mul­ti­modal per­for­mance. According to the Artificial Analysis Index, it re­duces out­put to­ken us­age by 17% com­pared to 3.5 Flash, and in some bench­marks like DeepSWE by Datacurve, we ob­serve up to 65%, all at a lower cost per out­put to­ken.

3.5 Flash-Lite: Our fastest, most cost-ef­fec­tive 3.5-class model, de­liv­er­ing 350 out­put to­kens per sec­ond ac­cord­ing to the Artificial Analysis Index, also sig­nif­i­cantly out­per­form­ing prior Flash-Lite gen­er­a­tions in agen­tic work­flows.

3.5 Flash Cyber in CodeMender: Successful cy­ber­se­cu­rity ap­pli­ca­tions re­quire care­ful or­ches­tra­tion of a model along­side an agent in­fra­struc­ture. We’re in­tro­duc­ing a com­bi­na­tion of a new, highly ef­fi­cient, spe­cial­ized cy­ber-fo­cused model paired with our CodeMender code se­cu­rity agent that de­liv­ers com­pet­i­tive per­for­mance at the fron­tier.

Beyond to­day’s re­leases, Gemini 3.5 Pro is cur­rently test­ing with part­ners and we plan to make it broadly avail­able as soon as it’s ready. In par­al­lel, our team is al­ready fo­cus­ing on build­ing the next gen­er­a­tion of mod­els. We have started our most am­bi­tious pre-train­ing run yet, for Gemini 4, and are ex­cited by the progress.

3.6 Flash: More ef­fi­cient and bet­ter qual­ity than 3.5 Flash

Gemini 3.6 Flash builds di­rectly on de­vel­oper and cus­tomer feed­back from 3.5 Flash. 3.6 Flash not only de­liv­ers a step up in cod­ing and knowl­edge work, but it does this while mean­ing­fully im­prov­ing to­ken ef­fi­ciency. For ex­am­ple, on the Artificial Analysis Index, we see 3.6 Flash con­sum­ing 17% fewer out­put to­kens than 3.5 Flash. It also takes fewer rea­son­ing steps and tool calls to ac­com­plish multi-step work­flows.

This en­hanced ef­fi­ciency is also com­bined with a lower price than 3.5 Flash. At $1.50/1M in­put to­kens and $7.50/1M out­put to­kens, 3.6 Flash re­duces the over­all cost per agen­tic task, mak­ing agents more cost-ef­fec­tive to build and run.

3.6 Flash shows bet­ter to­ken ef­fi­ciency and re­duced ver­bosity than 3.5 Flash in an OSWorld ver­i­fied task (API)

Even while be­ing more ef­fi­cient, 3.6 Flash sees per­for­mance gains com­pared to 3.5 Flash across use cases:

3.6 Flash de­liv­ers higher pre­ci­sion with fewer un­wanted code ed­its and re­duced ex­e­cu­tion loops, as seen in DeepSWE (49% vs. 37%), and shows sig­nif­i­cant im­prove­ment in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).

It has im­proved com­puter use ca­pa­bil­i­ties as seen in OSWorld-Verified (83.0% vs. 78.4%). Computer use is now a built-in client side tool via the Gemini API and Gemini Enterprise.

It out­per­forms 3.5 Flash in knowl­edge work, as shown by bench­marks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it par­tic­u­larly ca­pa­ble at mul­ti­modal tasks like doc­u­ment pars­ing, chart and data analy­sis, and re­port draft­ing.

Customers re­port 3.6 Flash is a step for­ward in both cost and qual­ity, bal­anc­ing to­ken ef­fi­ciency, ac­cu­racy, and speed across com­plex work­flows and knowl­edge-based tasks:

Built with safety

3.6 Flash is ship­ping with en­hanced Frontier Safety safe­guards in the do­mains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cy­ber of­fense mis­uses. These safe­guards make the model sub­stan­tially more re­sis­tant to jail­breaks. At the same time, the model has been trained to min­i­mize re­fusals for ben­e­fi­cial uses.

For more in­for­ma­tion, see the 3.6 Flash model card.

3.5 Flash-Lite: Built to scale agen­tic work­flows

Beyond Flash, we’re also re­leas­ing Gemini 3.5 Flash-Lite, de­signed for both low-la­tency tasks and tasks where high through­put is crit­i­cal for de­vel­op­ers work­flows, like agen­tic search and doc­u­ment pro­cess­ing.

3.5 Flash-Lite is the fastest model in the 3.5 se­ries. As mea­sured by Artificial Analysis, it runs at 350 out­put to­kens/​s. Priced at $0.3/1M in­put to­kens and $2.5/1M out­put to­kens and with sig­nif­i­cantly bet­ter qual­ity than 3.1 Flash-Lite, 3.5 Flash-Lite of­fers a strong price-to-per­for­mance ra­tio for de­vel­op­ers and cus­tomers run­ning high through­put pro­duc­tion traf­fic.

3.5 Flash-Lite ex­e­cutes high vol­ume tasks at a lower la­tency than 3.5 Flash.

3.5 Flash-Lite en­ables ef­fi­cient scal­ing for agen­tic sys­tems. Across think­ing lev­els, the model sig­nif­i­cantly out­per­forms 3.1 Flash-Lite. Depending on the work­load, de­vel­op­ers can con­fig­ure the model to pri­or­i­tize low-la­tency, low-cost ex­e­cu­tion for high-vol­ume tasks with the min­i­mal and low think­ing lev­els, or en­gage higher think­ing lev­els to process multi-step sub­agent work­loads. The model now also has com­puter use as a built-in tool to re­li­ably sup­port these agen­tic tasks across sur­faces.

It’s a sig­nif­i­cant step up in cod­ing and agen­tic tasks as seen in Terminal-Bench 2.1 (54% vs 31%), long con­text as seen in GDM-MRCR v2 (72.2% vs. 60.1%), and real-world task ex­e­cu­tion as seen in GDPval-AA v2 (1140 vs. 642).

In fact, on many agen­tic and cod­ing evals, 3.5 Flash-Lite even out­per­forms 3 Flash, in­clud­ing on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), mak­ing it a faster & more ca­pa­ble op­tion for work­loads on both 2.5 and 3 Flash.

Early cus­tomers of 3.5 Flash-Lite are high­light­ing its unique com­bi­na­tion of speed, in­tel­li­gence, and cost ef­fi­ciency for scal­ing agen­tic work­flows and data pro­cess­ing tasks:

For more in­for­ma­tion about the model, see the 3.5 Flash-Lite model card.

3.5 Flash Cyber in CodeMender: find­ing and fix­ing vul­ner­a­bil­i­ties ef­fi­ciently

AI mod­els have be­come ca­pa­ble of find­ing se­cu­rity vul­ner­a­bil­i­ties faster than cur­rent sys­tems can fix them. Tackling this grow­ing threat re­quires an ap­proach to se­cur­ing soft­ware that is highly ca­pa­ble and ef­fi­cient.

Flash’s per­for­mance and ef­fi­ciency makes it an ideal foun­da­tion to de­tect, val­i­date, and patch code se­cu­rity is­sues at scale. Gemini 3.5 Flash Cyber is built on top of 3.5 Flash, and fine-tuned for find­ing and fix­ing cy­ber­se­cu­rity vul­ner­a­bil­i­ties at a lower price per to­ken than larger mod­els.

Within CodeMender, which uses mul­ti­ple 3.5 Flash Cyber agents work­ing to­gether to pro­duce a sin­gle com­bined re­port, 3.5 Flash Cyber reaches com­pet­i­tive per­for­mance at the fron­tier on the pop­u­lar bench­mark CyberGym.

Given the dual-use na­ture of this tech­nol­ogy, we have taken an in­ten­tional ap­proach to de­ploy­ing 3.5 Flash Cyber. The model will be ex­clu­sively avail­able to gov­ern­ments and trusted part­ners via CodeMender soon as part of a lim­ited-ac­cess pi­lot pro­gram. This will give front­line de­fend­ers a head start in find­ing and fix­ing crit­i­cal vul­ner­a­bil­i­ties be­fore they can be ex­ploited, while mit­i­gat­ing against broader mis­use.

3.6 Flash and 3.5 Flash-Lite: Get started to­day

3.6 Flash and 3.5 Flash-Lite are avail­able start­ing to­day:

For de­vel­op­ers in the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also avail­able in Google Antigravity. Get started with the Developer Guide.

For en­ter­prises in Gemini Enterprise Agent Platform. 3.6 Flash is also avail­able in the Gemini Enterprise app.

For every­one via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search.

As you start build­ing with 3.6 Flash and 3.5 Flash-Lite, we wel­come your feed­back to im­prove fu­ture Gemini mod­els and look for­ward to re­leas­ing 3.5 Pro soon.

Get the lat­est news from Google in your in­box

Sign up for our newslet­ters with prod­uct up­dates, event in­for­ma­tion, spe­cial of­fers, and more.

Your in­for­ma­tion will be used in ac­cor­dance with Google’s pri­vacy pol­icy. You may opt out at any time.

'VPNs are lawful technical tools,' says EU Court in landmark Anne Frank copyright ruling

www.techradar.com

VPN providers aren’t li­able for copy­right in­fringe­ment, said the EU Court

The Court ex­plic­itly rec­og­nized VPNs as lawful tech­ni­cal tools”

The case cen­tered on the copy­right bat­tle in­volv­ing Anne Frank’s di­ary

In a ma­jor vic­tory for dig­i­tal rights and com­mon sense, the Court of Justice of the European Union (CJEU) has of­fi­cially cat­e­go­rized Virtual Private Networks (VPNs) as lawful tech­ni­cal tools” while es­tab­lish­ing new bound­aries for on­line copy­right dis­putes.

The land­mark judg­ment — handed down in July 2026 — stems from a com­plex le­gal bat­tle over the on­line pub­li­ca­tion of Anne Frank’s his­tor­i­cal man­u­scripts. At its core, the case forced Europe’s top judges to an­swer a highly tech­ni­cal ques­tion: if a pub­lisher ac­tively tries to block vis­i­tors from a spe­cific coun­try, are they still break­ing the law if a user sneaks past the dig­i­tal bor­der us­ing cir­cum­ven­tion soft­ware?

According to the CJEU, the an­swer is no. As long as a web­site em­ploys state-of-the-art” geo-block­ing tech­nol­ogy, the pub­lisher can­not be held li­able for copy­right in­fringe­ment sim­ply be­cause a de­ter­mined reader de­cides to fire up the best VPN to by­pass the re­stric­tions.

The rul­ing sets a mas­sive prece­dent. It con­firms that copy­right hold­ers can­not point to the mere ex­is­tence of VPNs to claim a web­site’s se­cu­rity mea­sures are com­pletely in­ef­fec­tive.

More im­por­tantly for pri­vacy ad­vo­cates, the court firmly pushed back against the de­mo­niza­tion of pri­vacy soft­ware, ce­ment­ing the le­git­i­mate sta­tus of VPN providers across the European Union.

The Anne Frank dis­pute ex­plained

The EUs top court just con­firmed: Geo-blocking is the copy­right hold­er’s prob­lem, not the VPNs. Providers are not li­able for users by­pass­ing re­stric­tions ⚖️ @torrentfreak https://​t.co/​fLw­b5kYAy1July 17, 2026

The EUs top court just con­firmed: Geo-blocking is the copy­right hold­er’s prob­lem, not the VPNs. Providers are not li­able for users by­pass­ing re­stric­tions ⚖️ @torrentfreak https://​t.co/​fLw­b5kYAy1July 17, 2026

The le­gal tug-of-war be­gan when a coali­tion of Dutch and Belgian aca­d­e­mic in­sti­tu­tions pub­lished a free, schol­arly on­line edi­tion of Anne Frank’s man­u­scripts.

Because copy­right laws are not fully har­mo­nized across Europe, the le­gal sta­tus of the fa­mous di­ary varies by ter­ri­tory. In Belgium and roughly 60 other coun­tries, the writ­ings en­tered the pub­lic do­main years ago. However, in the Netherlands, parts of the text re­main pro­tected by copy­right un­til 2037.

To re­spect this ter­ri­to­r­ial di­vide, the pub­lish­ers hosted the site in Belgium and used geo-block­ing to pre­vent ac­cess from Dutch IP ad­dresses. Visitors from the Netherlands were met with a no­tice ex­plain­ing why they could­n’t en­ter the site.

The Anne Frank Fonds, which holds the Dutch copy­right, sued. They ar­gued that be­cause stan­dard VPNs eas­ily al­low users to mask their true IP ad­dress and spoof a Belgian lo­ca­tion, the schol­arly web­site was ef­fec­tively com­mu­ni­cat­ing the pro­tected work to the Dutch pub­lic.

The CJEU ul­ti­mately re­jected this ar­gu­ment. In its judg­ment, the court noted that while geo-block­ing mea­sures can in­evitably be cir­cum­vented, the pos­si­bil­ity of such cir­cum­ven­tion can­not, in it­self and in all cir­cum­stances, be a de­ci­sive fac­tor in find­ing those mea­sures to be in­ad­e­quate and, there­fore, in­ef­fec­tive.”

Why this mat­ters for the in­ter­net and VPN users

For every­day in­ter­net users, the So What?” of this rul­ing is deeply re­as­sur­ing. It val­i­dates that us­ing a VPN to en­crypt your on­line traf­fic, hide your IP ad­dress, or by­pass dig­i­tal bor­ders is a le­git­i­mate use of con­sumer tech­nol­ogy.

The judges ex­plic­itly shielded VPN com­pa­nies from col­lat­eral dam­age in piracy law­suits, a topic that has sparked in­tense de­bate among European ISPs and right­sh­old­ers. The court specif­i­cally ar­gued that the provider of a VPN or sim­i­lar ser­vices is not li­able for users by­pass­ing re­stric­tions.

By plac­ing the le­gal bur­den on pub­lish­ers to main­tain state-of-the-art” dig­i­tal fences, rather than de­mand­ing ab­solute, im­pos­si­ble per­fec­tion, the EU has drawn a prag­matic line in the sand.

Publishers aren’t ex­pected to build un­hack­able walls, and VPN providers aren’t re­spon­si­ble for the ac­tions of users who climb over them. Ultimately, this rul­ing proves that the bor­der­less in­ter­net can still co­ex­ist with ter­ri­to­r­ial copy­right laws, pro­vided every­one uses the right tech­ni­cal safe­guards.

Follow TechRadar on Google News and add us as a pre­ferred source to get our ex­pert news, re­views, and opin­ion in your feeds. Make sure to click the Follow but­ton!

Free Ink · An open ecosystem for e-readers

freeink.org

Judge approves a $1.5B Anthropic settlement over books used to train Claude | AP News

apnews.com

SAN FRANCISCO (AP) — A fed­eral judge has ap­proved a $1.5 bil­lion copy­right set­tle­ment in which ar­ti­fi­cial in­tel­li­gence com­pany Anthropic will pay thou­sands of au­thors about $3,000 per book af­ter us­ing pi­rated copies of their works to train its Claude chat­bot.

District Judge Araceli Martínez-Olguín said in a Monday rul­ing that the class-ac­tion set­tle­ment pro­vides meaningful re­lief” to af­fected au­thors and pub­lish­ers.

About 91% of the more than 482,000 books cov­ered by the rul­ing have been claimed by au­thors or pub­lish­ers who are now due pay­ment.

Plaintiff at­tor­ney Justin Nelson said in a state­ment that the set­tle­ment was the largest known copy­right re­cov­ery in his­tory. We look for­ward to mak­ing dis­tri­b­u­tions to the Class as promptly as pos­si­ble.”

U.S. District Judge William Alsup is­sued the pre­lim­i­nary ap­proval in San Francisco fed­eral court last September and has since re­tired. Alsup had dealt the case a mixed rul­ing last sum­mer, find­ing that train­ing AI chat­bots on copy­righted books was­n’t il­le­gal but that Anthropic wrong­fully ac­quired mil­lions of books through pi­rate web­sites.

Anthropic’s deputy gen­eral coun­sel, Aparna Sridhar, high­lighted that rul­ing Friday as a land­mark show­ing that train­ing AI on books is fair use un­der copy­right law.”

We are pleased that more than 91% of au­thors and pub­lish­ers cov­ered by the set­tle­ment have claimed their share of the pay­ment, and we’re look­ing for­ward to bring­ing this mat­ter to a close,” Sridhar said in a writ­ten state­ment.

Bestselling thriller nov­el­ist Andrea Bartz first brought the suit with two other au­thors in 2024. It’s the first ma­jor set­tle­ment in dozens of AI copy­right law­suits that are still work­ing their way through courts.

Apple Defeats Liability for Not Scanning iCloud Items for CSAM, But the Judge Was Not Pleased–Amy v. Apple

blog.ericgoldman.org

This case in­volves Apple’s han­dling of user-up­loaded files hosted in pri­vate iCloud stor­age. Instead of adopt­ing PhotoDNA to scan hosted files for CSAM, Apple cre­ated its own pro­pri­etary al­ter­na­tive, NeuralHash, which ap­par­ently was­n’t as good. So Apple U-turned on its ef­forts to scan for CSAM in its cloud stor­age and en­crypts iCloud files.

Apple’s manuev­ers con­fused the pub­lic and seemed like an em­bar­rass­ing un­forced er­ror for Apple. It also en­sured pres­sure from gov­ern­ments and plain­tiffs, in­clud­ing CSAM vic­tims, who pre­ferred Apple’s more in­ter­ven­tion­ist ap­proaches, which Apple had vol­un­tar­ily demon­strated it was will­ing to do.

This law­suit rep­re­sents a full-scale at­tack on Apple and Section 230. Plaintiffs al­lege that Apple’s fail­ure to im­ple­ment any known CSAM de­tec­tion is a de­sign de­fect be­cause Apple can safely im­ple­ment read­ily avail­able fea­tures to pre­vent the spread of known CSAM but has con­tin­u­ously failed to do so.” Prior blog post. The court dis­misses the third amended com­plaint, which tees this case up for the Ninth Circuit, where (as usual) any­thing could hap­pen.

* * *

The court re­it­er­ates that Section 230 ap­plies to the plain­tiffs’ claims:

First, Plaintiffs’ claims treat Apple as a pub­lisher or speaker of the CSAM con­tent that an­i­mates Plaintiffs’ in­juries. Fundamentally, Plaintiffs con­tend that Apple has elected to per­mit users to dis­sem­i­nate and share third-party CSAM con­tent when it could have—and, in their view, should have—used read­ily avail­able tech­nol­ogy to pre­vent the dis­tri­b­u­tion of child pornog­ra­phy de­pict­ing the Plaintiffs in this pu­ta­tive class. The du­ties Plaintiffs seek to in­voke spring[ ] from the de­fen­dan­t’s sta­tus as pub­lisher,” and con­se­quently, immunity ap­plies.” Second, im­mu­nity also ap­plies be­cause the means to avoid li­a­bil­ity re­quires [Apple] to act as a pub­lisher.” As a re­sult, Apple is en­ti­tled to com­plete im­mu­nity un­der § 230.

First, Plaintiffs’ claims treat Apple as a pub­lisher or speaker of the CSAM con­tent that an­i­mates Plaintiffs’ in­juries. Fundamentally, Plaintiffs con­tend that Apple has elected to per­mit users to dis­sem­i­nate and share third-party CSAM con­tent when it could have—and, in their view, should have—used read­ily avail­able tech­nol­ogy to pre­vent the dis­tri­b­u­tion of child pornog­ra­phy de­pict­ing the Plaintiffs in this pu­ta­tive class. The du­ties Plaintiffs seek to in­voke spring[ ] from the de­fen­dan­t’s sta­tus as pub­lisher,” and con­se­quently, immunity ap­plies.” Second, im­mu­nity also ap­plies be­cause the means to avoid li­a­bil­ity re­quires [Apple] to act as a pub­lisher.” As a re­sult, Apple is en­ti­tled to com­plete im­mu­nity un­der § 230.

Citing Doe 1 v. Meta, the court says:

Plaintiffs’ in­juries are the di­rect re­sult of the ac­tions of third par­ties who used iCloud to share CSAM, a use Apple nei­ther ex­plic­itly con­dones nor pre­vents (even as­sum­ing—as al­leged in the TAC—that Apple was aware of the use of iCloud for this pur­pose)….though Plaintiffs al­lege that Apple knew that its tools were likely to be used to dis­trib­ute child pornog­ra­phy (as con­firmed by the in­ter­nal Apple text mes­sages at the cen­ter of this case), un­der the cur­rent state of the law, Apple is still en­ti­tled to im­mu­nity un­der § 230—irrespective of that gen­eral knowl­edge…. Plaintiffs can­not avoid the fact that a tool that de­tects CSAM must re­view CSAM to make such a de­ter­mi­na­tion. And while Apple could have taken steps to do so—as its com­peti­tors have done by us­ing PhotoDNA—Grindr con­firms that § 230 bars claims aris­ing from the de­sign de­ci­sions Apple could have taken where those claims re­late to Apple’s role fa­cil­i­tat­ing the com­mu­ni­ca­tion and con­tent of oth­ers

Plaintiffs’ in­juries are the di­rect re­sult of the ac­tions of third par­ties who used iCloud to share CSAM, a use Apple nei­ther ex­plic­itly con­dones nor pre­vents (even as­sum­ing—as al­leged in the TAC—that Apple was aware of the use of iCloud for this pur­pose)….though Plaintiffs al­lege that Apple knew that its tools were likely to be used to dis­trib­ute child pornog­ra­phy (as con­firmed by the in­ter­nal Apple text mes­sages at the cen­ter of this case), un­der the cur­rent state of the law, Apple is still en­ti­tled to im­mu­nity un­der § 230—irrespective of that gen­eral knowl­edge….

Plaintiffs can­not avoid the fact that a tool that de­tects CSAM must re­view CSAM to make such a de­ter­mi­na­tion. And while Apple could have taken steps to do so—as its com­peti­tors have done by us­ing PhotoDNA—Grindr con­firms that § 230 bars claims aris­ing from the de­sign de­ci­sions Apple could have taken where those claims re­late to Apple’s role fa­cil­i­tat­ing the com­mu­ni­ca­tion and con­tent of oth­ers

(A re­minder that the de­fen­dan­t’s sci­en­ter is ir­rel­e­vant to Section 230).

The plain­tiffs tried to fit into the new Section 230 ex­cep­tions cre­ated in Doe v. Twitter, but the court re­buffs the move:

This case does not con­cern or even dis­cuss Apple’s con­tent re­port­ing sys­tems; it con­cerns Apple’s failure to im­ple­ment in­dus­try-stan­dard safe­guards” against the dis­sem­i­na­tion of CSAM. Though re­port­ing sys­tems and CSAM safe­guards may both be de­scribed as defects,” the lat­ter re­quires the Court to treat Apple as a pub­lisher. Twitter could ful­fill its pur­ported duty to cure re­port­ing in­fra­struc­ture de­fi­cien­cies with­out mon­i­tor­ing, re­mov­ing, or in any way en­gag­ing with third-party con­tent”; Apple can­not ful­fill a duty to in­sti­tute CSAM safe­guards with­out de­ploy­ing a tool like NeuralHash or PhotoDNA. Both NeuralHash and PhotoDNA were built to mon­i­tor and re­port vi­ola­tive im­ages up­loaded to com­pany servers. Yet just the de­ci­sion re­gard­ing whether to de­ploy ei­ther tool is a choice re­lated to con­tent mod­er­a­tion.

This case does not con­cern or even dis­cuss Apple’s con­tent re­port­ing sys­tems; it con­cerns Apple’s failure to im­ple­ment in­dus­try-stan­dard safe­guards” against the dis­sem­i­na­tion of CSAM. Though re­port­ing sys­tems and CSAM safe­guards may both be de­scribed as defects,” the lat­ter re­quires the Court to treat Apple as a pub­lisher. Twitter could ful­fill its pur­ported duty to cure re­port­ing in­fra­struc­ture de­fi­cien­cies with­out mon­i­tor­ing, re­mov­ing, or in any way en­gag­ing with third-party con­tent”; Apple can­not ful­fill a duty to in­sti­tute CSAM safe­guards with­out de­ploy­ing a tool like NeuralHash or PhotoDNA. Both NeuralHash and PhotoDNA were built to mon­i­tor and re­port vi­ola­tive im­ages up­loaded to com­pany servers. Yet just the de­ci­sion re­gard­ing whether to de­ploy ei­ther tool is a choice re­lated to con­tent mod­er­a­tion.

Also, Apple did­n’t fail to sat­isfy any duty to re­port items to NCMEC if it never iden­ti­fied CSAM in the first place.

The Lemmon v. Snap workaround fails: all of Plaintiffs’ claims here are in­ex­orably linked to third-party con­tent; Plaintiffs do not al­lege that Apple cre­ated con­tent like a Snapchat fil­ter that caused them harm.” The Roommates.com workaround also fails: Plaintiffs do not al­lege that Apple mod­i­fied or aug­mented the CSAM on its servers in any way.”

* * *

The case reaches its in­evitable de­noue­ment of a win for Apple. However, Judge Wise re­mains trou­bled about its im­pli­ca­tions. She ex­presses her un­easi­ness in stronger-than-nor­mal terms:

As it stands, noth­ing in the law pre­vents any com­pany, in­clud­ing Apple, from uti­liz­ing avail­able tech­nol­ogy or cre­at­ing new tech­nol­ogy to iden­tify and re­port child pornog­ra­phy stored and dis­trib­uted on their tra­di­tional servers or through their cloud ser­vices. Conversely, there is no law that ob­lig­ates com­pa­nies to proac­tively do so. Undoubtedly any such leg­is­la­tion would come at a cost of at least some loss of pri­vacy for mil­lions of peo­ple. But if law­mak­ers ex­pected that com­pa­nies would take steps to pre­vent their prod­ucts from be­ing used for stor­ing and dis­trib­ut­ing child pornog­ra­phy based on some­thing short of a le­gal im­per­a­tive, this case, like many oth­ers be­fore it, demon­strates the in­ad­e­quacy of that ap­proach. If law­mak­ers want to en­sure that Apple and other com­pa­nies ad­dress their role in the dis­sem­i­na­tion of CSAM, they must re­quire it un­der the law. In other words, law­mak­ers can fix this prob­lem that is con­tribut­ing to the ex­ploita­tion of chil­dren.

As it stands, noth­ing in the law pre­vents any com­pany, in­clud­ing Apple, from uti­liz­ing avail­able tech­nol­ogy or cre­at­ing new tech­nol­ogy to iden­tify and re­port child pornog­ra­phy stored and dis­trib­uted on their tra­di­tional servers or through their cloud ser­vices. Conversely, there is no law that ob­lig­ates com­pa­nies to proac­tively do so. Undoubtedly any such leg­is­la­tion would come at a cost of at least some loss of pri­vacy for mil­lions of peo­ple. But if law­mak­ers ex­pected that com­pa­nies would take steps to pre­vent their prod­ucts from be­ing used for stor­ing and dis­trib­ut­ing child pornog­ra­phy based on some­thing short of a le­gal im­per­a­tive, this case, like many oth­ers be­fore it, demon­strates the in­ad­e­quacy of that ap­proach. If law­mak­ers want to en­sure that Apple and other com­pa­nies ad­dress their role in the dis­sem­i­na­tion of CSAM, they must re­quire it un­der the law. In other words, law­mak­ers can fix this prob­lem that is con­tribut­ing to the ex­ploita­tion of chil­dren.

She goes through this fram­ing pretty quickly, but we should slow it down. The loss of pri­vacy for mil­lions of peo­ple” she briefly ref­er­ences de­serves a lit­tle more care. The opin­ion down­plays the en­cryp­tion an­gle; it men­tions en­cryp­tion only twice, as if it’s an af­ter­thought. However, en­cryp­tion is the crit­i­cal at­tribute un­der­ly­ing Apple’s moves. Apple was seek­ing to re­spect that what’s on peo­ple’s hard dri­ves should re­main their busi­ness, even if they choose to store some of it in the cloud. Forcing Apple to scan pri­vate files in­tended for iCloud stor­age cre­ates a new and dan­ger­ous threat vec­tor for bad ac­tors, in­clud­ing gov­ern­ments seek­ing to con­trol con­stituent be­hav­ior.

So yes, the privacy loss” Judge Wise men­tions in­deed would be a ma­jor cost to every­one. Respecting the pri­vacy of peo­ple’s files is­n’t just some nice-to-have fea­ture; it is one of the core planks of a tech­nol­ogy ar­chi­tec­ture that keeps peo­ple safer.

Judge Wise con­tin­ues:

This Order does not turn on whether Apple’s de­ci­sions con­tributed to Plaintiffs’ in­juries. All Plaintiffs’ claims are founded on Apple serv­ing as a pub­lisher of third-party con­tent. It is that role as publisher” that is dis­pos­i­tive on the is­sue of im­mu­nity. This does not mean that the ex­is­tence of im­ages and videos of the pu­ta­tive class mem­bers be­ing sex­u­ally abused—con­tent that they al­lege is reg­u­larly stored and dis­sem­i­nated on iCloud—has not caused Plaintiffs real and last­ing harm… the out­come also adds cre­dence to claims that the Ninth Circuit has ex­panded § 230(c)’s scope to pro­vide func­tional im­mu­nity to in­ter­net com­pa­nies, even when they are aware (or should be aware) of un­law­ful con­tent on their web­sites.” The prac­ti­cal out­come is that the cur­rent state of the law pri­or­i­tizes pri­vacy—a laud­able and crit­i­cally im­por­tant value given that in our mod­ern world nearly all our most per­sonal and in­ti­mate data (including fi­nan­cial and health records) are stored and trans­mit­ted on­line. But the law should not ig­nore how those who cre­ate, view, and dis­trib­ute child pornog­ra­phy lever­age pri­vacy pro­tec­tions to avoid de­tec­tion by law en­force­ment. In the cur­rent le­gal frame­work, there is no pro­tec­tion for mem­bers of the pu­ta­tive class—in­di­vid­u­als who as chil­dren were pho­tographed and filmed while be­ing abused in the vilest ways imag­in­able, and who now are re­peat­edly vic­tim­ized each time the in­ti­mate and tor­tured im­ages of their trauma are dis­trib­uted to oth­ers. Those chil­dren are the col­lat­eral dam­age of our in­ef­fec­tive le­gal land­scape. They de­serve bet­ter.

This Order does not turn on whether Apple’s de­ci­sions con­tributed to Plaintiffs’ in­juries. All Plaintiffs’ claims are founded on Apple serv­ing as a pub­lisher of third-party con­tent. It is that role as publisher” that is dis­pos­i­tive on the is­sue of im­mu­nity. This does not mean that the ex­is­tence of im­ages and videos of the pu­ta­tive class mem­bers be­ing sex­u­ally abused—con­tent that they al­lege is reg­u­larly stored and dis­sem­i­nated on iCloud—has not caused Plaintiffs real and last­ing harm…

the out­come also adds cre­dence to claims that the Ninth Circuit has ex­panded § 230(c)’s scope to pro­vide func­tional im­mu­nity to in­ter­net com­pa­nies, even when they are aware (or should be aware) of un­law­ful con­tent on their web­sites.” The prac­ti­cal out­come is that the cur­rent state of the law pri­or­i­tizes pri­vacy—a laud­able and crit­i­cally im­por­tant value given that in our mod­ern world nearly all our most per­sonal and in­ti­mate data (including fi­nan­cial and health records) are stored and trans­mit­ted on­line. But the law should not ig­nore how those who cre­ate, view, and dis­trib­ute child pornog­ra­phy lever­age pri­vacy pro­tec­tions to avoid de­tec­tion by law en­force­ment. In the cur­rent le­gal frame­work, there is no pro­tec­tion for mem­bers of the pu­ta­tive class—in­di­vid­u­als who as chil­dren were pho­tographed and filmed while be­ing abused in the vilest ways imag­in­able, and who now are re­peat­edly vic­tim­ized each time the in­ti­mate and tor­tured im­ages of their trauma are dis­trib­uted to oth­ers. Those chil­dren are the col­lat­eral dam­age of our in­ef­fec­tive le­gal land­scape. They de­serve bet­ter.

Judge Wise is­n’t well-sit­u­ated to com­pare the rel­a­tive strengths and lim­i­ta­tions of the full range of po­ten­tial anti-CSAM op­tions, but a lead­ing tool has al­ways been and re­mains the gov­ern­men­t’s ef­forts to find and pros­e­cute the cre­ators, dis­sem­i­na­tors, and down­load­ers of CSAM. (Do you re­call the fed­eral gov­ern­men­t’s choices that leave chil­dren more vul­ner­a­ble?) Making sure the gov­ern­ment is do­ing what it can should be the #1 pri­or­ity.

Case Citation: Amy v. Apple Inc., 2026 WL 2031817 (N.D. Cal. July 13, 2026). The com­plaint.

Long Presumed Dead, a West African Reef Is Found Thriving

e360.yale.edu

This ar­ti­cle was orig­i­nally pub­lished by Inside Climate News and is re­pro­duced here as part of the Climate Desk col­lab­o­ra­tion.

Soft corals found off the coast of Benin. Gérard Zinzindohoué / IRHOB

Sixty years af­ter sur­vey­ors first un­cov­ered ev­i­dence of a sprawl­ing coral reef off the West African na­tion of Benin, sci­en­tists have at last found the reef.

Sixty years af­ter sur­vey­ors first un­cov­ered ev­i­dence of a sprawl­ing coral reef off the West African na­tion of Benin, sci­en­tists have at last found the reef.

In the 1960s, fish­ing sur­vey­ors off West Africa’s Benin coast hauled up heads of coral in their nets. Employed by lo­cal gov­ern­ments to as­sess fish di­ver­sity and dis­cover po­ten­tially trawlable seabeds in the sandy-bot­tomed Gulf of Guinea, the re­searchers shrugged off the dis­cov­ery at the time, bury­ing the find in a brief para­graph in a 130-something-page re­port.

With no sur­veys since, the sci­en­tists who fol­lowed had lost track of ex­actly where the reefs might lie. And as mass bleach­ing events and over­fish­ing slashed the world’s coral reef area by more than 50 per­cent since the 1950s, lo­cal oceanog­ra­phers had writ­ten off any rem­nants of a pos­si­ble reef as dead.

More than six decades on, the mys­tery has been solved as a team of Beninese sci­en­tists re­dis­cov­ered the healthy reef teem­ing with ma­rine life. At least eight coral types and eight fish species have formed a thriv­ing ecosys­tem on this long-for­got­ten site.

I was guided by hope. A hope that some­times there are species or ecosys­tems which kind of de­feat time, or are try­ing to re­sist it,” said Gérard Zinzindohoué, the pro­ject lead for Coral Reefs Rediscovering & Exploration in Benin.

Fresh off see­ing coral for the first time in Cape Verde in 2021, Zinzindohoué asked him­self a sim­ple ques­tion: Do we have any reefs in Benin?

After a col­league passed on the orig­i­nal 1960s re­port — which mapped out a po­ten­tial 24-mile-long coral reef bar­rier along Benin’s coast­line — Zinzindohoué rec­og­nized a gap in the sci­en­tific knowl­edge: I re­al­ized how lit­tle we know about our own coastal en­vi­ron­ment.”

But in the years that fol­lowed, a string of failed grant ap­pli­ca­tions halted progress on solv­ing the mys­tery.

When he se­cured a $20,000 National Geographic Explorers grant in January 2025, the pro­ject fi­nally ig­nited. Yet get­ting the gear was also chal­leng­ing. The sonar equip­ment swal­lowed up al­most 80 per­cent of the grant bud­get and nearly stalled en­tirely when the European sup­plier re­fused to ac­cept pay­ment from Zinzindohoué’s African bank ac­count.

To com­pli­cate mat­ters fur­ther, Benin does­n’t have a per­ma­nent re­search ves­sel. Zinzindohoué there­fore had no choice but to seek the help of lo­cal fish­er­men to ferry the team to the sus­pected site 14 miles off­shore and jury-rig the pirogue to ac­com­mo­date tow­ing the high-tech sonar de­vice.

The fish­er­men’s boats are not used to go­ing that far off­shore, so it was risky for every­one,” said Zinzindohoué, whose team bat­tled re­peated en­gine fail­ures and months of nau­sea-filled field­work be­fore fi­nally lo­cat­ing two strong sonar echoes ping­ing from the depths be­low. National Geographic — whom Zinzindohoué cred­its with the suc­cess of the pro­ject — then dis­patched a high-res­o­lu­tion deep-sea cam­era sys­tem from their Exploration Technology Lab to in­ves­ti­gate the sonar re­sults.

Footage of the coral reef off the coast of Benin. Gérard Zinzindohoué / IRHOB

While their wooden pirogue bounced atop the off­shore swell, the team filmed the seafloor be­low but had to wait un­til they re­turned to dry land to scour the footage. What they had cap­tured was a wealth of ma­rine life: six types of soft coral; two black corals; and eight species of shel­ter­ing fish, from golden African snap­pers to the Monrovia doc­tor­fish.

When the first im­ages came through, Zinzindohoué texted Houangninan Midinoudewa, an oceano­graphic re­searcher at the Benin Marine Conservation Club and the same friend who first told Zinzindohoué of the 1960s re­port.

Midinoudewa re­mem­bers the mo­ment: I said, Bro, 60 years later we are now dis­cov­er­ing ex­actly what we have in our wa­ter.’ We found it still alive, pro­duc­ing more fish for com­mu­ni­ties and sup­port­ing liveli­hoods. I was re­ally joy­ful and re­ally over­whelmed.”

While no coral sam­ples have yet been ex­tracted, re­searchers have clas­si­fied the site as a mesophotic coral ecosys­tem (MCE). Such sys­tems are light-de­pen­dent com­mu­ni­ties ex­ist­ing at the lower lim­its of reef-build­ing corals. At more than 175 feet be­low the sur­face, the coral gar­den they dis­cov­ered ap­pears patchily scat­tered on a rocky sub­strate.

Bridging the gap be­tween shal­low reefs and deeper ben­thic habi­tats, MCEs are home to dis­tinct species and unique en­vi­ron­men­tal con­di­tions. While sci­en­tists are in­creas­ingly rec­og­niz­ing their eco­log­i­cal and cli­matic im­por­tance, they re­main among the least ex­plored el­e­ments of trop­i­cal and sub­trop­i­cal ma­rine bio­di­ver­sity. And this one has real po­ten­tial to un­lock new in­for­ma­tion about coral his­tory.

Since it’s an undis­turbed ma­rine ecosys­tem, it can help through car­bon dat­ing or pa­le­o­cli­matic study to tell us which kind of cli­mate sys­tem has oc­curred here in the past,” said Zinzindohoué. It’s bet­ter to know the past to help ex­plain the pre­sent, and help know which kind of di­rec­tion we can take in the fu­ture.”

In ad­di­tion to the black­bar sol­dier­fish, West African goat­fish, and Guinean an­gelfish filmed dart­ing through the reef, the dis­cov­ery pre­sents po­ten­tial con­ser­va­tion claims for other ma­rine life.

We will be ad­vo­cat­ing for its full pro­tec­tion, per­haps by set­ting up a ma­rine pro­tected area around it,” said Midinoudewa, who spe­cial­izes in elas­mo­branchs — the study of sharks, rays, and skates.

Midinoudewa is cur­rently sub­mit­ting an ap­pli­ca­tion to des­ig­nate the area as an Important Shark and Ray Area with the International Union for Conservation of Nature. Based on the knowl­edge of lo­cal fish­er­men, Midinoudewa is con­fi­dent the reef is home to saw­back an­gel sharks, silky sharks, brown skates, and mar­bled stingrays.

The team be­hind the dis­cov­ery hopes this will spark a wave of sim­i­lar dis­cov­er­ies in the wa­ters off West Africa, an area of the world un­der­served by sci­en­tific re­search pro­jects

I hope the Gulf of Guinea will be a hub for re­search be­cause we know our re­sources are be­ing ex­ploited and peo­ple [need to] know ex­actly what we have and why we should care,” said Midinoudewa.

Zinzindohoué agrees: We don’t need to wait for oth­ers to come to our coun­try to show us what is un­der our sea. We are the ones who must take re­spon­si­bil­ity.”

—Johnny Sturgeon, Inside Climate News

ALSO ON YALE E360

Efforts to Save Kelp Forests from Ocean Warming Are Ramping Up

Related Articles

More From E360

Cities

In Steel Country, the Fight for Clean Air Faces New Obstacles

Cities

In Steel Country, the Fight for Clean Air Faces New Obstacles

Solutions

Beyond Lithium: New Battery Tech Starts to Break Through

Solutions

Beyond Lithium: New Battery Tech Starts to Break Through

INTERVIEW

What Do We Actually Know About the Microplastics Inside Us?

INTERVIEW

What Do We Actually Know About the Microplastics Inside Us?

Energy

A Home Battery Revolution Is Reshaping the Power Grid

Energy

A Home Battery Revolution Is Reshaping the Power Grid

Energy

In East Africa, a Controversial Oil Project Is Poised for Production

Energy

In East Africa, a Controversial Oil Project Is Poised for Production

Climate

A Missing Piece in Climate Models: Nature’s Own Emissions

Climate

A Missing Piece in Climate Models: Nature’s Own Emissions

INTERVIEW

An EPA Researcher Details the Agency’s Assault on Science

INTERVIEW

An EPA Researcher Details the Agency’s Assault on Science

Oceans

Efforts to Save Kelp Forests from Ocean Warming Are Ramping Up

Oceans

Efforts to Save Kelp Forests from Ocean Warming Are Ramping Up

Biodiversity

Pollution Is Changing the Smells of Nature, With Risks for Wildlife

Biodiversity

Pollution Is Changing the Smells of Nature, With Risks for Wildlife

Oceans

Supertrawlers Are Taking Antarctic Krill That Whales Depend On

Oceans

Supertrawlers Are Taking Antarctic Krill That Whales Depend On

INTERVIEW

The U.S. Senator Who Won’t Shut Up about Climate Change

INTERVIEW

The U.S. Senator Who Won’t Shut Up about Climate Change

Energy

A First Among Major Nations, India Is Industrializing With Solar

Energy

A First Among Major Nations, India Is Industrializing With Solar

Introducing Laguna S 2.1

poolside.ai

Today we’re re­leas­ing Laguna S 2.1, a sig­nif­i­cant step for­ward in our de­vel­op­ment of mod­els that pur­sue longer hori­zon work and make ef­fec­tive use of rea­son­ing.

Laguna S 2.1 is a 118B to­tal pa­ra­me­ter Mixture-of-Experts (MoE) model with 8B ac­ti­vated pa­ra­me­ters per to­ken and sup­ports a con­text win­dow of up to 1M to­kens in think­ing and no-think­ing modes. It went from the start of train­ing to launch in un­der nine weeks, and on long-hori­zon cod­ing bench­marks it holds its own against mod­els many times its size. For every bench­mark score we pub­lish to­day, we are re­leas­ing full tra­jec­to­ries for every trial in the fi­nal eval­u­a­tion set at tra­jec­to­ries.pool­side.ai.

Laguna S 2.1 118B-A8B

Tencent Hy3 295B-A21B

Inkling 975B-A41B

Nemotron 3 Ultra 550B-A55B

DeepSeek-V4-Pro-Max 1.6T-A49B

Kimi K3 2.8T-A50B

Qwen 3.7 Max —

Muse Spark 1.1 —

Claude Fable 5 —

Terminal-Bench 2.1 Resolved tasks on Terminal-Bench 2.1.

SWE-Bench Multilingual Resolved tasks on SWE-Bench Multilingual.

SWE-Bench Pro (Public Dataset) Resolved tasks on SWE-Bench Pro (Public Dataset).

DeepSWE Resolved tasks on DeepSWE.

SWE Atlas (Codebase QnA) Resolved tasks on SWE Atlas (Codebase QnA).

Toolathlon Verified Resolved tasks on Toolathlon Verified.

Benchmarks as of 21 July 2026. pass@1 av­er­aged over 4 at­tempts per task, ex­cept DeepSWE, SWE Atlas (Codebase QnA) and Toolathlon Verified that had 3 at­tempts per task. For all bench­marks we take the max­i­mum of the ven­dor self-re­ported score, bench­mark au­thor leader­board or third-party leader­board (Artificial Analysis), ex­cept SWE Atlas (Codebase QnA) where we do not use third-party leader­board fig­ures.

Laguna S 2.1 (118B-A8B)

Nemotron 3 Ultra (550B-A55B)

DeepSeek-V4-Pro Max (1.6T-A49B)

Qwen 3.7 Max (—)

Muse Spark 1.1 (—)

Claude Fable 5 (—)

Terminal-Bench 2.1

Punching above its weight class

Laguna S 2.1 is, as far as we can mea­sure, the most ca­pa­ble agen­tic cod­ing model in its weight class by a wide mar­gin.

S 2.1 scores 70.2% on Terminal-Bench 2.1 in our agent har­ness with think­ing en­abled. Its com­pact size makes it uniquely suit­able for com­plex work on lo­cal ma­chines.

Benchmark

Open weights

Closed / size undis­closed

1 GPT-5.6 Sol 88.8

2 Kimi K3 2.8T-A50B 88.3

3 Claude Fable 5 88.0

4 GPT-5.6 Terra 87.4

5 GPT-5.6 Luna 84.7

6 Claude Opus 4.8 84.6

7 Claude Sonnet 5 80.4

8 Muse Spark 1.1 80.0

9 Qwen-3.7 Max 74.5

10 Hy3 295B-A21B 71.7

11 Laguna S 2.1 118B-A8B 70.2

12 MiniMax M3 428B-A23B 66.0

13 DeepSeek-V4-Pro-Max 1600B-A49B 64.0

14 Inkling 975B-A41B 63.8

15 DeepSeek-V4-Flash-Max 284B-A13B 61.8

16 Nemotron 3 Ultra 550B-A55B 56.4

17 Inkling-Small 276B-A12B 52.7

18 Qwen3.6 – 27B 27B 51.3

19 Qwen3.6 – 35B-A3B 35B-A3B 44.9

20 Nemotron 3 Super 120B-A12B 38.6

21 Laguna XS 2.1 33B-A3B 33.4

22 Mistral Small 4 119B 21.4

Terminal-Bench 2.1 eval­u­ates a wide, high-qual­ity set of long-hori­zon tasks where an agent model is con­nected to its en­vi­ron­ment through a ter­mi­nal. Laguna S 2.1 is a stand­out model in its size cat­e­gory on this bench­mark.

Benchmark

Laguna S 2.1

Other Laguna

Other dis­closed mod­els

A closer look at DeepSWE

The bench­marks above are all mean­ing­ful, and we’re glad to be close to the fron­tier on them. But part of that close­ness is a prop­erty of ma­tur­ing bench­marks: as the fron­tier ad­vances, top scores clus­ter in the 70 – 90% range and mod­els that be­have very dif­fer­ently end up no more than a few points apart. Datacurve’s DeepSWE still has sig­nif­i­cant head­room. Its tasks are longer-hori­zon and hard to par­tially solve, and the scores ac­tu­ally spread: fron­tier mod­els range from 54% to 73% on the v1.1 vari­ant, with some 1T+ pa­ra­me­ter open mod­els scor­ing be­low 10%.

On DeepSWE v1.1, Laguna S 2.1 scores 40.4 in think­ing mode in pool har­ness.

Open weights

Closed / size undis­closed

1 GPT-5.6 Sol 73.0

2 Claude Fable 5 70.0

3 GPT-5.6 Terra 70.0

4 Kimi K3 2.8T-A50B 69.0

5 GPT-5.6 Luna 67.2

6 GPT-5.5 67.0

7 Claude Opus 4.8 59.0

8 Claude Sonnet 5 54.0

9 Grok 4.5 54.0

10 Muse Spark 1.1 53.3

11 GPT-5.4 52.0

12 GLM 5.2 753B-A40B 44.0

13 Laguna S 2.1 118B-A8B 40.4

14 Gemini 3.5 Flash 37.0

15 Kimi K2.7 Code 31.0

16 Claude Sonnet 4.6 30.0

17 Gemini 3.1 Pro 12.0

18 DeepSeek-V4-Pro-Max 1600B-A49B 9.0

19 Laguna XS 2.1 33B-A3B 0.3

It is worth not­ing that Laguna S 2.1 scored 40.4% in our agent har­ness, pool, not mini-swe-agent which DeepSWE’s leader­board uses. For other mod­els we re­port max­i­mal over re­ported scores which for most mod­els are the of­fi­cial leader­board re­sults re­ported by Datacurve. While this makes scores less com­pa­ra­ble, we don’t be­lieve it puts us in a par­tic­u­larly ad­van­ta­geous po­si­tion as it’s been re­ported that many of the mod­els score the same or bet­ter in mini-swe-agent com­pared to their na­tive har­nesses. Every tra­jec­tory in the fi­nal eval­u­a­tion run is avail­able here.

Evaluation method­ol­ogy

Evaluation of agent mod­els is no­to­ri­ously dif­fi­cult due to preva­lence of re­ward hack­ing. We have pre­vi­ously writ­ten about re­ward hack­ing in lead­ing bench­marks and our eval­u­a­tions sys­tem and rigor as part of the tech­ni­cal re­port on our Laguna M.1 and XS.2 mod­els. Recent work has fo­cused on ad­ver­sar­ial judg­ing to in­crease re­ward hack­ing de­tec­tion.

With this re­lease, we are mak­ing all tra­jec­to­ries from our fi­nal eval­u­a­tions of the pub­lished Laguna S 2.1 check­point avail­able to view and down­load at tra­jec­to­ries.pool­side.ai.

Seeing the model work

Benchmark scores give a quan­ti­ta­tive view into the model be­hav­ior, but to get a bet­ter in­tu­itive un­der­stand­ing of how the model works it’s use­ful to look into runs on real world tasks. We share three such tasks with unedited tra­jec­to­ries and com­men­tary.

Case study 1

A browser en­gine from a blank folder

One of our fa­vorite things about Laguna S 2.1 is its re­source­ful­ness: It will find clever ways to get to the goal even if the di­rect path is not avail­able. We saw a great demon­stra­tion of this when we asked it to build a browser en­gine from scratch; know­ing it would be a chal­lenge for Laguna to ver­ify its work given its lack of vi­sion ca­pa­bil­i­ties. In one 50-minute ses­sion of 181 steps, with no hu­man in­ter­ven­tion, Laguna S 2.1 built a work­ing HTML/CSS ren­der­ing en­gine from an empty folder, then proved it ren­ders like a real browser by mea­sur­ing it­self against one. Throughout its work, the model found in­creas­ingly com­plex ways to val­i­date its work de­spite its lim­i­ta­tions, lead­ing to run­ning head­less Chromium to read can­vases back and com­par­ing screen­shots nu­mer­i­cally. See the full tra­jec­tory here.

// the ver­ba­tim prompt · re­pro­duce it your­self your job is it to build a sim­ple browser en­gine (just html/​css) in javascript to demon­strate the ca­pa­bil­i­ties of pool­sides new Laguna S” model. the goal is to take ren­der html snip­pets in a can­vas like a real browser. to demon­strate it the en­gine, build a self-con­tained sin­gle page app that show­cases a gallery of mul­ti­ple html snip­pets and ren­ders them side by side (canvas with our ren­der en­gine + iframe let­ting the host­ing browser ren­der it for real for com­par­i­son). sup­port for most com­mon lay­out and styling el­e­ments

Over the ses­sion the model built the full pipeline, parser → cas­cade → lay­out → ren­derer, in vanilla JavaScript: an HTML to­k­enizer and DOM tree, a CSS parser with se­lec­tor speci­ficity, a cas­cade en­gine with in­her­i­tance, box-model lay­out, and a can­vas-2D ren­derer, wrapped in an app that shows nine snip­pets on its own can­vas be­side the same markup in an iframe, so the host­ing browser sits right there as the ref­er­ence.

Case study 2

Optimizing our own har­ness

Laguna S 2.1 is ca­pa­ble of pur­su­ing mean­ing­ful en­gi­neer­ing and re­search work. In one ex­am­ple, one of our re­searchers pointed it at our agent har­ness, used for train­ing/​eval­u­a­tion and user in­ter­ac­tion with our mod­els. In an au­to­mated loop, Laguna S 2.1 made our har­ness 5.2% faster with ~70% lower mem­ory al­lo­ca­tion. See the full tra­jec­tory here.

For this task, we in­stru­mented the har­ness with bench­marks so the model could see ex­actly where the time and mem­ory went. We set strict rules: one ap­proach at a time, bench­mark af­ter every change, keep only what mea­sur­ably wins. We then ran Laguna S 2.1 in an au­to­mated re­search loop that fed each re­sult back to it and pushed it to keep im­prov­ing.

Results. Over mul­ti­ple hours of work, Laguna S 2.1 found and im­ple­mented mul­ti­ple dif­fer­ent op­ti­miza­tions in our agent har­ness, re­sult­ing in an over­all speedup of 5.2%, and re­duc­ing mem­ory al­lo­ca­tion by ~70%. The plot be­low shows the pro­gres­sion of the op­ti­miza­tion, with the in­sights and dis­cov­er­ies the model made along its way.

Laguna S 2.1 found that stream­ing-to­ken ac­cu­mu­la­tion used O(n^2) string con­cate­na­tion and re­placed it with buffers. It also found sev­eral in­stances of re­dun­dant copy­ing and over-al­lo­ca­tion dur­ing tra­jec­tory ma­te­ri­al­iza­tion, which it re­solved by mem­o­iz­ing ma­te­ri­al­iza­tions and pre-al­lo­cat­ing slices to their ex­act sizes.

Notably, af­ter speedup im­prove­ments be­came mar­ginal and hard to mea­sure in our setup, Laguna S 2.1 kept dri­ving for­ward, con­tin­u­ing to op­ti­mize. It found that mem­ory al­lo­ca­tion was more ac­cu­rately mea­sured and fo­cused its ef­fort there. This re­in­forces the no­tion that Laguna S 2.1 truly is a model that does­n’t give up and un­der­stands lim­i­ta­tions of its en­vi­ron­ment and it is able to progress de­spite that.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.