10 interesting stories served every morning and every evening.

openai.com

Just a moment...

ads.openai.com

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

fireworks.ai

K3 is a fron­tier qual­ity open model at a frac­tion of the cost. Even big­ger is that it com­ple­ments Fable pre­dictably, which makes it pos­si­ble to get the high­est qual­ity in­tel­li­gence by rout­ing tasks.

🧭 tl;dr: We ran Kimi K3 (open) against Fable 5 (closed) on ~1,000 agen­tic tasks find­ing:

We achieved 93% ac­cu­racy with rout­ing be­tween K3 and Fable.

Results were up to ~50X more cost ef­fec­tive than Fable alone on long agen­tic loops, and con­sis­tently lower cost across every use case.

How We Measured

We av­er­aged bench­marks, each aimed at a dif­fer­ent kind of work, and ran K3 and Fable 5 through the same har­ness. About 1,030 tasks in all, in real agent loops.

One quick de­f­i­n­i­tion be­fore we get into the re­sults. Oracle rout­ing is a method for mea­sur­ing the best the­o­ret­i­cal per­for­mance by run­ning the task through each model and then pick­ing the cheap­est cor­rect op­tion (the cost/​per­for­mance ceil­ing). In a prac­ti­cal router, you don’t get to run your task against mul­ti­ple mod­els. The router makes a pre­dic­tion of which model has the best cost and qual­ity trade off, but ul­ti­mately it’s a guess.

In this study, or­a­cle rout­ing demon­strated K3 is se­lected for 72 – 96% of tasks. This sug­gests a near-per­fect router might be achiev­able, by learn­ing the dif­fer­ence be­tween day-to-day tasks and the true long tail of fron­tier work. It will re­quire an or­der of mag­ni­tude more rout­ing data, and real world per­for­mance to say de­fin­i­tively.

K3 is a good model.

From a 10,000 foot view, it can be easy to look at both mod­els and call the head-to-head a tie. For ex­am­ple, if you look at SWE, the head­line bench­mark, K3 gets 92.4%, Fable 92.6%. Across the five types of tasks we bench­marked on, the two mod­els tend to stay within a few points of each other, with Fable pulling slightly ahead on its cod­ing-lan­guage breadth (Multi-lang).

It’s easy to stop there and say they’re roughly even”. The news is that they have dis­cretely bet­ter per­for­mance across dif­fer­ent task types.

Two Models is Better than One

If you take a peek in­side a sin­gle bench­mark, there’s more to see than just a top-line ac­cu­racy num­ber. Take SWE, where the two are dead even over­all. If you split SWE by prob­lem do­main you can see where each model shines. K3 is sharpest on sym­bolic math and dev tool­ing; Fable wins on web & data vi­su­al­iza­tion work. The same pat­tern runs through the multi-lan­guage set, where Fable’s breadth car­ries Java, Python and C++, while K3 draws even on JavaScript and Rust.

For long-hori­zon work at a ter­mi­nal, dri­ving a shell and prod­ding at sys­tems across dozens of turns, K3 showed its true col­ors. It cleared a batch of tasks Fable never cracked: a 7z hash, FEAL crypt­analy­sis, leaked se­crets, a live vul­ner­a­bil­ity, run­away async jobs.

K3 can be up to 50x lower cost on Fireworks. 🫳🎤

While qual­ity is a near-tie at a high level, price is­n’t close.

So where’s this huge price gap com­ing from? to­ken pric­ing, prompt caching, and ef­fort-per-task. On SWE for ex­am­ple, K3 works much harder than Fable: roughly 55 turns and 1.3M to­kens a task ver­sus 21 turns and 130K. On the long ter­mi­nal tasks it’s the other way around: Fable is the one that spi­rals, run­ning up 64 turns and 1.5M to­kens (sometimes straight into a time­out).

Prompt caching does most of the work of turn­ing that ef­fort into K3′s price ad­van­tage: even when K3 reads ten times the to­kens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a trade­off. Tasks with ex­tra turns gen­er­ally mean more wall-clock time per run i.e. slower runs. If you need an an­swer in two sec­onds, that mat­ters; if you’re run­ning agents in the back­ground at scale, a bill that’s a frac­tion of the size mat­ters a lot more.

Don’t pick a model. Route.

If you send every task to who­ever han­dles it best, you don’t land some­where be­tween the two mod­els, you land above both.

Per-task rout­ing al­ways out per­forms any sin­gle model run:

The or­a­cle router choose K3, 72 – 96% of task traf­fic. By ar­chi­tect­ing a router this way, you end up with over­all qual­ity above ei­ther model alone at a cost close to just us­ing just the cost-op­ti­mized one.

K3 is cost op­ti­mized on all work types

Put both qual­ity and cost on one plot. K3 in blue lands to the left (the more cost-ef­fec­tive side) of Fable in red in all five task-fam­i­lies. Accuracy trades back and forth: Fable pulls ahead on multi-lan­guage, K3 on ter­mi­nal and le­gal, the rest roughly level.

Single Models Are Wasteful and No Longer SoTA

Kimi K3 + Fable routed to­gether un­locks their best qual­i­ties at the best price.

The sin­gle model provider, to­ken maxxing days, are com­ing to an end. The task-level data says these mod­els are spe­cial­ists at very dif­fer­ent prices. The best AI no longer comes out of a sin­gle lab, it’s a mix­ture of mod­els.

What this means in prac­tice:

Open as the de­fault. A 50x lower cost open model like K3 should be your base case, since the or­a­cle sends it most of the traf­fic any­way.

The router is your moat. A router must be tai­lored to your work­load and learn­ing that task/​model split con­tin­u­ously is the best chance you’ll have at stay­ing ahead.

Free Ink · An open ecosystem for e-readers

freeink.org

'VPNs are lawful technical tools,' says EU Court in landmark Anne Frank copyright ruling

www.techradar.com

VPN providers aren’t li­able for copy­right in­fringe­ment, said the EU Court

The Court ex­plic­itly rec­og­nized VPNs as lawful tech­ni­cal tools”

The case cen­tered on the copy­right bat­tle in­volv­ing Anne Frank’s di­ary

In a ma­jor vic­tory for dig­i­tal rights and com­mon sense, the Court of Justice of the European Union (CJEU) has of­fi­cially cat­e­go­rized Virtual Private Networks (VPNs) as lawful tech­ni­cal tools” while es­tab­lish­ing new bound­aries for on­line copy­right dis­putes.

The land­mark judg­ment — handed down in July 2026 — stems from a com­plex le­gal bat­tle over the on­line pub­li­ca­tion of Anne Frank’s his­tor­i­cal man­u­scripts. At its core, the case forced Europe’s top judges to an­swer a highly tech­ni­cal ques­tion: if a pub­lisher ac­tively tries to block vis­i­tors from a spe­cific coun­try, are they still break­ing the law if a user sneaks past the dig­i­tal bor­der us­ing cir­cum­ven­tion soft­ware?

According to the CJEU, the an­swer is no. As long as a web­site em­ploys state-of-the-art” geo-block­ing tech­nol­ogy, the pub­lisher can­not be held li­able for copy­right in­fringe­ment sim­ply be­cause a de­ter­mined reader de­cides to fire up the best VPN to by­pass the re­stric­tions.

The rul­ing sets a mas­sive prece­dent. It con­firms that copy­right hold­ers can­not point to the mere ex­is­tence of VPNs to claim a web­site’s se­cu­rity mea­sures are com­pletely in­ef­fec­tive.

More im­por­tantly for pri­vacy ad­vo­cates, the court firmly pushed back against the de­mo­niza­tion of pri­vacy soft­ware, ce­ment­ing the le­git­i­mate sta­tus of VPN providers across the European Union.

The Anne Frank dis­pute ex­plained

The EUs top court just con­firmed: Geo-blocking is the copy­right hold­er’s prob­lem, not the VPNs. Providers are not li­able for users by­pass­ing re­stric­tions ⚖️ @torrentfreak https://​t.co/​fLw­b5kYAy1July 17, 2026

The EUs top court just con­firmed: Geo-blocking is the copy­right hold­er’s prob­lem, not the VPNs. Providers are not li­able for users by­pass­ing re­stric­tions ⚖️ @torrentfreak https://​t.co/​fLw­b5kYAy1July 17, 2026

The le­gal tug-of-war be­gan when a coali­tion of Dutch and Belgian aca­d­e­mic in­sti­tu­tions pub­lished a free, schol­arly on­line edi­tion of Anne Frank’s man­u­scripts.

Because copy­right laws are not fully har­mo­nized across Europe, the le­gal sta­tus of the fa­mous di­ary varies by ter­ri­tory. In Belgium and roughly 60 other coun­tries, the writ­ings en­tered the pub­lic do­main years ago. However, in the Netherlands, parts of the text re­main pro­tected by copy­right un­til 2037.

To re­spect this ter­ri­to­r­ial di­vide, the pub­lish­ers hosted the site in Belgium and used geo-block­ing to pre­vent ac­cess from Dutch IP ad­dresses. Visitors from the Netherlands were met with a no­tice ex­plain­ing why they could­n’t en­ter the site.

The Anne Frank Fonds, which holds the Dutch copy­right, sued. They ar­gued that be­cause stan­dard VPNs eas­ily al­low users to mask their true IP ad­dress and spoof a Belgian lo­ca­tion, the schol­arly web­site was ef­fec­tively com­mu­ni­cat­ing the pro­tected work to the Dutch pub­lic.

The CJEU ul­ti­mately re­jected this ar­gu­ment. In its judg­ment, the court noted that while geo-block­ing mea­sures can in­evitably be cir­cum­vented, the pos­si­bil­ity of such cir­cum­ven­tion can­not, in it­self and in all cir­cum­stances, be a de­ci­sive fac­tor in find­ing those mea­sures to be in­ad­e­quate and, there­fore, in­ef­fec­tive.”

Why this mat­ters for the in­ter­net and VPN users

For every­day in­ter­net users, the So What?” of this rul­ing is deeply re­as­sur­ing. It val­i­dates that us­ing a VPN to en­crypt your on­line traf­fic, hide your IP ad­dress, or by­pass dig­i­tal bor­ders is a le­git­i­mate use of con­sumer tech­nol­ogy.

The judges ex­plic­itly shielded VPN com­pa­nies from col­lat­eral dam­age in piracy law­suits, a topic that has sparked in­tense de­bate among European ISPs and right­sh­old­ers. The court specif­i­cally ar­gued that the provider of a VPN or sim­i­lar ser­vices is not li­able for users by­pass­ing re­stric­tions.

By plac­ing the le­gal bur­den on pub­lish­ers to main­tain state-of-the-art” dig­i­tal fences, rather than de­mand­ing ab­solute, im­pos­si­ble per­fec­tion, the EU has drawn a prag­matic line in the sand.

Publishers aren’t ex­pected to build un­hack­able walls, and VPN providers aren’t re­spon­si­ble for the ac­tions of users who climb over them. Ultimately, this rul­ing proves that the bor­der­less in­ter­net can still co­ex­ist with ter­ri­to­r­ial copy­right laws, pro­vided every­one uses the right tech­ni­cal safe­guards.

Follow TechRadar on Google News and add us as a pre­ferred source to get our ex­pert news, re­views, and opin­ion in your feeds. Make sure to click the Follow but­ton!

OverpAId — Fire Your CEO. Hire The Future.

overpaid.lol

Introducing the world’s first Chief Executive Replacement Engine

Your CEO costs $22,000,000 a year.We cost $4,699. Once.

OverpAId is an Artificial Intelligence built from the ground up to do your CEOs en­tire job — strat­egy, vision,” mo­ti­va­tional all-hands emails — bet­ter, faster, and with­out ever once ask­ing the board for a big­ger jet. Runs on a sin­gle desk-sized AI com­puter. Real hard­ware, real price, zero mys­tique.

No golden para­chute re­quired. No sev­er­ance pack­age. No emo­tional sup­port LinkedIn post.

$18.9M

Average S&P 500 CEO to­tal com­pen­sa­tion, in a good year for every­one ex­cept the work­force

290 : 1

Typical CEO-to-median-worker pay ra­tio at large pub­lic com­pa­nies

24/7/365

OverpAId’s up­time. Your CEOs up­time: some­where be­tween at Davos” and processing.”

0

Corporate re­treats OverpAId needs in Aspen to reconnect with the mis­sion”

As Featured In (Not Really)

FORBES (Nobody Reads It) THE WALL STREET JOURNAL (Wouldn’t Say Jack) TECHCRUNCH (Crunched By Layoffs) BLOOMBERG (Allegedly) FAST COMPANY (Slow, Actually)

Live Activity

What’s Happening Right Now

A com­pletely real-time, def­i­nitely-not-ran­dom­ized feed of ex­ec­u­tive ac­tiv­ity vs. OverpAId ac­tiv­ity.

The Problem

Let’s talk about the ele­phant in the board­room.

Over the last four decades, CEO pay at the largest com­pa­nies has grown roughly 1,000%+, while typ­i­cal worker pay has crawled for­ward at a frac­tion of that rate — de­spite worker pro­duc­tiv­ity climb­ing the en­tire time. Somewhere along the way, the story be­came: pay the per­son at the top enough, and the value will trickle down to every­one else. It has­n’t. It does­n’t. It never re­ally did.

Meanwhile, the ac­tual day-to-day de­ci­sions dri­ving most com­pa­nies — re­source al­lo­ca­tion, pat­tern recog­ni­tion across moun­tains of data, should we do the thing the data clearly says to do” — are ex­actly the kind of de­ci­sions soft­ware has got­ten ex­tremely good at. So we built the ob­vi­ous, ex­tremely petty, deeply sat­is­fy­ing next step.

Meanwhile, Back At The Earnings Call

The Layoff Two-Step.

Across tech, re­tail, me­dia, lo­gis­tics, and fi­nance, a very spe­cific script has taken over: cut a wave of front­line and mid-level jobs, say AI as many times as pos­si­ble in the press re­lease, and qui­etly reroute the freed-up pay­roll into GPU leases, data cen­ter build­outs, and agentic AI li­cens­ing fees. The work­force gets optimized” to pay for the AI. The AI then gets credit for re­plac­ing the work­force. And the ex­ec­u­tive team that ap­proved both line items — the lay­offs and the AI bud­get — stays ex­actly where it was, at ex­actly its pre­vi­ous salary. In tech alone, well over half a mil­lion jobs have been cut across suc­ces­sive waves of these an­nounce­ments, a grow­ing share of them ex­plic­itly at­trib­uted to AI-driven ef­fi­ciency,” while ag­gre­gate CEO pay at the same com­pa­nies kept climb­ing right along­side the AI cap­i­tal ex­pen­di­ture.

Humbled and hon­ored to step into this role at such a piv­otal mo­ment for our com­pany. I’ve spent the last two weeks lis­ten­ing — to cus­tomers, to our board, to my­self — and I can say with to­tal con­vic­tion: our peo­ple are our great­est as­set. (This post was sched­uled be­fore this morn­ing’s an­nounce­ment. We are aware. We are mov­ing for­ward.)

💜 2,847   💬 412 (mostly Glassdoor re­views)   🔁 89

Executive Leadership 0% re­duc­tion

Senior Directors -8%

Middle Management -22%

Frontline & Support Staff -34%

The only layer im­mune to efficiency” is the one that ap­proves it.

Here’s the part that should bother you more than the lay­offs them­selves: these com­pa­nies al­ready be­lieve an AI agent can do a per­son’s job well enough to elim­i­nate the po­si­tion en­tirely. They just keep draw­ing that line one layer too low. If an agent can run a sup­port queue, man­age a sup­ply chain, or ship half a code­base, it can ob­vi­ously han­dle approve the re­org” and read the an­a­lyst note out loud on the earn­ings call.” Somehow that layer never makes the slide. That’s not a co­in­ci­dence. That’s the de­sign.

OverpAId flips the script on the one line item that’s al­ways ex­empt from the AI trans­for­ma­tion every­one else just got handed. Finally: a work­force re­duc­tion, funded by an AI ini­tia­tive, that ac­tu­ally starts at the top.

Meanwhile, Back At The Real Estate Portfolio

The Return-To-Office Two-Step

A re­mark­ably con­sis­tent pat­tern: com­pa­nies spend years prov­ing re­mote teams ship fine, then man­date a re­turn to of­fice cit­ing culture” and collaboration” — on a time­line that tracks sus­pi­ciously well with lease re­newals, down­town va­cancy head­lines, and com­mer­cial prop­erty val­u­a­tions, and not at all with any ac­tual drop in out­put. Office va­cancy in ma­jor U.S. down­towns has hov­ered near 19 – 20% for years, man­dates in­cluded. The desks aren’t empty be­cause peo­ple won’t come back. They were never go­ing to be full enough to mat­ter.

The Offsite That Prompted All This (Itemized)

Private jet char­ter, round trip: $340,000

3-night re­sort block, ex­ec­u­tive suites: $128,000

Team align­ment” mixol­ogy class: $6,200

Keynote speaker (was on a pod­cast once): $75,000

Branded fleece vests, size: only Medium: $14,000

The tell is al­ways the same: no com­pany has ever man­dated a re­turn to of­fice be­cause re­mote pro­duc­tiv­ity got worse. They man­dated it be­cause an as­set on the books needed a pulse in the lobby to jus­tify its val­u­a­tion — and mov­ing four thou­sand em­ploy­ees turned out to be eas­ier than ad­mit­ting a fif­teen-year lease was a mis­take.

OverpAId has no com­mute, no badge, and no as­signed desk — and, not co­in­ci­den­tally, no opin­ion what­so­ever about any­one’s down­town park­ing garage rev­enue.

For Your Next All-Hands

Corporate Jargon Bingo

Print this out. Bring it to your next town hall, standup, or quick sync.” OverpAId has never once gen­er­ated any of the fol­low­ing phrases un­prompted. Humans — usu­ally the ones with the biggest pack­ages — still do, con­stantly, ap­par­ently for free.

Circle Back

Move The Needle

Low-Hanging Fruit

Boil The Ocean

Bandwidth

Take This Offline

Double-Click On That

North Star

Paradigm Shift

Growth Hacking

Best-In-Class

Value-Add

Synergy (Free Space)

Deep Dive

Culture Fit

Think Outside The Box

Actionable Insights

Alignment

Bleeding Edge

Disruptive Innovation

Level Set

Ideate

Operationalize

Stakeholder Buy-In

Blue Ocean Strategy

Hard Stop

10x

Unicorn

TAM

Product-Market Fit

Down Round

Runway

Blitzscale

Vesting Cliff

Overheard, ver­ba­tim, in an ac­tual meet­ing: Let’s cir­cle back of­fline if you have the spare cy­cles so we can hop on a quick call for a touch­point.” Translation: email me later. Six buzz­words. One sen­tence. Zero in­for­ma­tion trans­ferred. OverpAId would have just said that.

Five in a row and, legally, you’re al­lowed to leave the meet­ing. (We checked. You’re not. But you should be.)

An Important Distinction

Not every job is a spread­sheet in a trench coat.

Before you print this out and sta­ple it to your nurse’s badge — no. OverpAId is not com­ing for the peo­ple who do the ac­tual work. It is com­ing, with ex­treme prej­u­dice, for ex­actly one cat­e­gory of job: the one that spent the last forty years in­sist­ing every­one else’s job was re­place­able.

🛡️ Cannot Be Abstracted Away

Ask an AI to do these and it will, at best, pro­duce a very con­fi­dent hal­lu­ci­na­tion.

🩺 A nurse catch­ing a pa­tien­t’s con­di­tion change be­fore the chart does

🏗️ An en­gi­neer de­bug­ging a live out­age at 3 a.m., be­cause the fix can’t wait for sprint plan­ning

🚑 A doc­tor mak­ing a call in the ER with in­com­plete in­for­ma­tion and a body on the table

👩‍🏫 A teacher notic­ing which kid in the back row stopped rais­ing their hand

🔧 A tech­ni­cian whose hands ac­tu­ally touch the ma­chine that ac­tu­ally breaks

🚒 Anyone whose job in­volves a body, a pa­tient, a cus­tomer, or a dead­line mea­sured in min­utes

🎯 Extremely, Suspiciously Abstractable

Ask an AI to do these and, un­com­fort­ably, it al­ready can. Better.

📈 Reading a re­port some­one else wrote, then re­peat­ing the con­clu­sion in a town hall

✅ Approving a de­ci­sion your own data team qui­etly made three weeks ago

🎤 Taking credit for quar­terly num­bers on an earn­ings call

📧 Replying let’s cir­cle back” to an email that needed a yes or no

Judge approves a $1.5B Anthropic settlement over books used to train Claude | AP News

apnews.com

SAN FRANCISCO (AP) — A fed­eral judge has ap­proved a $1.5 bil­lion copy­right set­tle­ment in which ar­ti­fi­cial in­tel­li­gence com­pany Anthropic will pay thou­sands of au­thors about $3,000 per book af­ter us­ing pi­rated copies of their works to train its Claude chat­bot.

District Judge Araceli Martínez-Olguín said in a Monday rul­ing that the class-ac­tion set­tle­ment pro­vides meaningful re­lief” to af­fected au­thors and pub­lish­ers.

About 91% of the more than 482,000 books cov­ered by the rul­ing have been claimed by au­thors or pub­lish­ers who are now due pay­ment.

Plaintiff at­tor­ney Justin Nelson said in a state­ment that the set­tle­ment was the largest known copy­right re­cov­ery in his­tory. We look for­ward to mak­ing dis­tri­b­u­tions to the Class as promptly as pos­si­ble.”

U.S. District Judge William Alsup is­sued the pre­lim­i­nary ap­proval in San Francisco fed­eral court last September and has since re­tired. Alsup had dealt the case a mixed rul­ing last sum­mer, find­ing that train­ing AI chat­bots on copy­righted books was­n’t il­le­gal but that Anthropic wrong­fully ac­quired mil­lions of books through pi­rate web­sites.

Anthropic’s deputy gen­eral coun­sel, Aparna Sridhar, high­lighted that rul­ing Friday as a land­mark show­ing that train­ing AI on books is fair use un­der copy­right law.”

We are pleased that more than 91% of au­thors and pub­lish­ers cov­ered by the set­tle­ment have claimed their share of the pay­ment, and we’re look­ing for­ward to bring­ing this mat­ter to a close,” Sridhar said in a writ­ten state­ment.

Bestselling thriller nov­el­ist Andrea Bartz first brought the suit with two other au­thors in 2024. It’s the first ma­jor set­tle­ment in dozens of AI copy­right law­suits that are still work­ing their way through courts.

Introducing Laguna S 2.1

poolside.ai

Today we’re re­leas­ing Laguna S 2.1, a sig­nif­i­cant step for­ward in our de­vel­op­ment of mod­els that pur­sue longer hori­zon work and make ef­fec­tive use of rea­son­ing.

Laguna S 2.1 is a 118B to­tal pa­ra­me­ter Mixture-of-Experts (MoE) model with 8B ac­ti­vated pa­ra­me­ters per to­ken and sup­ports a con­text win­dow of up to 1M to­kens in think­ing and no-think­ing modes. It went from the start of train­ing to launch in un­der nine weeks, and on long-hori­zon cod­ing bench­marks it holds its own against mod­els many times its size. For every bench­mark score we pub­lish to­day, we are re­leas­ing full tra­jec­to­ries for every trial in the fi­nal eval­u­a­tion set at tra­jec­to­ries.pool­side.ai.

Laguna S 2.1 118B-A8B

Tencent Hy3 295B-A21B

Inkling 975B-A41B

Nemotron 3 Ultra 550B-A55B

DeepSeek-V4-Pro-Max 1.6T-A49B

Kimi K3 2.8T-A50B

Qwen 3.7 Max —

Muse Spark 1.1 —

Claude Fable 5 —

Terminal-Bench 2.1 Resolved tasks on Terminal-Bench 2.1.

SWE-Bench Multilingual Resolved tasks on SWE-Bench Multilingual.

SWE-Bench Pro (Public Dataset) Resolved tasks on SWE-Bench Pro (Public Dataset).

DeepSWE Resolved tasks on DeepSWE.

SWE Atlas (Codebase QnA) Resolved tasks on SWE Atlas (Codebase QnA).

Toolathlon Verified Resolved tasks on Toolathlon Verified.

Benchmarks as of 21 July 2026. pass@1 av­er­aged over 4 at­tempts per task, ex­cept DeepSWE, SWE Atlas (Codebase QnA) and Toolathlon Verified that had 3 at­tempts per task. For all bench­marks we take the max­i­mum of the ven­dor self-re­ported score, bench­mark au­thor leader­board or third-party leader­board (Artificial Analysis), ex­cept SWE Atlas (Codebase QnA) where we do not use third-party leader­board fig­ures.

Laguna S 2.1 (118B-A8B)

Nemotron 3 Ultra (550B-A55B)

DeepSeek-V4-Pro Max (1.6T-A49B)

Qwen 3.7 Max (—)

Muse Spark 1.1 (—)

Claude Fable 5 (—)

Terminal-Bench 2.1

Punching above its weight class

Laguna S 2.1 is, as far as we can mea­sure, the most ca­pa­ble agen­tic cod­ing model in its weight class by a wide mar­gin.

S 2.1 scores 70.2% on Terminal-Bench 2.1 in our agent har­ness with think­ing en­abled. Its com­pact size makes it uniquely suit­able for com­plex work on lo­cal ma­chines.

Benchmark

Open weights

Closed / size undis­closed

1 GPT-5.6 Sol 88.8

2 Kimi K3 2.8T-A50B 88.3

3 Claude Fable 5 88.0

4 GPT-5.6 Terra 87.4

5 GPT-5.6 Luna 84.7

6 Claude Opus 4.8 84.6

7 Claude Sonnet 5 80.4

8 Muse Spark 1.1 80.0

9 Qwen-3.7 Max 74.5

10 Hy3 295B-A21B 71.7

11 Laguna S 2.1 118B-A8B 70.2

12 MiniMax M3 428B-A23B 66.0

13 DeepSeek-V4-Pro-Max 1600B-A49B 64.0

14 Inkling 975B-A41B 63.8

15 DeepSeek-V4-Flash-Max 284B-A13B 61.8

16 Nemotron 3 Ultra 550B-A55B 56.4

17 Inkling-Small 276B-A12B 52.7

18 Qwen3.6 – 27B 27B 51.3

19 Qwen3.6 – 35B-A3B 35B-A3B 44.9

20 Nemotron 3 Super 120B-A12B 38.6

21 Laguna XS 2.1 33B-A3B 33.4

22 Mistral Small 4 119B 21.4

Terminal-Bench 2.1 eval­u­ates a wide, high-qual­ity set of long-hori­zon tasks where an agent model is con­nected to its en­vi­ron­ment through a ter­mi­nal. Laguna S 2.1 is a stand­out model in its size cat­e­gory on this bench­mark.

Benchmark

Laguna S 2.1

Other Laguna

Other dis­closed mod­els

A closer look at DeepSWE

The bench­marks above are all mean­ing­ful, and we’re glad to be close to the fron­tier on them. But part of that close­ness is a prop­erty of ma­tur­ing bench­marks: as the fron­tier ad­vances, top scores clus­ter in the 70 – 90% range and mod­els that be­have very dif­fer­ently end up no more than a few points apart. Datacurve’s DeepSWE still has sig­nif­i­cant head­room. Its tasks are longer-hori­zon and hard to par­tially solve, and the scores ac­tu­ally spread: fron­tier mod­els range from 54% to 73% on the v1.1 vari­ant, with some 1T+ pa­ra­me­ter open mod­els scor­ing be­low 10%.

On DeepSWE v1.1, Laguna S 2.1 scores 40.4 in think­ing mode in pool har­ness.

Open weights

Closed / size undis­closed

1 GPT-5.6 Sol 73.0

2 Claude Fable 5 70.0

3 GPT-5.6 Terra 70.0

4 Kimi K3 2.8T-A50B 69.0

5 GPT-5.6 Luna 67.2

6 GPT-5.5 67.0

7 Claude Opus 4.8 59.0

8 Claude Sonnet 5 54.0

9 Grok 4.5 54.0

10 Muse Spark 1.1 53.3

11 GPT-5.4 52.0

12 GLM 5.2 753B-A40B 44.0

13 Laguna S 2.1 118B-A8B 40.4

14 Gemini 3.5 Flash 37.0

15 Kimi K2.7 Code 31.0

16 Claude Sonnet 4.6 30.0

17 Gemini 3.1 Pro 12.0

18 DeepSeek-V4-Pro-Max 1600B-A49B 9.0

19 Laguna XS 2.1 33B-A3B 0.3

It is worth not­ing that Laguna S 2.1 scored 40.4% in our agent har­ness, pool, not mini-swe-agent which DeepSWE’s leader­board uses. For other mod­els we re­port max­i­mal over re­ported scores which for most mod­els are the of­fi­cial leader­board re­sults re­ported by Datacurve. While this makes scores less com­pa­ra­ble, we don’t be­lieve it puts us in a par­tic­u­larly ad­van­ta­geous po­si­tion as it’s been re­ported that many of the mod­els score the same or bet­ter in mini-swe-agent com­pared to their na­tive har­nesses. Every tra­jec­tory in the fi­nal eval­u­a­tion run is avail­able here.

Evaluation method­ol­ogy

Evaluation of agent mod­els is no­to­ri­ously dif­fi­cult due to preva­lence of re­ward hack­ing. We have pre­vi­ously writ­ten about re­ward hack­ing in lead­ing bench­marks and our eval­u­a­tions sys­tem and rigor as part of the tech­ni­cal re­port on our Laguna M.1 and XS.2 mod­els. Recent work has fo­cused on ad­ver­sar­ial judg­ing to in­crease re­ward hack­ing de­tec­tion.

With this re­lease, we are mak­ing all tra­jec­to­ries from our fi­nal eval­u­a­tions of the pub­lished Laguna S 2.1 check­point avail­able to view and down­load at tra­jec­to­ries.pool­side.ai.

Seeing the model work

Benchmark scores give a quan­ti­ta­tive view into the model be­hav­ior, but to get a bet­ter in­tu­itive un­der­stand­ing of how the model works it’s use­ful to look into runs on real world tasks. We share three such tasks with unedited tra­jec­to­ries and com­men­tary.

Case study 1

A browser en­gine from a blank folder

One of our fa­vorite things about Laguna S 2.1 is its re­source­ful­ness: It will find clever ways to get to the goal even if the di­rect path is not avail­able. We saw a great demon­stra­tion of this when we asked it to build a browser en­gine from scratch; know­ing it would be a chal­lenge for Laguna to ver­ify its work given its lack of vi­sion ca­pa­bil­i­ties. In one 50-minute ses­sion of 181 steps, with no hu­man in­ter­ven­tion, Laguna S 2.1 built a work­ing HTML/CSS ren­der­ing en­gine from an empty folder, then proved it ren­ders like a real browser by mea­sur­ing it­self against one. Throughout its work, the model found in­creas­ingly com­plex ways to val­i­date its work de­spite its lim­i­ta­tions, lead­ing to run­ning head­less Chromium to read can­vases back and com­par­ing screen­shots nu­mer­i­cally. See the full tra­jec­tory here.

// the ver­ba­tim prompt · re­pro­duce it your­self your job is it to build a sim­ple browser en­gine (just html/​css) in javascript to demon­strate the ca­pa­bil­i­ties of pool­sides new Laguna S” model. the goal is to take ren­der html snip­pets in a can­vas like a real browser. to demon­strate it the en­gine, build a self-con­tained sin­gle page app that show­cases a gallery of mul­ti­ple html snip­pets and ren­ders them side by side (canvas with our ren­der en­gine + iframe let­ting the host­ing browser ren­der it for real for com­par­i­son). sup­port for most com­mon lay­out and styling el­e­ments

Over the ses­sion the model built the full pipeline, parser → cas­cade → lay­out → ren­derer, in vanilla JavaScript: an HTML to­k­enizer and DOM tree, a CSS parser with se­lec­tor speci­ficity, a cas­cade en­gine with in­her­i­tance, box-model lay­out, and a can­vas-2D ren­derer, wrapped in an app that shows nine snip­pets on its own can­vas be­side the same markup in an iframe, so the host­ing browser sits right there as the ref­er­ence.

Case study 2

Optimizing our own har­ness

Laguna S 2.1 is ca­pa­ble of pur­su­ing mean­ing­ful en­gi­neer­ing and re­search work. In one ex­am­ple, one of our re­searchers pointed it at our agent har­ness, used for train­ing/​eval­u­a­tion and user in­ter­ac­tion with our mod­els. In an au­to­mated loop, Laguna S 2.1 made our har­ness 5.2% faster with ~70% lower mem­ory al­lo­ca­tion. See the full tra­jec­tory here.

For this task, we in­stru­mented the har­ness with bench­marks so the model could see ex­actly where the time and mem­ory went. We set strict rules: one ap­proach at a time, bench­mark af­ter every change, keep only what mea­sur­ably wins. We then ran Laguna S 2.1 in an au­to­mated re­search loop that fed each re­sult back to it and pushed it to keep im­prov­ing.

Results. Over mul­ti­ple hours of work, Laguna S 2.1 found and im­ple­mented mul­ti­ple dif­fer­ent op­ti­miza­tions in our agent har­ness, re­sult­ing in an over­all speedup of 5.2%, and re­duc­ing mem­ory al­lo­ca­tion by ~70%. The plot be­low shows the pro­gres­sion of the op­ti­miza­tion, with the in­sights and dis­cov­er­ies the model made along its way.

Laguna S 2.1 found that stream­ing-to­ken ac­cu­mu­la­tion used O(n^2) string con­cate­na­tion and re­placed it with buffers. It also found sev­eral in­stances of re­dun­dant copy­ing and over-al­lo­ca­tion dur­ing tra­jec­tory ma­te­ri­al­iza­tion, which it re­solved by mem­o­iz­ing ma­te­ri­al­iza­tions and pre-al­lo­cat­ing slices to their ex­act sizes.

Notably, af­ter speedup im­prove­ments be­came mar­ginal and hard to mea­sure in our setup, Laguna S 2.1 kept dri­ving for­ward, con­tin­u­ing to op­ti­mize. It found that mem­ory al­lo­ca­tion was more ac­cu­rately mea­sured and fo­cused its ef­fort there. This re­in­forces the no­tion that Laguna S 2.1 truly is a model that does­n’t give up and un­der­stands lim­i­ta­tions of its en­vi­ron­ment and it is able to progress de­spite that.

LG to Ban Residential Proxies from Smart TV Apps

krebsonsecurity.com

The home ap­pli­ance gi­ant LG Electronics USA said this week it plans to sus­pend any apps built for its smart TVs that turn one’s tele­vi­sion into an al­ways-on res­i­den­tial proxy node. The move comes less than a month af­ter re­searchers found that more than 42 per­cent of games and other apps avail­able for down­load on LGs we­bOS store al­low un­known third-par­ties to route their Internet traf­fic through a user’s TV.

Proxy SDK preva­lence among smart TV apps for LG (webOS) and Samsung (Tizen OS) tele­vi­sions. Image: Spur.us.

On July 2, we fea­tured re­search by the se­cu­rity firm Spur that ex­am­ined the preva­lence of res­i­den­tial proxy soft­ware de­vel­op­ment kits (SDKs) in smart TV apps. Spur found more than 42 per­cent of apps avail­able for down­load on LG smart TVs in­clude SDKs that turn one’s tele­vi­sion in a proxy node in­def­i­nitely, and that more than a quar­ter of the apps made for Samsung’s Tizen op­er­at­ing sys­tem had sim­i­lar res­i­den­tial proxy com­po­nents.

Responding to ques­tions about Spur’s re­search, LG Senior Vice President John Taylor told KrebsOnSecurity the com­pany was work­ing with app de­vel­op­ers to re­move the res­i­den­tial proxy op­tion from their apps on the we­bOS plat­form. Developers that fail to com­ply, he said, will find their apps sus­pended.

A res­i­den­tial proxy net­work is not an in­tended use for LG smart TVs, and LG Electronics is work­ing with de­vel­op­ers to re­move the res­i­den­tial proxy op­tion from their apps on the we­bOS plat­form,” Taylor said. If this op­tion is not re­moved, these apps will be sus­pended.”

Taylor said LG is com­mit­ted to keep­ing res­i­den­tial proxy net­works out of its smart TV apps go­ing for­ward, and that the com­pa­ny’s re­view of those apps is well un­der­way now.”

As part of our on­go­ing ef­forts to en­hance plat­form qual­ity and the user ex­pe­ri­ence, LG will con­tinue to strengthen our eval­u­a­tion process for de­vel­oper-sub­mit­ted apps, in­clud­ing those that in­cor­po­rate res­i­den­tial proxy SDKs,” Taylor wrote in an emailed state­ment.

App mak­ers look­ing for ways to mon­e­tize their cre­ations can turn to res­i­den­tial proxy providers, which pay de­vel­op­ers to in­clude SDKs that turn the user’s de­vice into a res­i­den­tial proxy node that is rented to pay­ing cus­tomers. In the case of LG and Samsung smart TVs, Spur found res­i­den­tial proxy SDKs bun­dled with every­thing from sim­ple games like Pac-Man to screen­savers and file util­i­ties.

A Pac-Man smart TV app from Bright Data of­fers users the choice be­tween view­ing ads in the game or agree­ing to al­low their TV to serve as a res­i­den­tial proxy node. Image: Spur.us.

Spur’s re­port found the res­i­den­tial proxy net­work Bright Data ac­counted for a ma­jor­ity of proxy SDKs across both Samsung and LG smart TVs. In a state­ment shared with KrebsOnSecurity, Bright Data said its net­work is built on con­sent and re­spon­si­bil­ity and op­er­ates by LG and Samsung terms.

Every peer opts in through a ded­i­cated screen and re­ceives value in re­turn; every cus­tomer is vet­ted, and our prac­tices have now un­der­gone a sec­ond in­de­pen­dent au­dit by PwC,” the state­ment reads. We re­main com­mit­ted to an open, trans­par­ent in­ter­net where le­git­i­mate busi­nesses, re­searchers, and in­sti­tu­tions can re­spon­si­bly ac­cess data that lives in the pub­lic do­main.”

Bright Data and other proxy providers named in Spur’s re­port all say they fol­low rig­or­ous know-your-cus­tomer processes to val­i­date le­git­i­mate uses of their ser­vices, which is of­ten heav­ily tied to con­tent-scrap­ing ac­tiv­i­ties by said cus­tomers. The proxy com­pa­nies also say they in­cor­po­rate tech­no­log­i­cal coun­ter­mea­sures to pre­vent proxy ser­vice cus­tomers from be­ing able to in­ter­act with and con­trol other de­vices on the proxy user’s lo­cal net­work.

Spur ar­gues the prob­lem is not that res­i­den­tial proxy net­works ex­ist, but rather that they are be­ing em­bed­ded at scale in de­vices that most con­sumers do not think of as com­put­ers and are not equipped to au­dit.

A one-time con­sent prompt buried in a TV app is not a sub­sti­tute for mean­ing­ful trans­parency, on­go­ing con­trol, and plat­form over­sight,” Spur’s Trevor Sutter wrote. The risk is am­pli­fied when con­sent comes from in­di­vid­u­als within the house­hold who use the de­vice but should­n’t give con­sent, such as mi­nors.”

LGs an­nounce­ment that it is culling res­i­den­tial proxy SDKs from its app store is wel­come news, but the com­pany re­cently came un­der fire for an­other ques­tion­able part­ner­ship: Pimping McAfee se­cu­rity prod­ucts via soft­ware dri­vers in­cluded in its high-end LCD mon­i­tors.

Earlier this week, the Youtube chan­nel Gamers Nexus showed that cer­tain LG LCD mon­i­tors will au­to­mat­i­cally in­stall an app that pro­motes paid McAfee an­tivirus sub­scrip­tions, and that the app ar­rives through Windows Update with­out an ap­proval prompt.

Update, July 22, 1:06 p.m. ET: Added state­ment from Bright Data.

Bento Slides

bento.page

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.