10 interesting stories served every morning and every evening.

Our position on open-weights models

www.anthropic.com

A post by Dario Amodei, Anthropic CEO

Over the last few days there has been a lot of dis­cus­sion about open-weights mod­els, es­pe­cially those from China. Reports sug­gest that some US of­fi­cials are con­sid­er­ing ban­ning the use of Chinese open-weights mod­els by US com­pa­nies. In re­sponse, many tech com­pa­nies have signed a let­ter sup­port­ing open-weights mod­els, and some peo­ple have even ac­cused Anthropic of want­ing to ban open-weights mod­els as a means of pro­tect­ing our busi­ness. Anyone who has read my past writ­ing should know that I don’t re­gard such bans as a use­ful mea­sure, but let me state it clearly so that there is no doubt: Anthropic has never ad­vo­cated for a ban on open-weights mod­els.

Open-weights mod­els that don’t have dan­ger­ous ca­pa­bil­i­ties are a pub­lic good: they don’t cost any­thing be­sides the com­pute needed to run them, and they pro­vide value to busi­nesses, de­vel­op­ers, and re­searchers.

Protectionist bans would not ad­dress my most se­ri­ous na­tional se­cu­rity con­cerns. Specifically, I am wor­ried about two night­mare sce­nar­ios. I laid these out in my es­say The Adolescence of Technology six months ago1, and have held these po­si­tions con­sis­tently for many years:

My pri­mary con­cern is the risk that au­thor­i­tar­ian gov­ern­ments—not solely the Chinese Communist Party (CCP), al­though the CCP is clearly the most ca­pa­ble threat—build AI mod­els that are more pow­er­ful than those built by the US, and use them to achieve per­ma­nent mil­i­tary su­pe­ri­or­ity or per­pe­trate in­cred­i­bly deep re­pres­sion of their own peo­ple. This con­cern is widely shared within the US gov­ern­ment: Vice President Vance warned in Paris last year that authoritarian regimes have stolen and used AI to strengthen their mil­i­tary, in­tel­li­gence, and sur­veil­lance ca­pa­bil­i­ties,” and the Intelligence Community’s 2026 Annual Threat Assessment found that other global pow­ers’ ro­bust progress in AI is chal­leng­ing US eco­nomic com­pet­i­tive­ness and na­tional se­cu­rity ad­van­tages.” It is ir­rel­e­vant whether these mod­els are re­leased with open weights, and cer­tainly ir­rel­e­vant whether they are used by US busi­nesses. In fact, the most dan­ger­ous model may be one that is trained in se­cret and handed only to the People’s Liberation Army for use in drones and the Ministry of State Security for sur­veil­lance and re­pres­sion.

My sec­ondary con­cern is the risk that pow­er­ful AI mod­els may be mis­used to carry out cy­ber­at­tacks or bi­o­log­i­cal at­tacks, and may have se­ri­ous align­ment prob­lems. Open-weights mod­els—it does not mat­ter whether they come from China or any­where else—do po­ten­tially pre­sent a higher risk than closed mod­els, be­cause it is very dif­fi­cult to ap­ply guardrails to them or mon­i­tor their us­age, and once weights are re­leased they can­not be with­drawn2. But ban­ning the use of these mod­els by US busi­nesses does noth­ing to ad­dress this risk, be­cause bad ac­tors are un­likely to be le­git­i­mate US busi­nesses. It would pro­tect US AI com­pa­nies from com­pe­ti­tion, but that has never been my goal.

To ad­dress these con­cerns, I do sup­port the fol­low­ing three mea­sures, which I and Anthropic have con­sis­tently ad­vo­cated for:

We should not sell pow­er­ful chips or chip­mak­ing equip­ment to China, and we should crack down on the ram­pant smug­gling3 and workarounds used to ob­tain ac­cess to such chips. China has lim­ited do­mes­tic pro­duc­tion ca­pac­ity, and there­fore, due to the scal­ing laws, can­not build more pow­er­ful mod­els than the US with­out US chips. This is the most ef­fi­cient and di­rect way to block threat #1, and by ham­per­ing the train­ing of mod­els that are out of reach of US law, it also in­di­rectly helps with threat #2.

We should crack down on in­dus­trial-scale dis­til­la­tion op­er­a­tions. Distillation is a much more com­pute-ef­fi­cient process than train­ing mod­els from scratch. It al­lows China to build much bet­ter mod­els than its num­ber of chips would or­di­nar­ily en­able, and thus par­tially evade chip bans. Distillation does not al­low the CCP to ob­tain equiv­a­lent or su­pe­rior AI ca­pa­bil­i­ties to the US, but it can bring the Chinese fron­tier to within a few months of the US fron­tier. It is true that many of the com­pa­nies car­ry­ing out these op­er­a­tions re­lease open-weights mod­els—but the open weights are far less rel­e­vant than the fact that the op­er­a­tions are backed by an au­thor­i­tar­ian state seek­ing to over­take the US at the fron­tier. We should have pol­icy in­ter­ven­tions to de­ter this be­hav­ior. A blan­ket ban on open-weights mod­els is nei­ther the cor­rect rem­edy nor some­thing we have called for4.

All suf­fi­ciently ca­pa­ble mod­els, open and closed, should go through manda­tory safety test­ing. The best way to ad­dress threat #2 is to just di­rectly test mod­els for cy­ber, bi­o­log­i­cal, and align­ment risks be­fore re­lease. I think this idea is ac­tu­ally close to a con­sen­sus: I have been heart­ened both that the Trump ad­min­is­tra­tion has moved in this di­rec­tion in re­cent months, and by re­cent in­dus­try pro­pos­als that would ap­ply such test­ing to the most ca­pa­ble mod­els re­gard­less of their coun­try of ori­gin or whether they are open or closed (while ex­empt­ing less ca­pa­ble mod­els, such as those from star­tups and acad­e­mia, en­tirely). Whether open mod­els do or don’t pose an in­creased risk, and whether that risk can be mit­i­gated, is some­thing that should emerge from test­ing, rather than be de­cided in ad­vance—and there may be promis­ing meth­ods for im­prov­ing the safety of open-weights mod­els, in­clud­ing re­cent re­search from AE Studio and Anthropic on mod­u­lar train­ing strate­gies. Note that to be ef­fec­tive, test­ing would need to be global, which means even the CCP would need to be on board. I think this may ac­tu­ally be pos­si­ble: as I wrote in The Adolescence of Technology, lim­ited co­op­er­a­tion around pre­vent­ing AI bi­o­log­i­cal weapons may be pos­si­ble be­cause it is in China’s in­ter­est too.

This brings me to the open let­ter. I agree with much of it: open weights ex­pand ac­cess to the AI econ­omy, they strengthen com­pe­ti­tion at least for some use cases, and they give cus­tomers greater con­trol. Concerns about dis­til­la­tion should be ad­dressed through tar­geted le­gal and com­mer­cial frame­works—the same mea­sure I de­scribed above. But I don’t agree with the let­ter’s as­ser­tions that open-weights mod­els nec­es­sar­ily make it eas­ier to de­velop safe­guards or that broad ac­cess to ca­pa­bil­i­ties nec­es­sar­ily helps de­fend­ers more than at­tack­ers. It seems at least as likely to me that the op­po­site will be true. For ex­am­ple, I worry that bi­ol­ogy will have a strong at­tacker-de­fender asym­me­try, where suf­fi­ciently ca­pa­ble mod­els may be able to quickly weaponize pan­demic-level viruses with widely avail­able ma­te­ri­als, whereas de­fense against these agents is a multi-year op­er­a­tional task in the best case (as we saw with Operation Warp Speed)5. Questions like this should be em­pir­i­cally an­swered by rig­or­ous pre-re­lease test­ing, not as­sumed in ad­vance.

To sum­ma­rize my and Anthropic’s po­si­tion, we have not and are not ad­vo­cat­ing for a ban on open-weights mod­els as a cat­e­gory. We should in­stead fo­cus on keep­ing pow­er­ful chips out of au­thor­i­tar­ian hands, stop­ping in­dus­trial-scale dis­til­la­tion, and re­quir­ing safety test­ing of all suf­fi­ciently ca­pa­ble mod­els, open and closed.

*Edit 28 July: Updated to note that the cited re­search on mod­u­lar train­ing strate­gies was a col­lab­o­ra­tion be­tween Anthropic and AE Studio.

Related con­tent

Cognizant and Anthropic ex­pand their part­ner­ship to bring Claude to en­ter­prise clients

Read more

Introducing Claude Opus 5

Opus 5 is a step change im­prove­ment for the Opus tier pow­er­ing long-run­ning agents while de­liv­er­ing im­prove­ments in cod­ing and pro­fes­sional work.

Read more

A re­search agenda for the Economic Futures Research Fund

We’re shar­ing the re­search agenda for the Anthropic Economic Futures Research Fund.

Read more

www.data.jma.go.jp

topへ

inc.com

www.inc.com

Please en­able JS and dis­able any ad blocker

Exclusive | Netflix exec goes ballistic after being fired for stunning 'trust exercise' confession at retreat: suit

nypost.com

A Netflix ex­ec­u­tive was fired from his $1.1 mil­lion a year job af­ter re­veal­ing dur­ing a trust ex­er­cise” at a work re­treat that he had taken med­ically pre­scribed ke­t­a­mine, a law­suit has claimed.

Kevin Baillie, who was vice pres­i­dent and head of cre­ative at Eyeline Studios, is su­ing the com­pany af­ter it launched an in­ves­ti­ga­tion into his com­ments that ul­ti­mately ended in his fir­ing, the pa­pers say.

Baillie, who’s been on the vi­sual ef­fects team for Pirates of the Caribbean” and the Harry Potter fran­chise, says he took the drug un­der med­ical su­per­vi­sion in October and November of 2022 at a Santa Barbara clinic.

He sought the treat­ment for clin­i­cal de­pres­sion af­ter the death of his mother, ac­cord­ing to the suit.

During what’s called a Vulnerability-Trust ex­er­cise” at a January 2026 re­treat at the ex­clu­sive Sendero Ranch, a Northern California prop­erty owned by Netflix, Baillie shared with his col­leagues that he had un­der­gone the treat­ment, the suit says.

Baillie claims he ex­plained the rea­son why he had taken the drug but was in­ves­ti­gated by Netflix. On March 18, 2026 a com­pany in­ves­ti­ga­tor brought the in­ci­dent up, in a man­ner sug­gest­ing sus­pi­cion of recre­ational drug use,” the suit reads.

The ex­ec­u­tive was fired in April with the com­pa­ny’s at­tor­ney con­firm­ing the ke­t­a­mine ther­apy is­sue has fac­tored into the ter­mi­na­tion,” the pa­pers say, and go on to sug­gest Baille was de­nied up to a year of sev­er­ance pay.

Baillie says in the law­suit the scope” of the in­ves­ti­ga­tion re­lated to al­leged pro­fan­ity and drink­ing. He had been warned dur­ing his per­for­mance re­view that he should drop one or two less f-bombs but don’t stop en­tirely.”

It goes on to say that at the same re­treat Baillie drank a Guinness stand­ing on his head, af­ter shar­ing that he had learned the trick from his for­mer fa­ther in law dur­ing a con­ver­sa­tion in­spired by the trust ses­sion.

His col­league im­me­di­ately asked for a demon­stra­tion, rather than with­hold the open­ness that the ses­sion had en­cour­aged, he per­formed the trick,” the pa­pers say.

Sign up for the California Morning Report newslet­ter

California’s top news, sports and en­ter­tain­ment de­liv­ered to your in­box every day.

Thanks for sign­ing up!

Baille also paints a pic­ture of an al­legedly al­co­hol-fu­eled com­pany en­vi­ron­ment en­cour­aged by the Eyeline Studios CEO Jeff Shapiro.

The law­suit al­leged alcohol con­sump­tion was com­pany-spon­sored, lead­er­ship-mod­eled and con­doned”, with mul­ti­ple ex­am­ples of how Shapiro set the cul­tural tone con­cern­ing al­co­hol at the ex­ec­u­tive level”.

The doc­u­ments claimed Shapiro on one oc­ca­sion pur­chased beer at a cor­ner store and brought it to a com­pany car ride to the Visual Effects Society Awards for staff to share.

Baillie claims he also wit­nessed the CEO con­sume al­co­hol with Netflix and Eyeline em­ploy­ees at his own wel­come din­ner in September 2024, the Netflix Annual Business Review events in March 2025 and even a Lakers game at­tended by Netflix ex­ecs in­clud­ing Shapiro’s di­rect su­per­vi­sor in February 2026.

All up, Baillie’s at­tor­neys pro­vided over half a dozen ex­am­ples of the CEO be­ing pre­sent at a work event with a drink in his hand, ac­cord­ing to court pa­pers.

Download The California Post App, fol­low us on so­cial, and sub­scribe to our newslet­ters

California Post News: Facebook, Instagram, TikTok, X, YouTube, WhatsApp, LinkedInCalifornia Post Sports Facebook, Instagram, TikTok, YouTube, XCalifornia Post Opinion California Post Newsletters: Sign up here!Cal­i­for­nia Post App: Download here!Home de­liv­ery: Sign up here!Page Six Hollywood: Sign up here!

In ad­di­tion to host­ing mul­ti­ple par­ties, Shapiro also had a per­sonal bar in his of­fice from which he served al­co­hol (to Baille) in­clud­ing af­ter a suc­cess­ful meet­ing with Netflix’s CEO Ted Sarandos,” ac­cord­ing to the pa­pers.

Baillie is ask­ing for a jury trial, com­pen­satory dam­ages, lost wages, dam­ages for emo­tional dis­tress, and puni­tive dam­ages.

Netflix and Eyeline were con­tacted for com­ment.

advanced-context-engineering-for-coding-agents/benchmarking-opus-5-on-slop-code-bench.md at main · humanlayer/advanced-context-engineering-for-coding-agents

github.com

Benchmarking Opus 5 on SlopCodeBench

we got bet­ter bench­marks

I’ve writ­ten be­fore some­thing along the lines of:

THERE ARE NO GOOD BENCHMARKS for a mod­el’s abil­ity to main­tain code­base qual­ity

THERE ARE NO GOOD BENCHMARKS for a mod­el’s abil­ity to main­tain code­base qual­ity

That was­n’t en­tirely true. I love noth­ing more than bury­ing a good lede.

Last Friday I dug into SlopCodeBench, a new-ish (March 2026) long-hori­zon cod­ing bench­mark from @GOrlanski’s lab at UW Madison. It ad­dresses the thing that both­ers me most about cod­ing bench­marks - that even larger” more com­plex bench­marks still di­vulge the whole prob­lem up front:

In con­trast, each chal­lenge in SlopCodeBench has mul­ti­ple checkpoints” - the model does­n’t know the whole prob­lem up front, it has to evolve the code­base over time as new re­quire­ments are di­vulged.

It’s a good pa­per. It’s not that long. You should read it.

What’s cool about this bench­mark is that it is un­sat­u­rated - at the time of run­ning, the best mod­els avail­able, GPT-5.4 and Opus 4.6, got 11% and 17% strict pass rates, re­spec­tively.

bench­ing opus 5 on slop­codebench

On Friday I ran three claude mod­els (Opus 4.8, Sonnet 5, and Opus 5) through a sub­set of SlopCodeBench and watched it live for six hours. Opus 5 wins tech­ni­cally but none of them did a very good job IMO. Will post more re­sults soon with Fable and 5.6 Sol in the mix.

The big head­line is that Opus 5 got a 24% on the small sub­set of the bench­mark that I ran - not much higher than Opus 4.6′s 17% strict pass rate in the orig­i­nal pa­per. All of the tested mod­els showed a pretty sig­nif­i­cant in­crease in ver­bosity, com­plex­ity, and a bunch of other code smell met­rics over the course of each chal­lenge, with Opus 5 writ­ing five times the num­ber of func­tions/​callables than Opus 4.8 over the course of the same set of chal­lenges.

My per­sonal read of this 23% pass rate is that SlopCodeBench fi­nally gives some sig­nal for a thing I’ve only ever been able to ar­gue from vibes - that for real-shaped soft­ware en­gi­neer­ing work, build­ing one is­sue at a time, to­day’s mod­els can’t be re­lied on to run lights-off with­out steer­ing.

the bench­mark sub­set

i had claude pick out 3 prob­lems from the repo, 17 check­points to­tal, a mix of easy/​medium/​hard la­beled prob­lems:

cir­cuit_e­val — easy (8 check­points)

data­base_mi­gra­tion — medium (5 check­points)

dy­nam­ic_­con­fig_ser­vice_api — hard (4 check­points)

There’s an ap­pen­dix at the end with all 17 check­points ex­plained in de­tail but I won’t put that all here.

and then I ran them across all three mod­els, in par­al­lel, with a fresh con­text win­dow per check­point. All mod­els got the same prompts and ran in the claude code har­ness.

the met­ric I de­cided I care about is the strict pass: every­thing new is green in­clud­ing every re­gres­sion test that was in­her­ited from pre­vi­ous check­points.

A model fails a check­point if the so­lu­tion has a de­fect - de­fects are de­tected by tak­ing the mod­els out­put, a CLI to run or in some cases e.g. an api server to poke at, and run­ning a set of held-out black-box tests against the pro­duced en­try­point.

Model writes code for ck1

Eval har­ness runs black-box tests against ck1

Model writes code for ck2

Eval runs black-box tests for ck1 and ck2

etc

Again, the strict pass cri­te­ria means that if a model bun­gles some­thing in check­point 4, it can’t pass the fol­low­ing check­points be­cause that fail­ing part of the code car­ries for­ward (unless the model in­dav­er­tently fixes an eval case in check­point 6 that was bro­ken in check­point 4, but we did­n’t see this hap­pen in prac­tice).

For all 9 test runs, none of the mod­els made it to the end of any chal­lenge with every­thing pass­ing, even on the prob­lem marked as easy” dif­fi­culty.

while it ran

son­net’s first check­point was more ex­pen­sive but by the end of prob­lem 1, son­net be­came the cheap­est of the three. (it seems like once the ba­sics were built and the work turned into main­te­nance, then the cost sav­ings started to take over)

For the first chal­lenge, last gen­er­a­tion’s mod­els ac­cu­mu­lated de­fects steadily, Opus 5 a de­fect each on check­points 4 and 5.

for the first two hours opus 5 was the only model with any strict passes at all — three in a row to start.

Things evolved as we went. Claude dili­gently up­dated the html.

Compared against the other mod­els, Opus 5 was tech­ni­cally bet­ter on prob­lem 1 (circuit_eval). But af­ter ac­ing the first three check­points, every sub­se­quent so­lu­tion had at least one de­fect (failing test case).

fi­nal re­sult

If our de­f­i­n­i­tion of suc­cess is reached the fi­nal check­point with no de­fects” then opus 5 failed all three prob­lems, but it failed slightly-less-badly than the other mod­els.

For the costs vs de­fects re­port, I re­ally hate claude-isms but this one i de­cided to leave in:

every dol­lar bought cor­rect­ness. no­body bought enough of it.

every dol­lar bought cor­rect­ness. no­body bought enough of it.

(obviously this small sub­set of the bench can­not tell us de­fin­i­tively that spend­ing more $$ will lead to higher pass rates)

as far as strict passes go, Opus 5 got four of them (24% pass rate) (the first three ck of cir­cuit_e­val, plus data­base_mi­gra­tion ck1).

opus 4.8 and son­net 5 both got one strict pass (6% pass rate), the same data­base_mi­gra­tion ck1 that opus 5 got.

so the win­ner cleared 4/17, and 3 of those were the open­ing check­points of one prob­lem. It would ap­pear we have an un­sat­u­rated bench­mark for the next fron­tier of mod­els. nice work @GOrlanski and team.

the slop me­ter

I’m not to­tally sold on linting the slop away” just yet, be­cause I don’t think its yet pos­si­ble to de­ter­min­is­ti­cally parse the maintainability” of a par­tic­u­lar code­base check­point. But they are in­ter­est­ing to keep an

Code qual­ity met­rics are in­ter­est­ing to keep an eye on, and they’re prob­a­bly di­rec­tion­ally cor­rect, and

With SlopCodeBench, you get the re­sults af­ter each check­point across var­i­ous qual­ity met­rics. There are 41 of them in the re­sults file. Roughly grouped:

size — source lines, files, func­tions, meth­ods, classes, state­ments, and lines added and re­moved at that check­point

com­plex­ity — cy­clo­matic com­plex­ity mean, max, and spread, how many func­tions land in the high” and extreme” bands, how con­cen­trated the com­plex­ity is, max nest­ing depth, and mean func­tion length

du­pli­ca­tion — cloned lines, and clone lines as a share of source

de­com­po­si­tion — sin­gle-use func­tions, triv­ial wrap­pers, un­used vari­ables, lines per sym­bol

rule vi­o­la­tions — lint er­rors and how many are auto-fix­able, ast-grep hits against test slop rules, and the share of lines flagged ver­bose

de­pen­dency graph — prop­a­ga­tion cost (how far a change rip­ples), cyclic de­pen­dency mass, de­pen­dency en­tropy (probably the most in­ter­est­ing one to me)

Each of these is com­puted de­ter­min­is­ti­cally us­ing the cur­rent code state af­ter each check­point.

The chart be­low shows the spread among mod­els for ck1 score vs. ck8 for the cir­cuit_e­val chal­lenge. (That is, how much did the slop in­di­ca­tor in­crease over the lif­time of the chal­lenge check­points.) Most in­ter­est­ingly, most of the met­rics don’t tell the mod­els apart.

I like that these mea­sures are re­peat­able and don’t use a model for judge­ment. But the link be­tween any one of them and is this code­base easy to change and evolve” is not yet es­tab­lished.

more cor­rect­ness came at the cost of wayyy more code

But a lot of that was more tests” - the ac­tual pro­duc­tion vol­ume is closer to 1.8x for opus 5 vs opus 4.8.

My guess would be…ex­pen­sive ver­bosity here that did­n’t trans­late di­rectly to much bet­ter re­sults.

Will have to dig in more to know whether this is a model tic sig­nal or just a this is ac­tu­ally a re­ally hard prob­lem and war­rants this much code”.

al­most all the code writ­ten trig­gered the slop me­ter

For all mod­els, a huge ma­jor­ity of the code lines tripped at least one of the the bench­mark’s slop rules. The av­er­ages across the three prob­lems:

opus 4.8 — 98%

opus 5 — 93%

son­net 5 — 89%

And specif­i­cally, lines flagged as be­ing too ver­bose go up across the tra­jec­tory for every model, roughly 65% at ck1 to 80% by ck8, even for Opus 5.

I’d ac­tu­ally prob­a­bly say this is a sign that some of the code qual­ity mea­sures are a bit over-ag­gres­sive. I looked into ap­ply­ing the rule­set to our type­script monorepo, but the cur­rent slop-code-bench de­tec­tors are python only.

So I had 5.6-Sol cook up a sub­set of rules for type­script, but it only came up with 76 slop de­tec­tors (compared to the SCB python li­brary of 200+) - but found some di­rec­tional find­ings - the Opus 5 lights-off-gen­er­ated so­lu­tions have over 11 times more slop trig­gers per kLOC than our 99%-AI-generated-but-also-carefully-reviewed Typescript monorepo (yes that’s a 1000% in­crease).

Obviously there’s a moun­tain of as­ter­isks on this find­ing (fewer rules, haven’t re­viewed the par­ity, etc.) but it’s in­ter­est­ing to say the least.

these mod­els write a lot of func­tions

Another in­ter­est­ing data point - opus 5 wrote 5x more func­tions than the other two mod­els. But Opus 4.8 wrote a higher %% of sin­gle-use func­tions (almost 50% of its func­tions were called ex­actly once). And Sonnet 5′s share of sin­gle-use func­tions is the high­est at 71.5%.

FWIW I don’t think lots of small func­tions is bad. I take it with a grain of salt these days, but I used to be a die-hard Clean Code guy. Small de­scrip­tive func­tion names are way bet­ter than lots of com­ments, yada yada

Complexity grows over time for all mod­els

I’ve been say­ing mod­els de­grade code­base qual­ity over time for about a year now, mostly on vibes. But now we have some data.

Not a sin­gle model made it through all the chal­lenges with­out in­creas­ing com­plex­ity across check­points. While Opus 5 has the low­est mean com­plex­ity, it also wrote 2000 func­tions. There’s a trade­off here: lots of small func­tions or fewer big ones. I don’t think any of these com­plex­ity met­rics can stand alone, but they give us some kind of com­pos­ite sig­nal.

Both son­net and Opus 4.8 an­swered the in­creas­ing com­plex­ity of the chal­lenge check­points by mak­ing in­di­vid­ual func­tions big­ger rather than mov­ing things around. Opus 4.8 is the ex­treme, up 70% over eight check­points, and its sin­gle worst func­tion ended at a cy­clo­matic com­plex­ity of 93.

Duplication is where they split. Opus 4.8 goes from 4.6% to 16.8%, with an in­flec­tion at ck3 — ~roughly where new re­quire­ments start fight­ing the ini­tial de­sign.

Aside - here’s what the first three check­points of cir­cuit_e­val ask for (full list­ing for all chal­lenges in the ap­pen­dix at the end):

ck1 — a CLI with –help, –version, a JSON out­put mode, and a check com­mand that parses and val­i­dates a .circ cir­cuit file. Every sig­nal is a sin­gle bit.

ck2 — an eval com­mand: pass the cir­cuit some in­puts, get the out­puts back. Still one bit per sig­nal, stan­dard boolean op­er­a­tors.

ck3 — sig­nals be­come vec­tors. data[7:0] in­stead of data, plus slic­ing, in­dex­ing, con­cate­na­tion, new op­er­a­tors, and a width check on every operand.

By the end one line in six is a copy of an­other line. The other two mod­els came down over the same stretch.

However, opus 5 is ba­si­cally flat, 2.41 to 2.64. I’ve pre­vi­ously ar­gued that improving code­base qual­ity over time” had not moved much be­tween model gen­er­a­tions. So if you trust du­pli­ca­tion as a golden met­ric, you could ar­gue that we did get­ting in­cre­men­tally bet­ter in the last ~3 months. Big if though, and I think most soft­ware ar­chi­tec­ture ex­perts would agree that it’s not black and white.

the shape of a bet­ter or­a­cle for soft­ware qual­ity

While the code qual­ity met­rics are in­ter­est­ing, I don’t think they tell the whole story, and its easy for a model to re­ward hack any of them. Just like SWE-bench shaped prob­lems were the best ver­i­fier for solve a soft­ware prob­lem one time”, be­cause they map onto real world work at that zoom level”, I think pass all ver­i­fiers for an in­cre­men­tally-di­vulged spec” is a very re­al­is­tic eval for can a model main­tain a code­base over time”.

That is, a code­base be­com­ing hard to main­tain would lead to fail­ing check­points in later stages, so a higher strict pass rate is a sig­nal that the model is good at build­ing a code­base that is main­tain­able.

With fron­tier mod­els like Fable / Sol prov­ing to be ex­pert de­bug­gers and re­verse-en­gi­neers, in­cor­po­rat­ing cost/​time/​to­ken met­rics might be­come more im­por­tant over time - fron­tier mod­els like Fable and Sol can prob­a­bly get it done in the NASTIEST of code­bases, but I’d ven­ture that a well-fac­tored code­base will tend to lead to shorter, more to­ken-ef­f­i­cent solves for fu­ture prob­lems.

And while build a whole fea­ture across 8 check­points” is a lot slower than solve a 15min SWE-bench mul­ti­lin­gual prob­lem”, it can be ex­e­cuted un­at­tended and is sub­ject de­ter­min­is­tic be­hav­ior ver­i­fiers at the end. So IMO its a much bet­ter or­a­cle than e.g. does an­other model think this code is clean”.

I think any even bet­ter sig­nal that we can hope to get from a model that is re­ally good at main­tain­ing a code­base, is that we could try hav­ing a fron­tier model like opus 5, fa­ble 5, or gpt-5.6-sol write the first N check­points, and see if a dumber model like son­net 5 or gpt-5.6-terra can im­ple­ment check­point N+1.

This am­pli­fies the sig­nal of whether the smart mod­els did a good job main­tain­ing high-qual­ity code that is easy to change. Whether a small model like Sonnet or Terra or even Haiku can im­ple­ment check­point 8 im­pacts the smart mod­els’ score on check­points 1 – 7.

some­thing you can ac­tu­ally mea­sure

My per­sonal read of all this is that SlopCodeBench gives a sig­nal for some­thing I’ve so far only been able to ar­gue from ex­pe­ri­ence - that for real-shaped soft­ware en­gi­neer­ing work, build­ing one is­sue at a time, to­day’s mod­els can’t be re­lied on to run lights-off with­out steer­ing.

SCB is a mea­sure of the fu­ture that I’ll be keep­ing a close eye on. I would­n’t bet my code­base on a good score in Frontier Code, SWE-Marathon, or DeepSWE, but if/​when mod­els can score 80%+ on a (well-held-out) bench­mark like SlopCodeBench which mea­sures it­er­a­tion over time, I’ll feel a LOT bet­ter about set­ting them loose with the lights off.

I won’t posit when that will hap­pen, be­cause when” mat­ters less than hav­ing a good sig­nal to know that it’s hap­pen­ing. (Assuming no­body accidentally” trains on test in the mean­time).

what’s next / things i’d do dif­fer­ently

I’ll be read­ing some of the slop­codebench prob­lems more deeply for in­spi­ra­tion and to cu­rate a few that I think map well to the day-to-day build­ing we do here at @humanlayer_dev.

Claude de­cided to par­al­lelize by model, run­ning each through three chal­lenges in se­quence. We could have just as eas­ily done 3 mod­els x 3 chal­lenges in 9 par­al­lel ses­sions and fin­ished in 1 – 2 hours in­stead of 6.

As I said, I looked into ap­ply­ing the rule­set to our type­script monorepo, but the cur­rent slop-code-bench de­tec­tors are python only. It would be in­ter­est­ing to port those to TS and a few other lan­guages. I hate to be that guy but I’d bet python is a more slop-prone lan­guage than most.

A missing underscore sent innocent man to prison for 18 months

arstechnica.com

One miss­ing un­der­score in a Skyrim-themed user­name put an in­no­cent Nova Scotia man in prison for 18 months.

A 2018 child-lur­ing in­ves­ti­ga­tion, which be­gan in Madison, Wisconsin, and even­tu­ally ex­tended to Halifax, Canada, was based on a false premise.

Police were look­ing for a man us­ing the Kik mes­sag­ing ser­vice un­der the name fus__ro_dah” (two un­der­scores af­ter fus”), but they ac­ci­den­tally re­quested records for the user­name fus_ro_dah” (one un­der­score af­ter fus”). This one-char­ac­ter dif­fer­ence led them not to the per­pe­tra­tor but to a Canadian man named Brandon Klayme.

(Ars read­ers may rec­og­nize fus ro dah” as the Unrelenting Force dragon shout” from The Elder Scrolls V: Skyrim.)

Despite find­ing no ev­i­dence of the crime on his dig­i­tal de­vices, Canadian po­lice ar­rested Klayme in 2020 on child sex abuse charges. He was con­victed af­ter a trial in 2023 and sen­tenced in 2024 to 18 months in prison. He served the full term.

Even af­ter re­lease, Klayme con­tin­ued to fight his con­vic­tion. In the process of prepar­ing his ap­peal, the user­name mis­take that led to all these years of dis­rup­tion was fi­nally dis­cov­ered. On Thursday, the Nova Scotia Court of Appeal over­turned Klayme’s con­vic­tion, writ­ing: Mr. Klayme is fac­tu­ally in­no­cent of the of­fences. He should never have been charged, let alone con­victed.”

One un­der­score

The case be­gan in 2018. From August through December of that year, a 12-year-old Wisconsin girl com­mu­ni­cated with an adult male through the Kik mes­sag­ing ser­vice. During a check of the girl’s phone, her mother found an inappropriate” photo of the male and called lo­cal po­lice.

The Dane County Sheriff’s Department re­sponded. A deputy took the phone, and the de­part­ment ran a foren­sic search on it. The re­port iden­ti­fied 125 Kik mes­sages be­tween the girl and an adult with the user­name fus__ro_dah” (two un­der­scores af­ter fus”).

To iden­tify this per­son, the cops con­tacted Kik, but their sub­poena ac­ci­den­tally re­quested in­for­ma­tion about the Kik user fus_ro_dah” (one un­der­score af­ter fus”). Kik pro­vided Klayme’s email ad­dress in re­sponse.

Google records showed that this email ad­dress was used to ac­cess Google ser­vices from an IP ad­dress in Canada, so the Dane County in­ves­ti­ga­tors turned the case over to Halifax Regional Police. Halifax po­lice took the IP ad­dress they had been given to lo­cal Internet provider Bell Aliant. Bell con­nected the IP ad­dress to the phys­i­cal ad­dress of their sub­scriber, Brandon Klayme.

New HIV vaccine shows unprecedented success in preclinical study

www.lji.org

Highlights:

Scientists at La Jolla Institute for Immunology, Scripps Research have de­vel­oped an HIV vac­cine that trains im­mune cells to see past HIVs de­fenses.

This HIV vac­cine works by prompt­ing the body’s im­mune sys­tem to make sub­stan­tial num­bers of rarely seen broadly neu­tral­iz­ing” an­ti­bod­ies.

In this new study, this vac­cine re­sulted in the best HIV-fighting an­ti­body re­sponse ever seen in pri­mates. Human tri­als have now started.

LA JOLLA, CA—A new HIV vac­cine de­vel­oped by La Jolla Institute for Immunology (LJI), Scripps Research sci­en­tists, and IAVI has the po­ten­tial to pro­tect hu­mans from de­vel­op­ing HIV in­fec­tion and AIDS. This HIV vac­cine is the first to gen­er­ate a high num­ber of broadly neu­tral­iz­ing,” virus-fight­ing an­ti­bod­ies in pri­mates.

This feels like a huge suc­cess,” says LJI Professor and Chief Scientific Officer Shane Crotty, Ph.D., who co-led the re­search with Scripps Research Professor William Schief, Ph.D. We con­structed a suc­cess­ful vac­cine from the ground up, which re­quired a deep un­der­stand­ing of the im­mune sys­tem.”

This ground­break­ing re­search, pub­lished in Nature, is the re­sult of 14 years of col­lab­o­ra­tion be­tween La Jolla Institute for Immunology and Scripps Research, as part of the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD). This has been one of those Apollo moon mis­sion-type pro­jects, where there is an ex­cep­tional goal and the team has to ac­com­plish a myr­iad of dis­cov­er­ies and in­ven­tions along the way,” says Crotty.

Outsmarting HIV

The new vac­cine works by in­ter­ven­ing in a process called B cell mat­u­ra­tion. B cells make an­ti­bod­ies. Like many im­mune cells, B cells have an early naive” stage be­fore they are ready to make an­ti­bod­ies. B cells start to ma­ture once they get the sig­nal that a pathogen, such as a virus, is try­ing to at­tack. B cells see pieces of that pathogen’s mol­e­c­u­lar struc­ture and start pro­duc­ing an­ti­bod­ies that can bind to that struc­ture and halt in­fec­tion.

It can take a lit­tle while for B cells to find the right bullseye” on a pathogen. But B cells keep try­ing. As they ma­ture, B cells tweak their an­ti­body pro­duc­tion, re­fin­ing an­ti­body struc­tures to bind to a pathogen in just the right, vul­ner­a­ble spots.

Scientists de­scribe B cell de­vel­op­ment as a train­ing process or boot­camp. In most cases, the body is left with a well-honed B cell army.

HIV is hard to beat be­cause it does­n’t give B cells a chance to de­velop ef­fec­tive an­ti­bod­ies. The first prob­lem is that HIV dis­guises it­self from the im­mune sys­tem. The virus is wrapped in an ever-shift­ing cloak of sugar mol­e­cules, called gly­cans. This lets HIV sneak un­de­tected past hu­man cells, which are also cov­ered in gly­cans.

The sec­ond big prob­lem is that HIV mu­tates very quickly. The world­wide di­ver­sity of HIV mu­ta­tions is ex­tra­or­di­nary. Even the di­ver­sity within one in­di­vid­ual per­son liv­ing with HIV is dra­matic,” says LJI Instructor Patrick Madden, Ph.D., who served as study co-first au­thor with Jon Steichen, Ph.D., an in­sti­tute in­ves­ti­ga­tor at Scripps Research.

The third prob­lem is that HIV changes its shape when it in­fects hu­man cells. Even if B cells get a glimpse of its vi­ral struc­ture—snap!—the struc­ture changes.

Taken to­gether, these prob­lems rarely give B cells a chance to hone their an­ti­body re­sponses against HIV. Even if a B cell man­ages to make neu­tral­iz­ing an­ti­bod­ies, the virus can mu­tate or change its shape, ren­der­ing those an­ti­bod­ies use­less.

The LJI and Scripps Research teams spent years hunt­ing for broadly neu­tral­iz­ing” an­ti­bod­ies that can ac­tu­ally bind to HIV and rec­og­nize key vi­ral struc­tures, even if the rest of the virus mu­tates. These an­ti­bod­ies are very, very rare, but they can be found in blood sam­ples from a small num­ber of peo­ple liv­ing with HIV.

An ef­fec­tive HIV vac­cine would need to prompt the im­mune sys­tem to make these same broadly neu­tral­iz­ing an­ti­bod­ies. How could we flip the whole im­mune re­sponse on its head so the rare re­sponses be­come the com­mon re­sponses? That was a crit­i­cal chal­lenge we faced,” says Crotty.

Testing the new vac­cine

It was time to go back to B cell boot­camp. The sci­en­tists stud­ied what made the HIV-fighting B cells spe­cial. Then they re­versed the process to see ex­actly how those B cells ma­tured. By look­ing back at the mat­u­ra­tion process, the re­searchers could track how the B cells changed when they saw spe­cific pieces of the HIV struc­ture.

The team dis­cov­ered that B cells ma­tured to make broadly neu­tral­iz­ing an­ti­bod­ies af­ter they got an early look at parts of HIVs outer envelope” pro­tein. Because these vi­ral sites sparked an im­mune re­sponse, sci­en­tists would call them antigens.”

An ef­fec­tive HIV vac­cine would likely need to in­clude mod­els of these anti­gens. The anti­gens would work like mugshots of America’s most wanted. If B cells saw those anti­gens early and of­ten, they would get re­ally good at rec­og­niz­ing and even neu­tral­iz­ing HIV. We were try­ing to mimic the pro­gres­sion of those neu­tral­iz­ing an­ti­bod­ies,” says Madden.

In a feat of mol­e­c­u­lar en­gi­neer­ing, the Schief Lab de­vel­oped vac­cine mol­e­cules that re­sem­bled the real HIV anti­gens. The sci­en­tists then worked with Emory National Primate Research Center, to test this po­ten­tial HIV vac­cine in a non-hu­man pri­mate species called rhe­sus macaques.

The re­searchers first ad­min­is­tered a priming” vac­cine meant to ac­ti­vate each an­i­mal’s naive B cells. The an­i­mals then re­ceived a se­ries of shepherding” booster shots to help their B cells de­velop along the right path.

This se­ries of vac­ci­na­tions will guide, or walk’, a B cell from its naive state to its broadly neu­tral­iz­ing state,” says Madden.

This new type of vac­cine ap­proach is called germline tar­get­ing” be­cause it tar­gets naive B cells in their germline” or naive form, be­fore they be­gin their train­ing process.

The sci­en­tists found that around 44 per­cent of the an­i­mals went on to pro­duce broadly neu­tral­iz­ing an­ti­bod­ies against HIV in their blood. These an­ti­bod­ies were im­pres­sively abun­dant.

We suc­ceeded in tak­ing ul­tra-rare an­ti­body re­sponses and turn­ing them into com­mon re­sponses by the end of the vac­ci­na­tion process,” adds Crotty. In other re­search re­cently pub­lished, they re­ported a new strat­egy to ac­cel­er­ate re­lated vac­cine an­ti­body re­sponses [See Nature Immunology pa­per].

The team did­n’t test whether these an­ti­bod­ies could pre­vent in­fec­tion, but it’s sig­nif­i­cant that these an­ti­bod­ies could be found in the blood, where they could en­counter and po­ten­tially block HIV.

Bringing the HIV vac­cine to hu­mans

The Crotty Lab plans to in­ves­ti­gate how they might change the booster shot reg­i­men to make the HIV vac­cine even more ef­fec­tive. It was in­cred­i­ble to get those re­sults, but of course we’d like to see a re­sponse in 100 per­cent of the an­i­mals,” says Madden.

Importantly, the an­ti­bod­ies found in the an­i­mal sub­jects re­sem­bled the ex­act kinds of broadly neu­tral­iz­ing an­ti­bod­ies seen in those rare hu­mans who made their own neu­tral­iz­ing an­ti­bod­ies. It’s clear that our im­mune sys­tems can make these pow­er­ful an­ti­bod­ies, given the right train­ing.

We be­lieve this vac­cine ap­proach is even more likely to suc­ceed in hu­mans, be­cause of the im­muno­genet­ics,” Crotty says.

The prim­ing im­muno­gen used in this study was eval­u­ated in hu­mans in the HVTN 144 trial and is cur­rently be­ing tested in the Phase 1 trial IAVI G004. IAVI, Scripps Research, the HIV Vaccine Trials Network, and part­ners are now ad­vanc­ing plans to fur­ther eval­u­ate the full im­mu­niza­tion reg­i­men in a fu­ture hu­man clin­i­cal study.

Additional au­thors of the study, Vaccination elic­its HIV broadly neu­tral­iz­ing an­ti­bod­ies in pri­mates,” in­clude Claudia T. Flynn, Swastik Phulera, Monolina Shil, Oleksandr Kalyuzhniy, Alessia Liguori, Carolyne Kifude, Leigh M. Sewall, Christopher A. Cottrell, Krystal M. Ma, Sabyasachi Baboo, Jolene K. Diedrich, Katherine McKenney, Allan C. de­Camp, Diane G. Carnathan, Ivy Phung, Parham Ramezani-Rad, Ester Marina-Zárate, Brian Freeman, Zhenfei Xie, Jeong Hyun Lee, Troy Sincomb, Nicole Phelps, Danny Lu, Diana Goodwin, Ryan Tingle, Yumiko Adachi, Nushin Alavi, Jenny Tran, Andy S. Tran, Alyne Nascimento, Catherine Sovie, Daniel L. V. Bader, Hannah Voic, Xiaoya Zhou, Grace Pixton, Agnes Walsh, Mariane B. Melo, Torben Schiffner, Facundo D. Batista, Dennis R. Burton, Darrell J. Irvine, James C. Paulson, John R. Yates III, Gabriel Ozorowski, Andrew B. Ward, Guido Silvestri.

This work was sup­ported by National Institute of Allergy and Infectious Diseases (NIAID), of the National Institutes of Health, through grant UM1 Al100663 to the Scripps Center for HIV/AIDS Vaccine Immunology and Immunogen Discovery (CHAVI-ID), grant UM1 AI144462 to the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD), P51 OD011132 to Emory National Primate Research Center, and R01 AI113867; by the  Gates Foundation un­der the Collaboration for AIDS Vaccine Discovery (NAC INV-007522, INV-008813, INV-034657, and INV-064772), via the IAVI Neutralizing Antibody Center (NAC); and by the National Institute of Health grant S10OD025052.

Decathlon launches Wero payment option on decathlon.de

www.sgieurope.com

Decathlon did­n’t just add a new pay­ment but­ton to its German check­out page. It handed European banks their first big-box sport­ing goods test case at a mo­ment when com­pe­ti­tion in dig­i­tal pay­ments is in­ten­si­fy­ing and the push for EU-made pay­ment in­fra­struc­ture is grow­ing.

Decathlon switched on Wero on de­cathlon.de on July 20, mak­ing Germany the first mar­ket in the French group’s in­ter­na­tional foot­print to carry the European pay­ment scheme.

Wero is the re­tail pay­ment prod­uct of the European Payments Initiative (EPI), a con­sor­tium of banks and pay­ment providers backed by 18 share­holder in­sti­tu­tions and more than 50 mem­bers across the re­gion. It routes pay­ments di­rectly be­tween bank ac­counts in real time, au­tho­rized in the cus­tomer’s bank­ing app via fin­ger­print or face recog­ni­tion, so the mer­chant never re­ceives a card num­ber.

Archive re­port­ing shows the scale Wero is still build­ing to­ward: 43 mil­lion ac­tive users across the re­gion in August 2025, 46 mil­lion by November 2025 along­side the re­tail launch, and around 56 mil­lion by the time of the Decathlon roll­out in July 2026, with German reg­is­tra­tions ris­ing from 1.3 mil­lion to 1.8 mil­lion over the same pe­riod.

Decathlon joins a mer­chant ros­ter that al­ready in­cludes Eventim Live, with Lidl, Rossmann and Hornbach ex­pected to fol­low.

The in­sider why

For Decathlon, the in­cen­tive is not just op­tics. Patrick Müller, chief dig­i­tal and chief mar­ket­ing of­fi­cer at Decathlon Germany, tied the launch to the re­tail­er’s mem­ber­ship pro­gram, say­ing the con­nec­tion de­liv­ers a gen­uine added value” on every pur­chase. That frames Wero as a loy­alty and re­ten­tion tool, not only a check­out op­tion. Industry es­ti­mates com­monly place card net­work pro­cess­ing costs at around 1 to 2 per­cent of trans­ac­tion value in fees to Visa and Mastercard. An ac­count to ac­count rail that by­passes that layer can be a di­rect mar­gin lever for a re­tailer op­er­at­ing at scale on thin sport­ing goods mar­gins.

The roll­out plan ex­tends be­yond the web­site.

Decathlon says it is prepar­ing the tech­ni­cal in­te­gra­tion needed to bring Wero to its roughly 110 phys­i­cal stores in Germany, with the goal of a con­sis­tent pay­ment ex­pe­ri­ence across on­line and of­fline chan­nels. A joint ac­ti­va­tion cam­paign with EPI is sched­uled for October: a six week pro­mo­tion of­fer­ing a €10 voucher to cus­tomers who pay with Wero on de­cathlon.de, funded en­tirely by EPI rather than Decathlon.

The chal­lenge and the promise

EPIs will­ing­ness to fund cus­tomer in­cen­tives points to the net­work’s core prob­lem. Payment schemes live or die on how many mer­chants and con­sumers show up at the same time, and Wero still lacks in store point of sale func­tion­al­ity. That ca­pa­bil­ity is not ex­pected un­til 2026 and 2027. Joachim Schmalzl, EPIs su­per­vi­sory board chair­man, has ac­knowl­edged the ini­tia­tive still faces real hur­dles even as he called the ecom­merce launch a mile­stone to­ward a sov­er­eign European pay­ment sys­tem.”

Decathlon be­ing among the first to im­ple­ment the sys­tem can help test Wero’s promise: lower pro­cess­ing costs, tighter loy­alty in­te­gra­tion and one less trans­ac­tion fee flow­ing to a US based net­work, pro­vided Wero can close the gap be­tween its am­bi­tions and its still lim­ited in store reach.

Judge Rejects Google’s Attempt To DMCA Its Way Out Of Being Scraped

www.techdirt.com

from the pulling-up-the-lad­der dept

Back in December we called out Google for fil­ing a DMCA 1201 law­suit over com­pa­nies scrap­ing Google’s re­sults. Almost every­thing about the law­suit seemed prob­lem­atic, not the least of which is that Google’s en­tire busi­ness was built on scrap­ing the web. To sue an­other com­pany for scrap­ing Google just felt… ob­nox­ious. And now a judge has dis­missed the law­suit, though leav­ing it open for Google to re­file.

Some back­ground: now that we’re in the age of AI, ac­cess to all kinds of data has be­come more pre­cious, which means we’re see­ing more and more at­tempts to put a toll booth on parts of the open web, pri­mar­ily aimed at AI com­pa­nies. But the rest of us get locked out along the way. SerpAPI is one of the play­ers in the space which (as its name im­plies) ba­si­cally tries to cre­ate an unau­tho­rized API for search en­gine re­sult pages.

Last fall, Reddit sued SerpAPI and some oth­ers (including search AI com­pany Perplexity), claim­ing that be­cause SerpAPI was al­low­ing oth­ers (like Perplexity) to ac­cess Reddit con­tent via its scrape of Google, it was vi­o­lat­ing the DMCAs anti-cir­cum­ven­tion (DMCA 1201) clause. We found the whole thing to be an at­tack on the prin­ci­ples of the open web. It re­ally seemed weird. Reddit had no copy­right in­ter­est in its users’ posts (the users hold the copy­right) and SerpAPI was scrap­ing Google, not Reddit. Reddit has an API deal with Google, but none of the par­ties be­ing sued were par­ties to that deal. The whole thing was just we don’t like that this is hap­pen­ing, so we’re su­ing.”

Google’s case came a few months later and was quite sim­i­lar, fo­cused on SerpAPI. And while at least in this case (unlike Reddit) they could point out that SerpAPI was scrap­ing their own site, it still makes no sense to claim that scrap­ing an open web­site can be a 1201 anti-cir­cum­ven­tion vi­o­la­tion, no mat­ter what technological pro­tec­tion mea­sures” you throw up to try to block scrap­ing. The Reddit case con­tin­ues to move for­ward with the de­fen­dants fil­ing mo­tions to dis­miss, but the Google case has lapped them a bit, with the judge al­ready dis­miss­ing the com­plaint, and point­ing out (correctly!) that Google has no le­git­i­mate copy­right claim to make here.

While SerpAPI tried a va­ri­ety of dif­fer­ent ways to kill the law­suit, what seemed to stick is that Google was clearly stretch­ing the way the DMCA 1201 is sup­posed to work. Remember, 1201 is the anti-circumvention” part of the DMCA, and was ini­tially writ­ten to pro­tect DRM so that if peo­ple broke DRM (or even talked about how to break DRM) they could still be held li­able for copy­right in­fringe­ment just for the act of cir­cum­vent­ing the technological pro­tec­tion mea­sure.” This very broad and poorly worded law has cre­ated huge messes in its wake, in­clud­ing bla­tant abuses like com­pa­nies ar­gu­ing that you can’t use third-party printer ink or third-party garage door open­ers be­cause of flimsy technological pro­tec­tion mea­sures” put into those de­vices, even though the un­der­ly­ing cir­cum­ven­tion had noth­ing to do with copy­right.

The court also looks at one of those ear­lier cases (regarding Lexmark’s print­ers), but con­cludes it does­n’t ap­ply here — long story, not worth the de­tail, ex­cept to note that the prece­dent that mat­tered against Lexmark came from trade­mark law, not the DMCA, even though Lexmark had also tried (and failed) to use Section 1201 it­self.

However, SerpAPI (rightly) also pointed out that Google is over­claim­ing what SearchGuard” — the technological pro­tec­tion mea­sure” — ac­tu­ally pro­tects here. As the court ex­plains it, SearchGuard is ba­si­cally a kind of CAPTCHA:

SearchGuard works by send­ing a JavaScript challenge” to search queries that Google re­ceives from un­rec­og­nized sources to con­firm that they come from real users as op­posed to au­to­mated soft­ware. Id. ¶ 29. Google’s com­puter sys­tem trans­mits JavaScript code that calls upon the user’s browser to send Google a solve” for the chal­lenge, i.e., to send Google spe­cific in­for­ma­tion re­gard­ing the browser and user gen­er­at­ing the re­quest. Id. ¶ 29. For hu­man users, the solve” is rel­a­tively straight­for­ward; their browsers run the JavaScript code and send back the re­quired in­for­ma­tion seam­lessly, with­out dis­rupt­ing the user ex­pe­ri­ence. Id. ¶ 29. However, au­to­mated sys­tems that sub­mit au­to­mated queries at a mas­sive scale typ­i­cally can­not solve the SearchGuard chal­lenge. Id. As a re­sult, SearchGuard de­nies them ac­cess to Google’s Search re­sults.

SearchGuard works by send­ing a JavaScript challenge” to search queries that Google re­ceives from un­rec­og­nized sources to con­firm that they come from real users as op­posed to au­to­mated soft­ware. Id. ¶ 29. Google’s com­puter sys­tem trans­mits JavaScript code that calls upon the user’s browser to send Google a solve” for the chal­lenge, i.e., to send Google spe­cific in­for­ma­tion re­gard­ing the browser and user gen­er­at­ing the re­quest. Id. ¶ 29. For hu­man users, the solve” is rel­a­tively straight­for­ward; their browsers run the JavaScript code and send back the re­quired in­for­ma­tion seam­lessly, with­out dis­rupt­ing the user ex­pe­ri­ence. Id. ¶ 29. However, au­to­mated sys­tems that sub­mit au­to­mated queries at a mas­sive scale typ­i­cally can­not solve the SearchGuard chal­lenge. Id. As a re­sult, SearchGuard de­nies them ac­cess to Google’s Search re­sults.

But, as SerpAPI high­lighted, SearchGuard has lit­tle to do with copy­right. And that, at least, gets the court’s at­ten­tion:

SerpApi con­tends that Google’s claims un­der the DMCA are sub­ject to dis­missal be­cause SearchGuard is de­signed and func­tions to con­trol ac­cess to and pre­vent the scrap­ing of Google Search re­sults re­gard­less of whether they con­tain a copy­righted com­po­nent, and be­cause SearchGuard is not rea­son­ably tai­lored to con­trol ac­cess only with re­spect to any copy­righted com­po­nent that may be in­cluded in Google Search re­sults. The Court agrees with SerpApi in part. To the ex­tent that Google Search re­sults do not con­tain any copy­righted con­tent, SearchGuard can­not be said to ef­fec­tively con­trol ac­cess to a work pro­tected un­der the Copyright Act. Here, Google al­leges that SearchGuard con­trols ac­cess to Google Search re­sults, which are com­pi­la­tions of pub­licly-avail­able in­for­ma­tion that Google ob­tains from the in­ter­net and or­ga­nizes for pre­sen­ta­tion to users on google.com based on rel­e­vance. See Compl. ¶¶ 13, 14, 27. SearchGuard con­trols ac­cess to Google Search re­sults be­cause its purpose” is to pre­vent unau­tho­rized third par­ties from au­to­mat­i­cally ac­cess­ing Google’s Search re­sults” to scrape them, as such scrap­ing ac­tiv­i­ties im­pose a deadweight loss” on Google. See id. ¶¶ 24, 26 – 27, 29. However, Google does not al­lege that google.com or the Google Search re­sults dis­played therein are pro­tected un­der the Copyright Act. Importantly, Google al­leges that Google Search re­sults are often” ac­com­pa­nied by a Knowledge Panel” that may con­tain some copy­righted con­tent that Google li­censes from third par­ties, such as copy­righted im­ages. Google does not al­lege that the Knowledge Panel” is al­ways in­cluded in Google Search re­sults, or that the Knowledge Panel, if in­cluded in the Search re­sults, al­ways con­tains copy­righted con­tent. See id. ¶¶ 14 – 16. Accordingly, Google’s al­le­ga­tions in­di­cate a mix of con­tent, some with copy­righted ma­te­r­ial and oth­ers with­out.

SerpApi con­tends that Google’s claims un­der the DMCA are sub­ject to dis­missal be­cause SearchGuard is de­signed and func­tions to con­trol ac­cess to and pre­vent the scrap­ing of Google Search re­sults re­gard­less of whether they con­tain a copy­righted com­po­nent, and be­cause SearchGuard is not rea­son­ably tai­lored to con­trol ac­cess only with re­spect to any copy­righted com­po­nent that may be in­cluded in Google Search re­sults.

The Court agrees with SerpApi in part. To the ex­tent that Google Search re­sults do not con­tain any copy­righted con­tent, SearchGuard can­not be said to ef­fec­tively con­trol ac­cess to a work pro­tected un­der the Copyright Act. Here, Google al­leges that SearchGuard con­trols ac­cess to Google Search re­sults, which are com­pi­la­tions of pub­licly-avail­able in­for­ma­tion that Google ob­tains from the in­ter­net and or­ga­nizes for pre­sen­ta­tion to users on google.com based on rel­e­vance. See Compl. ¶¶ 13, 14, 27. SearchGuard con­trols ac­cess to Google Search re­sults be­cause its purpose” is to pre­vent unau­tho­rized third par­ties from au­to­mat­i­cally ac­cess­ing Google’s Search re­sults” to scrape them, as such scrap­ing ac­tiv­i­ties im­pose a deadweight loss” on Google. See id. ¶¶ 24, 26 – 27, 29. However, Google does not al­lege that google.com or the Google Search re­sults dis­played therein are pro­tected un­der the Copyright Act. Importantly, Google al­leges that Google Search re­sults are often” ac­com­pa­nied by a Knowledge Panel” that may con­tain some copy­righted con­tent that Google li­censes from third par­ties, such as copy­righted im­ages. Google does not al­lege that the Knowledge Panel” is al­ways in­cluded in Google Search re­sults, or that the Knowledge Panel, if in­cluded in the Search re­sults, al­ways con­tains copy­righted con­tent. See id. ¶¶ 14 – 16. Accordingly, Google’s al­le­ga­tions in­di­cate a mix of con­tent, some with copy­righted ma­te­r­ial and oth­ers with­out.

And that cuts against Google’s ar­gu­ment here:

Thus, be­cause the DMCA does not ap­ply where the work con­trolled by a tech­no­log­i­cal mea­sure is not pro­tected un­der the Copyright Act, Google’s claims un­der 17 U.S.C. § 1201(a)(1)(A) and 17 U.S.C. § 1201(a)(2) are sub­ject to dis­missal as a mat­ter of law to the ex­tent that they are premised on in­stances where SearchGuard con­trols ac­cess to Google Search re­sults that do not con­tain any copy­righted con­tent.

Thus, be­cause the DMCA does not ap­ply where the work con­trolled by a tech­no­log­i­cal mea­sure is not pro­tected un­der the Copyright Act, Google’s claims un­der 17 U.S.C. § 1201(a)(1)(A) and 17 U.S.C. § 1201(a)(2) are sub­ject to dis­missal as a mat­ter of law to the ex­tent that they are premised on in­stances where SearchGuard con­trols ac­cess to Google Search re­sults that do not con­tain any copy­righted con­tent.

Even more damn­ing for Google is that when it’s us­ing SearchGuard, that has lit­er­ally noth­ing to do with effectively con­trol­ling ac­cess to a [copyright-protected] work.” And that’s the en­tire point of 1201.

SerpApi ar­gues that Google’s claims un­der the DMCA fail be­cause it does not al­lege that it im­ple­mented SearchGuard to pro­tect a copy­righted work with the authority of the copy­right owner” as re­quired un­der 17 U.S.C. § 1201(a)(3)(B)…. The Court agrees. The plain lan­guage of 17 U.S.C. § 1201(a)(3)(B) makes clear that, for a tech­no­log­i­cal mea­sure to effectively con­trol[] ac­cess to a work” it must, among other things, require[] the ap­pli­ca­tion of in­for­ma­tion, or a process or a treat­ment, with the au­thor­ity of the copy­right owner, to gain ac­cess to the work.” See 17 U.S.C. § 1201(a)(3)(B). The Ninth Circuit has in­ter­preted the with the au­thor­ity of the copy­right owner” el­e­ment as re­quir­ing a plain­tiff to al­lege and later prove that the tech­no­log­i­cal mea­sure in ques­tion was im­ple­mented and func­tioned with the au­thor­ity of the copy­right owner.

SerpApi ar­gues that Google’s claims un­der the DMCA fail be­cause it does not al­lege that it im­ple­mented SearchGuard to pro­tect a copy­righted work with the authority of the copy­right owner” as re­quired un­der 17 U.S.C. § 1201(a)(3)(B)….

The Court agrees. The plain lan­guage of 17 U.S.C. § 1201(a)(3)(B) makes clear that, for a tech­no­log­i­cal mea­sure to effectively con­trol[] ac­cess to a work” it must, among other things, require[] the ap­pli­ca­tion of in­for­ma­tion, or a process or a treat­ment, with the au­thor­ity of the copy­right owner, to gain ac­cess to the work.” See 17 U.S.C. § 1201(a)(3)(B). The Ninth Circuit has in­ter­preted the with the au­thor­ity of the copy­right owner” el­e­ment as re­quir­ing a plain­tiff to al­lege and later prove that the tech­no­log­i­cal mea­sure in ques­tion was im­ple­mented and func­tioned with the au­thor­ity of the copy­right owner.

Google tried to ar­gue that it some­how has the sup­port of copy­right hold­ers to pro­tect their work with SearchGuard, but the court is not im­pressed.

Google’s ar­gu­ments do not com­pel a dif­fer­ent con­clu­sion. It con­tends that it is not re­quired to al­lege facts in­di­cat­ing that it had the au­thor­ity of the copy­right own­ers to im­ple­ment SearchGuard be­cause the phrase with the au­thor­ity of the copy­right owner” de­fines who may cir­cum­vent a tech­no­log­i­cal mea­sure to gain ac­cess to pro­tected work and does not de­fine who may de­ploy a tech­no­log­i­cal mea­sure to con­trol ac­cess to a pro­tected work…. This ar­gu­ment is un­avail­ing. Google’s au­thor­i­ties in­ter­pret a dif­fer­ent pro­vi­sion of the DMCA, namely 17 U.S.C. § 1201(a)(3)(A), which de­fines what it means to circumvent a tech­no­log­i­cal mea­sure.” See Disney Enters., Inc. v. VidAngel, Inc., 869 F.3d 848, 863 (9th Cir. 2017) (“Section 1201(a)(3)(A) ex­empts from cir­cum­ven­tion li­a­bil­ity only those whom a copy­right owner au­tho­rizes to cir­cum­vent an ac­cess con­trol mea­sure, not those whom a copy­right owner au­tho­rizes to ac­cess the work.”) (citation and in­ter­nal quo­ta­tion marks omit­ted); Universal City Studios, Inc. v. Corley, 273 F.3d 429, 444 (2d Cir. 2001) (“[S]ubsection 1201(a)(3)(A) frees an in­di­vid­ual to traf­fic in en­cryp­tion tech­nol­ogy de­signed or mar­keted to cir­cum­vent an en­cryp­tion mea­sure if the owner of the ma­te­r­ial pro­tected by the en­cryp­tion mea­sure au­tho­rizes that cir­cum­ven­tion.”). These au­thor­i­ties do not ad­dress the is­sue here, which is whether a tech­no­log­i­cal mea­sure must func­tion with the au­thor­ity of the copy­right owner” in or­der to effectively con­trol[] ac­cess to a work” un­der 17 U.S.C. § 1201(a)(3)(B).

Google’s ar­gu­ments do not com­pel a dif­fer­ent con­clu­sion. It con­tends that it is not re­quired to al­lege facts in­di­cat­ing that it had the au­thor­ity of the copy­right own­ers to im­ple­ment SearchGuard be­cause the phrase with the au­thor­ity of the copy­right owner” de­fines who may cir­cum­vent a tech­no­log­i­cal mea­sure to gain ac­cess to pro­tected work and does not de­fine who may de­ploy a tech­no­log­i­cal mea­sure to con­trol ac­cess to a pro­tected work…. This ar­gu­ment is un­avail­ing. Google’s au­thor­i­ties in­ter­pret a dif­fer­ent pro­vi­sion of the DMCA, namely 17 U.S.C. § 1201(a)(3)(A), which de­fines what it means to circumvent a tech­no­log­i­cal mea­sure.” See Disney Enters., Inc. v. VidAngel, Inc., 869 F.3d 848, 863 (9th Cir. 2017) (“Section 1201(a)(3)(A) ex­empts from cir­cum­ven­tion li­a­bil­ity only those whom a copy­right owner au­tho­rizes to cir­cum­vent an ac­cess con­trol mea­sure, not those whom a copy­right owner au­tho­rizes to ac­cess the work.”) (citation and in­ter­nal quo­ta­tion marks omit­ted); Universal City Studios, Inc. v. Corley, 273 F.3d 429, 444 (2d Cir. 2001) (“[S]ubsection 1201(a)(3)(A) frees an in­di­vid­ual to traf­fic in en­cryp­tion tech­nol­ogy de­signed or mar­keted to cir­cum­vent an en­cryp­tion mea­sure if the owner of the ma­te­r­ial pro­tected by the en­cryp­tion mea­sure au­tho­rizes that cir­cum­ven­tion.”). These au­thor­i­ties do not ad­dress the is­sue here, which is whether a tech­no­log­i­cal mea­sure must func­tion with the au­thor­ity of the copy­right owner” in or­der to effectively con­trol[] ac­cess to a work” un­der 17 U.S.C. § 1201(a)(3)(B).

Some of SerpAPI’s other ar­gu­ments fail, but for now all the DMCA claims are dis­missed, though Google can (and al­most cer­tainly will) re­file re­gard­ing some more nar­row claims. Specifically, Google can­not file claims re­gard­ing search re­sults for which it does not hold the copy­right, but could file more nar­row claims re­gard­ing con­tent where it does (such as the Knowledge Panel). That’s much more lim­ited, and about the only rea­son to keep the case go­ing is to be a nui­sance to SerpAPI.

That might be worth it to Google, which re­ally seems to dis­like SerpAPI be­ing out there and scrap­ing their re­sults. But it would be a much nar­rower case, and (in the­ory) SerpAPI could sim­ply change its scrap­ing to avoid Google-produced con­tent. Either way, all of this re­mains quite silly. Google’s en­tire busi­ness was built on scrap­ing the web. Suing some­one else for scrap­ing Google sure feels like pulling up the open in­ter­net lad­der up af­ter them­selves.

SerpAPI’s com­ments on the dis­missal make this point ex­plic­itly:

We’re pleased that the court re­jected Google’s at­tempts to ex­pand the DMCA to as­sert con­trol over ac­cess to pub­lic pages. The in­ter­net’s found­ing prin­ci­ple — open ac­cess to us­able in­for­ma­tion — is es­sen­tial to dri­ving in­no­va­tion and en­sur­ing every­one ben­e­fits from the promise of data. SerpApi will con­tinue sup­port­ing de­vel­op­ers, AI com­pa­nies, re­searchers, and busi­nesses that rely on ac­cess to pub­lic search in­for­ma­tion.

We’re pleased that the court re­jected Google’s at­tempts to ex­pand the DMCA to as­sert con­trol over ac­cess to pub­lic pages. The in­ter­net’s found­ing prin­ci­ple — open ac­cess to us­able in­for­ma­tion — is es­sen­tial to dri­ving in­no­va­tion and en­sur­ing every­one ben­e­fits from the promise of data. SerpApi will con­tinue sup­port­ing de­vel­op­ers, AI com­pa­nies, re­searchers, and busi­nesses that rely on ac­cess to pub­lic search in­for­ma­tion.

One would hope that this ini­tial dis­missal from the court gets the com­pany to re­think this anti-open-in­ter­net strat­egy, but some­how I fear the old adage of young com­pa­nies in­no­vate, old com­pa­nies lit­i­gate” is start­ing to seep into Google.

Filed Under: copy­right, dmca, dmca 1201, scrap­ing, search re­sults

Companies: google, red­dit, ser­papi

How to survive boiling water

taxa.substack.com

The story of MITs most no­to­ri­ous milk car­ton be­gins, as many good sto­ries do, in a col­lege dorm.

The milk in ques­tion was pur­chased in 1994 and re­dis­cov­ered in 1995 by an un­der­grad named Justin Cave. By that point, the re­port­edly lac­tose-in­tol­er­ant Cave had even less use for the milk he had aban­doned in his fridge ten months ear­lier. For rea­sons lost to his­tory, he did not throw the milk away. He threw it a birth­day party.

The Milk lived the rest of its life un­re­frig­er­ated, stored in a tall, sin­gle-walled jar. For twenty-seven years, the res­i­dents of Random Hall dorm gath­ered faith­fully to cel­e­brate its birth­day. At age 20, the Milk ap­plied to, and was re­jected from, MIT1. The jar was pe­ri­od­i­cally burped” to re­lease the gas pres­sure in­side, un­til the Milk reached its sta­ble fi­nal form — a cloudy brown liq­uid. When asked why the Milk was never thrown away, one res­i­dent of Random Hall replied: Why throw some­thing away when you can tell a story about it?”

Stuff I learned from things that nor­mally get thrown away” could be the ti­tle of many sci­en­tists’ mem­oirs, in­clud­ing Louis Pasteur’s. Winemaking pro­duces, well, wine, but it also pro­duces acidic crys­tals on the walls of the vats. These byprod­ucts were not dis­carded — they were stud­ied by Pasteur and his con­tem­po­raries. Pasteur’s ob­ser­va­tions both rev­o­lu­tion­ized our un­der­stand­ing of chem­istry and led him to the phe­nom­e­non that would de­fine his ca­reer and lay the foun­da­tion of mod­ern food safety: fer­men­ta­tion.

The mi­croor­gan­isms re­spon­si­ble for fer­men­ta­tion are vis­i­ble to our naked senses only through the tex­tures, col­ors and smells re­sult­ing from their col­lec­tive ef­forts. From Aristotle through the 1850s, it was as­sumed that some in­trin­sic prop­erty of a non-liv­ing start­ing sub­stance (like grain or milk) en­abled its spon­ta­neous fer­men­ta­tion into some­thing use­ful (like beer or yo­gurt) or its even­tual spoilage.

It was Pasteur who proved that liv­ing or­gan­isms were re­quired for the trans­for­ma­tions that took place dur­ing fer­men­ta­tion. He heated up nu­tri­ent-rich broths in cus­tom flasks that let gases, but not mi­crobes, flow in and out of the flasks. Pasteur then broke the neck off of one of the flasks to ex­pose the broth to the air. If the boiled broth could spon­ta­neously trans­form, it would do so with or with­out ex­po­sure to mi­crobes in the air and en­vi­ron­ment.

The flask with the neck bro­ken off grew cloudy and fer­mented as bac­te­ria bloomed, but the ster­ile one re­mained clear.

This find­ing was great news for Napoleon. The French were los­ing money, and per­haps more alarm­ingly, their rep­u­ta­tion, ex­port­ing wine to the British — the wine would spontaneously” go bad dur­ing ship­ping. The French gov­ern­ment of­fered a prize for a sci­en­tist to solve the case of the spoiled wine. With the knowl­edge of the mi­croor­gan­ism-dri­ven process of fer­men­ta­tion in hand, Pasteur did to the wine what he did to the broth, just more gen­tly — he heated the wine enough to kill mi­crobes with­out dam­ag­ing the wine’s fla­vor. Immortalized as pas­teur­iza­tion, this process was adapted shortly af­ter its in­ven­tion in 1865 to let us safely drink stored milk2.

Killing bac­te­ria thus be­came a ma­jor pre­oc­cu­pa­tion of mod­ern life. Its most vis­i­ble man­i­fes­ta­tion to­day might be tak­ing an­tibi­otics (first avail­able in the 1940s): about 7 out of 10 peo­ple in the U.S. were pre­scribed an­tibi­otics in 2024 ac­cord­ing to CDC data. A close sec­ond might be the dizzy­ing ar­ray of dis­in­fec­tant prod­ucts found in U.S. gro­cery stores.

What’s less vis­i­ble is the san­i­ti­za­tion in­fra­struc­ture that makes things like gro­cery stores or med­i­cine pos­si­ble at all. The com­pany Steris, one maker of high tem­per­a­ture, pres­sur­ized ster­il­iza­tion equip­ment and other med­ical in­stru­ments, is a $5 bil­lion an­nual rev­enue com­pany, with a $21 bil­lion mar­ket cap. The U.S. pas­teur­izes around 50 bil­lion liters of fluid milk every year. To pack­age salad greens like spinach, the greens are washed in a di­lute bleach so­lu­tion to kill any lin­ger­ing soil mi­crobes. I could go on.

But our war on bac­te­ria has its own war­ring in­dus­try. This in­dus­try has cap­tured the imag­i­na­tions of sci­en­tists, the food and bev­er­age in­dus­try, pharma com­pa­nies and doc­tors along with in­flu­encers, mar­ket­ing gu­rus and op­por­tunists of all fla­vors. This in­dus­try em­pha­sizes that some mi­crobes are friends, not foe, (true) and you should be eat­ing them in large quan­ti­ties, on pur­pose, all the time, and prefer­ably pay­ing more for prod­ucts that con­tain them (dubious). This is the pro­bi­otics in­dus­try.

Probiotic” is a bit of a mis­nomer — it means for life, or pro­mot­ing life, but the for­mal de­f­i­n­i­tion of a pro­bi­otic is an ac­tual liv­ing mi­croor­gan­ism. In sim­ple terms: tak­ing a pro­bi­otic is just eat­ing bac­te­ria on pur­pose. I say on pur­pose be­cause we con­sume mi­crobes ac­ci­den­tally all the time from our en­vi­ron­ment, largely obliv­i­ous to their ex­is­tence or ef­fects. The bac­te­ria we spend much of our time and en­ergy try­ing to kill are out­num­bered, at a species level, at least 1000 to 1 by a com­bi­na­tion of harm­less and ben­e­fi­cial bac­te­ria liv­ing in and on our bod­ies. It’s this lat­ter prop­erty of ben­e­fi­cial­ness that pro­bi­otics are try­ing to ex­ploit.

I say ex­ploit be­cause of a re­cent trip I took to the gro­cery store. I had a cold and was in search of lemon gin­ger tea. I bought a box of Bigelow, went home, boiled some wa­ter, poured it over a tea bag, waited a bit, added honey, took a sip, and al­most spit it out. The tea had its ex­pected notes of gin­ger, a hint of lemon, and some pow­dery, al­ka­line af­ter­taste that I could­n’t place. Frankly, it tasted ter­ri­ble. (Sorry, Bigelow).

I in­spected the box again. In my con­gested state, I had un­wit­tingly pur­chased a new of­fer­ing from the tea com­pany — Bigelow Lemon Ginger, with pro­bi­otics. What made this tea dif­fer­ent from all the other teas I hap­pily sipped on was that in ad­di­tion to nice-sound­ing things like lemon­grass and cin­na­mon, it con­tained bac­te­ria. Bacteria which I had just boiled, at a tem­per­a­ture 40oC hot­ter than pas­teur­iza­tion.

Did the tea taste bad be­cause I was drink­ing dead bac­te­ria wa­ter? And if that was the ul­ti­mate out­come of the nor­mal brew­ing process, why bother putting bac­te­ria in the tea at all?

I was at a lab happy hour when I men­tioned this to my PhD the­sis ad­vi­sor. I know, right?” she said, sud­denly an­i­mated. Probiotic teas taste SO BAD.” I was thrilled to have an­other wit­ness. Doesn’t it seem crazy to add in bac­te­ria that you’re just go­ing to boil and kill any­way?” I asked. Is it all a scam?” She was al­ready nod­ding. You have to won­der whether the bac­te­ria in the tea make it to the gut at all, and whether they do any­thing help­ful once they get there,” she said.

I’d be ly­ing if I said I re­mem­bered ex­actly what hap­pened next, or who sug­gested what. All I re­mem­ber is an idea. An idea to test this seem­ingly para­dox­i­cal mar­ket­ing tac­tic like the mi­cro­bi­ol­o­gists we are. The idea was sim­ple: What if we tried to grow the bac­te­ria from the tea bag, in the lab?

I went home. I stared at the box of bac­te­ria tea.

Why throw some­thing away when you can tell a story about it?

BC30™, the bac­te­r­ial strain in the tea, is short for Bacillus co­ag­u­lans GBI-30, 6086®. It re­ceived the FDAs GRAS (Generally Recognized as Safe3) des­ig­na­tion in 2012 and is found in over a thou­sand leading food, bev­er­age and pet food prod­ucts world­wide” ac­cord­ing to the pro­bi­otic’s web­site.

To coax these bac­te­ria to grow out of steeped tea, I needed to know three things:

Is this species safe to grow in the lab?

Is this species safe to grow in the lab?

What does it like to eat?

What does it like to eat?

What are its pre­ferred growth con­di­tions?

What are its pre­ferred growth con­di­tions?

In gen­eral, I try not to in­gest the bac­te­ria I grow in the lab — even ones with the low­est safety des­ig­na­tion, BSL-1. By na­ture of it be­ing a com­mer­cial pro­bi­otic, BC30 is both BSL-1 (safe to grow un­der nor­mal lab pre­cau­tions) and ed­i­ble.

But, I still would­n’t try this at home or eat bac­te­ria off of a cul­ture plate. Why? BC30s pre­ferred food source is not that dif­fer­ent from the pre­ferred food source of many other mi­croor­gan­isms: a sugar- and amino acid-rich nu­tri­ent medium called MRS (De Man, Rogosa and Sharpe) agar.

While MRS agar has some ad­just­ments to make it pref­er­en­tially ap­pe­tiz­ing to BC30 and its rel­a­tives, those rel­a­tives also in­clude Streptococcus pyo­genes (causes strep throat) and Bacillus cereus (causes food poi­son­ing). Without ster­ile tech­nique and rig­or­ous species-level con­fir­ma­tion, you can­not know for sure what is grow­ing on your plate.

With that said, the American Society for Microbiology’s blog sug­gested that were I suc­cess­ful in cul­tur­ing BC30, I would see growth of translu­cent white colonies on MRS agar plates af­ter 48 hours of in­cu­ba­tion at 30 – 33oC in the pres­ence of oxy­gen.

First, I needed to make tea.

I wanted the con­di­tions of the ex­per­i­ment to rep­re­sent a range of re­al­is­tic tea-drink­ing sce­nar­ios, from in­tended use to fla­grant im­pro­vi­sa­tion, and set up three steeps:

The Rule Follower — Brewed as di­rected for 4 min­utes in boil­ing wa­ter.

The Rule Follower — Brewed as di­rected for 4 min­utes in boil­ing wa­ter.

I for­got I made tea” — We’ve all been there. 15 min­utes, boil­ing wa­ter.

I for­got I made tea” — We’ve all been there. 15 min­utes, boil­ing wa­ter.

Cold brew an­ar­chist — Self-explanatory.

Cold brew an­ar­chist — Self-explanatory.

It was at this point I re­al­ized I needed a ster­ile-ish way to trans­port the steeped tea and tea bags from my house to the lab. Luckily, I had re­cently run a blind­folded vol­ume pour­ing ac­cu­racy com­pe­ti­tion at our de­part­men­tal re­treat and had left­over Falcon tubes still in their orig­i­nal pack­age. While the tea was def­i­nitely not ster­ile, I rea­soned that a lit­tle ex­tra asep­tic tech­nique would­n’t hurt. I poured the tea into the tubes over my kitchen stove, us­ing the open flame as a makeshift Bunsen burner.

In re­al­ity, lab came first. I had to make MRS agar plates be­fore I steeped the tea. Our lab does not use MRS broth very of­ten, and when I first looked for some all I found was a 10-year-old so­lid­i­fied block of MRS pow­der in our stock cab­i­net that was grow­ing large green spots in­side of its glass con­tainer. Behind it was one that looked mer­ci­fully nor­mal.

I mixed broth pow­der, agar and wa­ter in a glass bot­tle, loosely capped it, put it in a wa­ter bath and took it to the au­to­clave. Autoclave” is a nice word for gi­ant pres­sure cooker. Ours is made by the afore­men­tioned Steris. It rat­tled and hissed as its jaws opened to ac­cept my tray of cul­ture me­dia, which it then heated to 121oC for 45 min­utes, ster­il­iz­ing the liq­uid.

Back at my lab bench, when the molten MRS agar had cooled enough to han­dle, I lit a Bunsen burner next to a stack of empty plas­tic petri dishes and poured a layer of agar into each one. Left overnight, the plates so­lid­i­fied into nu­tri­ent-dense beds for BC30.

The next day, I took my tubes of tea to lab. I pipet­ted 400 mi­cro­liters (0.4mL) of each liq­uid tea con­di­tion onto a plate next to the Bunsen burner. I spread the liq­uid evenly across the plate with a hockey stick-shaped plas­tic spread­er4 and left the lids on the plates cracked open to dry near the flame.

But to an­swer my ques­tion, I needed one more test. If there were bac­te­ria in the tea bag ini­tially, but they died when boiled, then I might see bac­te­r­ial growth by plat­ing the dry in­gre­di­ents of an un­steeped tea bag, or the tea bag steeped in cold wa­ter. I cut open the tea bags and shook some of their con­tents onto the agar. I put my full set of plates, in­clud­ing a plain MRS plate to check its steril­ity, into the in­cu­ba­tor at 37oC — a stan­dard growth tem­per­a­ture, but a lit­tle warmer than rec­om­mended. I was skep­ti­cal that any­thing would grow. For the next two days, all I could do was wait and see.

The first thing I no­ticed when I took the plates out of the in­cu­ba­tor was the smell. I was in dis­be­lief when I saw lit­tle white colonies dot­ting al­most all the plates and opened one to get a closer look. A sickly sweet, gin­gery aroma wafted from the plate as I in­spected the translu­cent colonies — a byprod­uct of the bac­te­ria me­tab­o­liz­ing the sug­ars in the MRS plate. By all ac­counts, I was look­ing at BC30.

Colonies grew on all of the tea and tea bag plates, while my ster­ile con­trol plate re­mained bac­te­ria-free. The colonies from the boil­ing-wa­ter steeps and the tea bags were a va­ri­ety of sizes, in­clud­ing some that were sig­nif­i­cantly larger than oth­ers, while the colonies from the cold-wa­ter tea were uni­formly small.

Because I knew the vol­ume of tea I had put on each plate, I could cal­cu­late a stan­dard mea­sure­ment of bac­te­r­ial den­sity: colony-form­ing units (CFUs) per mil­li­liter. Contradictory to my ex­pec­ta­tions, I saw a five-fold in­crease in colonies from the tea steeped in boil­ing wa­ter rel­a­tive to the tea steeped in cold wa­ter for the four-minute con­di­tion. I saw the same pat­tern in the fif­teen-minute con­di­tion, with a nearly four-fold in­crease in boil­ing vs cold.

Before I could draw any con­clu­sions, I needed to know, for sure, that these colonies were Bacillus co­ag­u­lans. The most ro­bust way to check is by se­quenc­ing their DNA, but se­quenc­ing is ex­pen­sive. A sim­pler, cheaper way to check is with PCR, which am­pli­fies small re­gions of DNA unique to a species. I down­loaded the BC30 genome and se­lected two re­gions of its genome that did­n’t match other species in the NCBI data­base. Using Primer3, I gen­er­ated two pairs of PCR primers, short stretches of DNA to bind to ei­ther side of my re­gion of in­ter­est.

I picked the largest colony I could see from each plate (seven to­tal), sus­pended the cells in a small vol­ume of wa­ter, and set up stan­dard colony PCR re­ac­tions. The heat dur­ing the re­ac­tion bursts the cells, mak­ing the DNA avail­able for am­pli­fi­ca­tion. The com­pleted re­ac­tion was run through a porous gel with an elec­tric cur­rent and vi­su­al­ized with UV. If I saw bands on the gel for both primer sets, from to­tally dif­fer­ent parts of the BC30 genome, I could be con­fi­dent that this was, in fact, BC30.

I loaded the gel into the im­ager and hit run. There, in black re­lief against the grey back­ground of the gel, were my bands.

Bigelow knew some­thing I did­n’t5. It turns out that BC30, like many of its rel­a­tives, is a spore-form­ing bac­terium. When starved of nu­tri­ents, Bacillus co­ag­u­lans di­vides asym­met­ri­cally, pack­ing its ba­sic cel­lu­lar in­for­ma­tion into a spore with a thick pro­tec­tive coat. These spores are re­sis­tant to dessi­ca­tion, nu­tri­ent star­va­tion, ra­di­a­tion, chem­i­cal dis­in­fec­tants and ex­treme heat. It was these spores that were in the tea bag — spores that are per­fectly com­fort­able be­ing steeped in boil­ing wa­ter.

Like the seeds of plants, when the spores find them­selves in fa­vor­able con­di­tions for growth — say, on an MRS agar plate at a balmy 37oC — they ger­mi­nate back into ac­tively grow­ing cells. This is the premise of their abil­ity to func­tion as a pro­bi­otic. The spores are dor­mant and shelf-sta­ble in a tea bag, or any of the thou­sand prod­ucts ad­ver­tised to con­tain BC30, and will, in the­ory, ger­mi­nate upon ar­rival in the GI tract, where they can ex­ert some sort of ef­fect on the host that con­sumed them.

To pro­duce spores at scale, man­u­fac­tur­ers grow bac­te­ria in vats of nu­tri­ent-rich broth. If the nu­tri­ents are not re­plen­ished, the bac­te­ria even­tu­ally start to starve, trig­ger­ing the sporu­la­tion process. Around 24 hours later, the bac­te­r­ial broth is treated with en­zymes to kill any re­main­ing, non-sporu­lated cells. The mix­ture is con­cen­trated, washed with wa­ter, and fi­nally, in a fan­tas­tic twist of irony, pas­teur­ized.

The pitch for BC30 is that it im­proves digestive health” and protein ab­sorp­tion.” The re­ported end­points for di­ges­tive health on BC30s web­site are re­duc­tions in bowel move­ment fre­quency, ab­dom­i­nal pain and ab­dom­i­nal bloat­ing in adults with IBS. In the study pro­moted on the site, the base­line for the placebo group for ab­dom­i­nal pain and bloat­ing is, mys­te­ri­ously and re­spec­tively, 12.5% and 30% higher than the base­line for the BC30 treat­ment group. The placebo group ex­pe­ri­enced no change in sever­ity scores over the sub­se­quent course of treat­ment, while the BC30 group dropped to placebo lev­els af­ter a week and sta­bi­lized.

For one of the pro­tein ab­sorp­tion stud­ies, there is a small but sta­tis­ti­cally sig­nif­i­cant dif­fer­ence in amino acid lev­els in the blood, in­clud­ing when BC30 is paired with an­other one of its par­ent com­pa­ny’s prod­ucts, a nutritional milk pro­tein con­cen­trate” called Ultranor.

If these re­sults hold, they beg the ques­tion — could a prod­uct like Bigelow’s pro­bi­otic tea be able to pro­duce these ben­e­fi­cial ef­fects? Most of the clin­i­cal tri­als I could find, in­clud­ing the IBS study above, dosed peo­ple daily over the course of one to eight weeks with 1 bil­lion CFUs (spores) of BC30. Per my cal­cu­la­tions, a prop­erly steeped cup of pro­bi­otic tea yields around 30,000 CFUs: 0.003% of the clin­i­cally tested dose.

Granted, I am one per­son and this is one ex­per­i­ment. But there are in­de­pen­dent, con­flict­ing re­ports on whether BC30 sur­vives the GI tract at all. One study sug­gests that Bacillus pro­bi­otics don’t make it, while an­other re­ports about half of the ini­tial dose of spores sur­viv­ing tran­sit through an ar­ti­fi­cial hu­man gut sys­tem. Other re­search sug­gests that the ef­fect of the pro­bi­otic is not even due to the cells com­ing back to life, but due to an im­mune re­sponse against the dor­mant or veg­e­ta­tive cells.

These ob­ser­va­tions have con­se­quences for con­sumers be­ing parted from their money by un­sub­stan­ti­ated health claims. But they are in­ter­est­ing ob­ser­va­tions in their own right. Sporulating or­gan­isms’ im­per­vi­ous­ness to heat, while use­ful for com­mer­cial biotech ap­pli­ca­tions, causes prob­lems for the food in­dus­try. The food-poi­son­ing agent Bacillus cereus is a species nor­mally found in the soil. It can re­lease heat-re­sis­tant tox­ins if it mul­ti­plies in food, and live bac­te­ria can pro­duce tox­ins when they reach the small in­tes­tine. Even pas­teur­ized milk spoils even­tu­ally as heat-re­sis­tant spores, mostly soil Bacillus, be­gin to mul­ti­ply.

Bacillus co­ag­u­lans is also a soil bac­terium by na­ture6, and not a typ­i­cal res­i­dent of the com­mu­nity of mi­crobes in our gut (called the gut mi­cro­biome). It was dis­cov­ered in 1915, in canned milk that had spoiled and co­ag­u­lated. In spite of its ori­gins, BC30 seems in­ert as a pathogen, and re­ports of it caus­ing in­fec­tion are van­ish­ingly rare. BC30s safety track record is re­mark­able.

But safety is only the first step on the quest to use pro­bi­otics for good. New pro­bi­otic com­pa­nies like Pendulum and Seed mar­ket the fact that they are backed by clin­i­cal trial data — Pendulum for blood sugar con­trol in Type II di­a­betes, and Seed for gas, bloat­ing and reg­u­lar­ity in healthy adults. The back­bones of Pendulum and Seed’s prod­ucts are or­gan­isms found more com­monly in the gut, and, in­ter­est­ingly, both com­pa­nies fo­cus on multi-species prod­ucts, dos­ing pa­tients with minia­ture mi­cro­bial com­mu­ni­ties.

One fo­cus of my PhD lab is on ab­nor­mal path­o­genic be­hav­ior of nor­mally harm­less bac­te­r­ial res­i­dents of the gut, most com­monly in peo­ple who are al­ready quite sick. While un­happy mi­cro­bio­mes can be un­happy in their own way, we do not have a con­sen­sus on what a healthy” gut mi­cro­biome looks like, ei­ther. In col­lab­o­ra­tion with a con­ti­nent-wide con­sor­tium in Africa, our lab helped cat­a­logue the species found in healthy adult women across the con­ti­nent. We found over 1,000 new species rel­a­tive to what had been pre­vi­ously de­scribed in stud­ies fo­cused on Western coun­tries.

Companies try­ing to in­tro­duce tar­geted com­bi­na­tions of mi­crobes into the gut are thus for­ever shoot­ing at a mov­ing tar­get. Outside of spe­cific in­di­ca­tions for GI in­fec­tions, de­ter­min­ing whether to give (or take) a pro­bi­otic is a grey area. And for the com­mon GI com­plaints fo­cused on by the pro­bi­otic mar­ket, tar­get­ing the mi­cro­biome with ad­di­tional or­gan­isms may not be the an­swer at all. Rather, by un­der­stand­ing how bac­te­ria work to­gether in the mi­cro­biome, so­lu­tions may fa­vor chang­ing the meta­bolic en­vi­ron­ment of the gut to drive the for­ma­tion of species-ag­nos­tic guilds” that per­form spe­cific func­tions, likely via di­etary in­ter­ven­tions.

I never get tired of grow­ing bac­te­ria. For a colony to be vis­i­ble on a plate, it con­sists of at least a mil­lion, of­ten closer to a bil­lion, in­di­vid­ual cells. Learning how bac­te­ria grow and adapt does not di­min­ish the sense of won­der I feel when I ob­serve them — it only en­hances it. I like to think this same sense of won­der an­i­mated the sci­en­tist who first cul­tured Bacillus co­ag­u­lans out of canned milk that had spoiled. And I have to imag­ine some mix­ture of won­der, awe and hor­ror kept the Random Hall Milk alive for twenty-seven years.

The next Louis Pasteur could be a lac­tose-in­tol­er­ant un­der­grad, or a pro­cras­ti­nat­ing PhD stu­dent. It could be you. Pausing to look a lit­tle longer, to ask why the world is the way it is — this is how we up­end as­sump­tions of what is valu­able. What is worth look­ing at. Because in the end, trash is in the eye of the be­holder.

1

You can read the Milk’s ap­pli­ca­tion here.

2

As demon­strated by the Milk, even pas­teur­ized bev­er­ages spoil, a process sped up by ex­po­sure to the mi­crobes in the air but which will pro­ceed within an un­opened con­tainer any­way. How is this pos­si­ble?

It’s be­cause pas­teur­ized milk is not the same as ster­il­ized milk. Pasteurization heats to ~60C for a few min­utes. While the mi­crobes that we worry about caus­ing in­fec­tion can’t sur­vive this, some heat-tol­er­ant bac­te­ria and pro­teins can — those are what will even­tu­ally break down the milk, even if it is­n’t opened to the air. Heating to 140C, on the other hand, makes milk ef­fec­tively ster­ile, killing even the heat-tol­er­ant bac­te­ria. But this process, used to pro­duce ultra-high tem­per­a­ture” or UHT shelf-sta­ble milk, does some odd things to the pro­teins that sub­tly change the fla­vor, color, and tex­ture of the milk.

3

If a pro­bi­otic is mar­keted as a food or di­etary sup­ple­ment, as most are, then it is not re­quired to un­dergo a clin­i­cal trial in the U.S. but in­stead to sub­mit a GRAS no­ti­fi­ca­tion. The sec­ond most im­por­tant thing to know about the GRAS sys­tem is that it does not re­quire proof of ef­fi­cacy — only safety. The most im­por­tant thing to know is that the proof of safety is pro­vided by the com­pany re­quest­ing the GRAS des­ig­na­tion. This proof is then re­viewed by the FDA, to de­ter­mine whether the no­tice pro­vides a sufficient ba­sis for a GRAS de­ter­mi­na­tion” and whether information in the no­tice or oth­er­wise avail­able to FDA raises any safety con­cerns.

4

This is one of the most po­lar­iz­ing choices one can make as a bio­med­ical re­search sci­en­tist. The al­ter­na­tive to the hockey stick is to use glass beads that you au­to­clave then sprin­kle on the plate and roll around. People are very pas­sion­ate about their cho­sen method and will at­tempt to con­vert you.

5

Saw this weirdly ag­gres­sive Bigelow com­mer­cial at the gym. I don’t think they’re go­ing to spon­sor me af­ter this ar­ti­cle.

6

As our un­der­stand­ing of mi­croor­gan­isms evolves, so do our nam­ing and clas­si­fi­ca­tion con­ven­tions. Bacillus co­ag­u­lans is a more dis­tant rel­a­tive of Bacillus cereus and sim­i­lar soil mi­crobes than pre­vi­ously thought and has been re-clas­si­fied into a new genus called Weizmannia. Its full nomen­cla­ture his­tory can be found on the LPSN.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.