10 interesting stories served every morning and every evening.

www.data.jma.go.jp

topへ

New HIV vaccine shows unprecedented success in preclinical study

www.lji.org

Highlights:

Scientists at La Jolla Institute for Immunology, Scripps Research have de­vel­oped an HIV vac­cine that trains im­mune cells to see past HIVs de­fenses.

This HIV vac­cine works by prompt­ing the body’s im­mune sys­tem to make sub­stan­tial num­bers of rarely seen broadly neu­tral­iz­ing” an­ti­bod­ies.

In this new study, this vac­cine re­sulted in the best HIV-fighting an­ti­body re­sponse ever seen in pri­mates. Human tri­als have now started.

LA JOLLA, CA—A new HIV vac­cine de­vel­oped by La Jolla Institute for Immunology (LJI), Scripps Research sci­en­tists, and IAVI has the po­ten­tial to pro­tect hu­mans from de­vel­op­ing HIV in­fec­tion and AIDS. This HIV vac­cine is the first to gen­er­ate a high num­ber of broadly neu­tral­iz­ing,” virus-fight­ing an­ti­bod­ies in pri­mates.

This feels like a huge suc­cess,” says LJI Professor and Chief Scientific Officer Shane Crotty, Ph.D., who co-led the re­search with Scripps Research Professor William Schief, Ph.D. We con­structed a suc­cess­ful vac­cine from the ground up, which re­quired a deep un­der­stand­ing of the im­mune sys­tem.”

This ground­break­ing re­search, pub­lished in Nature, is the re­sult of 14 years of col­lab­o­ra­tion be­tween La Jolla Institute for Immunology and Scripps Research, as part of the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD). This has been one of those Apollo moon mis­sion-type pro­jects, where there is an ex­cep­tional goal and the team has to ac­com­plish a myr­iad of dis­cov­er­ies and in­ven­tions along the way,” says Crotty.

Outsmarting HIV

The new vac­cine works by in­ter­ven­ing in a process called B cell mat­u­ra­tion. B cells make an­ti­bod­ies. Like many im­mune cells, B cells have an early naive” stage be­fore they are ready to make an­ti­bod­ies. B cells start to ma­ture once they get the sig­nal that a pathogen, such as a virus, is try­ing to at­tack. B cells see pieces of that pathogen’s mol­e­c­u­lar struc­ture and start pro­duc­ing an­ti­bod­ies that can bind to that struc­ture and halt in­fec­tion.

It can take a lit­tle while for B cells to find the right bullseye” on a pathogen. But B cells keep try­ing. As they ma­ture, B cells tweak their an­ti­body pro­duc­tion, re­fin­ing an­ti­body struc­tures to bind to a pathogen in just the right, vul­ner­a­ble spots.

Scientists de­scribe B cell de­vel­op­ment as a train­ing process or boot­camp. In most cases, the body is left with a well-honed B cell army.

HIV is hard to beat be­cause it does­n’t give B cells a chance to de­velop ef­fec­tive an­ti­bod­ies. The first prob­lem is that HIV dis­guises it­self from the im­mune sys­tem. The virus is wrapped in an ever-shift­ing cloak of sugar mol­e­cules, called gly­cans. This lets HIV sneak un­de­tected past hu­man cells, which are also cov­ered in gly­cans.

The sec­ond big prob­lem is that HIV mu­tates very quickly. The world­wide di­ver­sity of HIV mu­ta­tions is ex­tra­or­di­nary. Even the di­ver­sity within one in­di­vid­ual per­son liv­ing with HIV is dra­matic,” says LJI Instructor Patrick Madden, Ph.D., who served as study co-first au­thor with Jon Steichen, Ph.D., an in­sti­tute in­ves­ti­ga­tor at Scripps Research.

The third prob­lem is that HIV changes its shape when it in­fects hu­man cells. Even if B cells get a glimpse of its vi­ral struc­ture—snap!—the struc­ture changes.

Taken to­gether, these prob­lems rarely give B cells a chance to hone their an­ti­body re­sponses against HIV. Even if a B cell man­ages to make neu­tral­iz­ing an­ti­bod­ies, the virus can mu­tate or change its shape, ren­der­ing those an­ti­bod­ies use­less.

The LJI and Scripps Research teams spent years hunt­ing for broadly neu­tral­iz­ing” an­ti­bod­ies that can ac­tu­ally bind to HIV and rec­og­nize key vi­ral struc­tures, even if the rest of the virus mu­tates. These an­ti­bod­ies are very, very rare, but they can be found in blood sam­ples from a small num­ber of peo­ple liv­ing with HIV.

An ef­fec­tive HIV vac­cine would need to prompt the im­mune sys­tem to make these same broadly neu­tral­iz­ing an­ti­bod­ies. How could we flip the whole im­mune re­sponse on its head so the rare re­sponses be­come the com­mon re­sponses? That was a crit­i­cal chal­lenge we faced,” says Crotty.

Testing the new vac­cine

It was time to go back to B cell boot­camp. The sci­en­tists stud­ied what made the HIV-fighting B cells spe­cial. Then they re­versed the process to see ex­actly how those B cells ma­tured. By look­ing back at the mat­u­ra­tion process, the re­searchers could track how the B cells changed when they saw spe­cific pieces of the HIV struc­ture.

The team dis­cov­ered that B cells ma­tured to make broadly neu­tral­iz­ing an­ti­bod­ies af­ter they got an early look at parts of HIVs outer envelope” pro­tein. Because these vi­ral sites sparked an im­mune re­sponse, sci­en­tists would call them antigens.”

An ef­fec­tive HIV vac­cine would likely need to in­clude mod­els of these anti­gens. The anti­gens would work like mugshots of America’s most wanted. If B cells saw those anti­gens early and of­ten, they would get re­ally good at rec­og­niz­ing and even neu­tral­iz­ing HIV. We were try­ing to mimic the pro­gres­sion of those neu­tral­iz­ing an­ti­bod­ies,” says Madden.

In a feat of mol­e­c­u­lar en­gi­neer­ing, the Schief Lab de­vel­oped vac­cine mol­e­cules that re­sem­bled the real HIV anti­gens. The sci­en­tists then worked with Emory National Primate Research Center, to test this po­ten­tial HIV vac­cine in a non-hu­man pri­mate species called rhe­sus macaques.

The re­searchers first ad­min­is­tered a priming” vac­cine meant to ac­ti­vate each an­i­mal’s naive B cells. The an­i­mals then re­ceived a se­ries of shepherding” booster shots to help their B cells de­velop along the right path.

This se­ries of vac­ci­na­tions will guide, or walk’, a B cell from its naive state to its broadly neu­tral­iz­ing state,” says Madden.

This new type of vac­cine ap­proach is called germline tar­get­ing” be­cause it tar­gets naive B cells in their germline” or naive form, be­fore they be­gin their train­ing process.

The sci­en­tists found that around 44 per­cent of the an­i­mals went on to pro­duce broadly neu­tral­iz­ing an­ti­bod­ies against HIV in their blood. These an­ti­bod­ies were im­pres­sively abun­dant.

We suc­ceeded in tak­ing ul­tra-rare an­ti­body re­sponses and turn­ing them into com­mon re­sponses by the end of the vac­ci­na­tion process,” adds Crotty. In other re­search re­cently pub­lished, they re­ported a new strat­egy to ac­cel­er­ate re­lated vac­cine an­ti­body re­sponses [See Nature Immunology pa­per].

The team did­n’t test whether these an­ti­bod­ies could pre­vent in­fec­tion, but it’s sig­nif­i­cant that these an­ti­bod­ies could be found in the blood, where they could en­counter and po­ten­tially block HIV.

Bringing the HIV vac­cine to hu­mans

The Crotty Lab plans to in­ves­ti­gate how they might change the booster shot reg­i­men to make the HIV vac­cine even more ef­fec­tive. It was in­cred­i­ble to get those re­sults, but of course we’d like to see a re­sponse in 100 per­cent of the an­i­mals,” says Madden.

Importantly, the an­ti­bod­ies found in the an­i­mal sub­jects re­sem­bled the ex­act kinds of broadly neu­tral­iz­ing an­ti­bod­ies seen in those rare hu­mans who made their own neu­tral­iz­ing an­ti­bod­ies. It’s clear that our im­mune sys­tems can make these pow­er­ful an­ti­bod­ies, given the right train­ing.

We be­lieve this vac­cine ap­proach is even more likely to suc­ceed in hu­mans, be­cause of the im­muno­genet­ics,” Crotty says.

The prim­ing im­muno­gen used in this study was eval­u­ated in hu­mans in the HVTN 144 trial and is cur­rently be­ing tested in the Phase 1 trial IAVI G004. IAVI, Scripps Research, the HIV Vaccine Trials Network, and part­ners are now ad­vanc­ing plans to fur­ther eval­u­ate the full im­mu­niza­tion reg­i­men in a fu­ture hu­man clin­i­cal study.

Additional au­thors of the study, Vaccination elic­its HIV broadly neu­tral­iz­ing an­ti­bod­ies in pri­mates,” in­clude Claudia T. Flynn, Swastik Phulera, Monolina Shil, Oleksandr Kalyuzhniy, Alessia Liguori, Carolyne Kifude, Leigh M. Sewall, Christopher A. Cottrell, Krystal M. Ma, Sabyasachi Baboo, Jolene K. Diedrich, Katherine McKenney, Allan C. de­Camp, Diane G. Carnathan, Ivy Phung, Parham Ramezani-Rad, Ester Marina-Zárate, Brian Freeman, Zhenfei Xie, Jeong Hyun Lee, Troy Sincomb, Nicole Phelps, Danny Lu, Diana Goodwin, Ryan Tingle, Yumiko Adachi, Nushin Alavi, Jenny Tran, Andy S. Tran, Alyne Nascimento, Catherine Sovie, Daniel L. V. Bader, Hannah Voic, Xiaoya Zhou, Grace Pixton, Agnes Walsh, Mariane B. Melo, Torben Schiffner, Facundo D. Batista, Dennis R. Burton, Darrell J. Irvine, James C. Paulson, John R. Yates III, Gabriel Ozorowski, Andrew B. Ward, Guido Silvestri.

This work was sup­ported by National Institute of Allergy and Infectious Diseases (NIAID), of the National Institutes of Health, through grant UM1 Al100663 to the Scripps Center for HIV/AIDS Vaccine Immunology and Immunogen Discovery (CHAVI-ID), grant UM1 AI144462 to the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD), P51 OD011132 to Emory National Primate Research Center, and R01 AI113867; by the  Gates Foundation un­der the Collaboration for AIDS Vaccine Discovery (NAC INV-007522, INV-008813, INV-034657, and INV-064772), via the IAVI Neutralizing Antibody Center (NAC); and by the National Institute of Health grant S10OD025052.

Initiative detail | European Citizens' Initiative

citizens-initiative.europa.eu

European Citizens’ Initiative

How to survive boiling water

taxa.substack.com

The story of MITs most no­to­ri­ous milk car­ton be­gins, as many good sto­ries do, in a col­lege dorm.

The milk in ques­tion was pur­chased in 1994 and re­dis­cov­ered in 1995 by an un­der­grad named Justin Cave. By that point, the re­port­edly lac­tose-in­tol­er­ant Cave had even less use for the milk he had aban­doned in his fridge ten months ear­lier. For rea­sons lost to his­tory, he did not throw the milk away. He threw it a birth­day party.

The Milk lived the rest of its life un­re­frig­er­ated, stored in a tall, sin­gle-walled jar. For twenty-seven years, the res­i­dents of Random Hall dorm gath­ered faith­fully to cel­e­brate its birth­day. At age 20, the Milk ap­plied to, and was re­jected from, MIT1. The jar was pe­ri­od­i­cally burped” to re­lease the gas pres­sure in­side, un­til the Milk reached its sta­ble fi­nal form — a cloudy brown liq­uid. When asked why the Milk was never thrown away, one res­i­dent of Random Hall replied: Why throw some­thing away when you can tell a story about it?”

Stuff I learned from things that nor­mally get thrown away” could be the ti­tle of many sci­en­tists’ mem­oirs, in­clud­ing Louis Pasteur’s. Winemaking pro­duces, well, wine, but it also pro­duces acidic crys­tals on the walls of the vats. These byprod­ucts were not dis­carded — they were stud­ied by Pasteur and his con­tem­po­raries. Pasteur’s ob­ser­va­tions both rev­o­lu­tion­ized our un­der­stand­ing of chem­istry and led him to the phe­nom­e­non that would de­fine his ca­reer and lay the foun­da­tion of mod­ern food safety: fer­men­ta­tion.

The mi­croor­gan­isms re­spon­si­ble for fer­men­ta­tion are vis­i­ble to our naked senses only through the tex­tures, col­ors and smells re­sult­ing from their col­lec­tive ef­forts. From Aristotle through the 1850s, it was as­sumed that some in­trin­sic prop­erty of a non-liv­ing start­ing sub­stance (like grain or milk) en­abled its spon­ta­neous fer­men­ta­tion into some­thing use­ful (like beer or yo­gurt) or its even­tual spoilage.

It was Pasteur who proved that liv­ing or­gan­isms were re­quired for the trans­for­ma­tions that took place dur­ing fer­men­ta­tion. He heated up nu­tri­ent-rich broths in cus­tom flasks that let gases, but not mi­crobes, flow in and out of the flasks. Pasteur then broke the neck off of one of the flasks to ex­pose the broth to the air. If the boiled broth could spon­ta­neously trans­form, it would do so with or with­out ex­po­sure to mi­crobes in the air and en­vi­ron­ment.

The flask with the neck bro­ken off grew cloudy and fer­mented as bac­te­ria bloomed, but the ster­ile one re­mained clear.

This find­ing was great news for Napoleon. The French were los­ing money, and per­haps more alarm­ingly, their rep­u­ta­tion, ex­port­ing wine to the British — the wine would spontaneously” go bad dur­ing ship­ping. The French gov­ern­ment of­fered a prize for a sci­en­tist to solve the case of the spoiled wine. With the knowl­edge of the mi­croor­gan­ism-dri­ven process of fer­men­ta­tion in hand, Pasteur did to the wine what he did to the broth, just more gen­tly — he heated the wine enough to kill mi­crobes with­out dam­ag­ing the wine’s fla­vor. Immortalized as pas­teur­iza­tion, this process was adapted shortly af­ter its in­ven­tion in 1865 to let us safely drink stored milk2.

Killing bac­te­ria thus be­came a ma­jor pre­oc­cu­pa­tion of mod­ern life. Its most vis­i­ble man­i­fes­ta­tion to­day might be tak­ing an­tibi­otics (first avail­able in the 1940s): there were ~700 an­tibi­otic pre­scrip­tions per 1000 peo­ple3 in 2024 ac­cord­ing to CDC data. A close sec­ond might be the dizzy­ing ar­ray of dis­in­fec­tant prod­ucts found in U.S. gro­cery stores.

What’s less vis­i­ble is the san­i­ti­za­tion in­fra­struc­ture that makes things like gro­cery stores or med­i­cine pos­si­ble at all. The com­pany Steris, one maker of high tem­per­a­ture, pres­sur­ized ster­il­iza­tion equip­ment and other med­ical in­stru­ments, is a $5 bil­lion an­nual rev­enue com­pany, with a $21 bil­lion mar­ket cap. The U.S. pas­teur­izes around 50 bil­lion liters of fluid milk every year. To pack­age salad greens like spinach, the greens are washed in a di­lute bleach so­lu­tion to kill any lin­ger­ing soil mi­crobes. I could go on.

But our war on bac­te­ria has its own war­ring in­dus­try. This in­dus­try has cap­tured the imag­i­na­tions of sci­en­tists, the food and bev­er­age in­dus­try, pharma com­pa­nies and doc­tors along with in­flu­encers, mar­ket­ing gu­rus and op­por­tunists of all fla­vors. This in­dus­try em­pha­sizes that some mi­crobes are friends, not foe, (true) and you should be eat­ing them in large quan­ti­ties, on pur­pose, all the time, and prefer­ably pay­ing more for prod­ucts that con­tain them (dubious). This is the pro­bi­otics in­dus­try.

Probiotic” is a bit of a mis­nomer — it means for life, or pro­mot­ing life, but the for­mal de­f­i­n­i­tion of a pro­bi­otic is an ac­tual liv­ing mi­croor­gan­ism. In sim­ple terms: tak­ing a pro­bi­otic is just eat­ing bac­te­ria on pur­pose. I say on pur­pose be­cause we con­sume mi­crobes ac­ci­den­tally all the time from our en­vi­ron­ment, largely obliv­i­ous to their ex­is­tence or ef­fects. The bac­te­ria we spend much of our time and en­ergy try­ing to kill are out­num­bered, at a species level, at least 1000 to 1 by a com­bi­na­tion of harm­less and ben­e­fi­cial bac­te­ria liv­ing in and on our bod­ies. It’s this lat­ter prop­erty of ben­e­fi­cial­ness that pro­bi­otics are try­ing to ex­ploit.

I say ex­ploit be­cause of a re­cent trip I took to the gro­cery store. I had a cold and was in search of lemon gin­ger tea. I bought a box of Bigelow, went home, boiled some wa­ter, poured it over a tea bag, waited a bit, added honey, took a sip, and al­most spit it out. The tea had its ex­pected notes of gin­ger, a hint of lemon, and some pow­dery, al­ka­line af­ter­taste that I could­n’t place. Frankly, it tasted ter­ri­ble. (Sorry, Bigelow).

I in­spected the box again. In my con­gested state, I had un­wit­tingly pur­chased a new of­fer­ing from the tea com­pany — Bigelow Lemon Ginger, with pro­bi­otics. What made this tea dif­fer­ent from all the other teas I hap­pily sipped on was that in ad­di­tion to nice-sound­ing things like lemon­grass and cin­na­mon, it con­tained bac­te­ria. Bacteria which I had just boiled, at a tem­per­a­ture 40oC hot­ter than pas­teur­iza­tion.

Did the tea taste bad be­cause I was drink­ing dead bac­te­ria wa­ter? And if that was the ul­ti­mate out­come of the nor­mal brew­ing process, why bother putting bac­te­ria in the tea at all?

I was at a lab happy hour when I men­tioned this to my PhD the­sis ad­vi­sor. I know, right?” she said, sud­denly an­i­mated. Probiotic teas taste SO BAD.” I was thrilled to have an­other wit­ness. Doesn’t it seem crazy to add in bac­te­ria that you’re just go­ing to boil and kill any­way?” I asked. Is it all a scam?” She was al­ready nod­ding. You have to won­der whether the bac­te­ria in the tea make it to the gut at all, and whether they do any­thing help­ful once they get there,” she said.

I’d be ly­ing if I said I re­mem­bered ex­actly what hap­pened next, or who sug­gested what. All I re­mem­ber is an idea. An idea to test this seem­ingly para­dox­i­cal mar­ket­ing tac­tic like the mi­cro­bi­ol­o­gists we are. The idea was sim­ple: What if we tried to grow the bac­te­ria from the tea bag, in the lab?

I went home. I stared at the box of bac­te­ria tea.

Why throw some­thing away when you can tell a story about it?

BC30™, the bac­te­r­ial strain in the tea, is short for Bacillus co­ag­u­lans GBI-30, 6086®. It re­ceived the FDAs GRAS (Generally Recognized as Safe4) des­ig­na­tion in 2012 and is found in over a thou­sand leading food, bev­er­age and pet food prod­ucts world­wide” ac­cord­ing to the pro­bi­otic’s web­site.

To coax these bac­te­ria to grow out of steeped tea, I needed to know three things:

Is this species safe to grow in the lab?

Is this species safe to grow in the lab?

What does it like to eat?

What does it like to eat?

What are its pre­ferred growth con­di­tions?

What are its pre­ferred growth con­di­tions?

In gen­eral, I try not to in­gest the bac­te­ria I grow in the lab — even ones with the low­est safety des­ig­na­tion, BSL-1. By na­ture of it be­ing a com­mer­cial pro­bi­otic, BC30 is both BSL-1 (safe to grow un­der nor­mal lab pre­cau­tions) and ed­i­ble.

But, I still would­n’t try this at home or eat bac­te­ria off of a cul­ture plate. Why? BC30s pre­ferred food source is not that dif­fer­ent from the pre­ferred food source of many other mi­croor­gan­isms: a sugar- and amino acid-rich nu­tri­ent medium called MRS (De Man, Rogosa and Sharpe) agar.

While MRS agar has some ad­just­ments to make it pref­er­en­tially ap­pe­tiz­ing to BC30 and its rel­a­tives, those rel­a­tives also in­clude Streptococcus pyo­genes (causes strep throat) and Bacillus cereus (causes food poi­son­ing). Without ster­ile tech­nique and rig­or­ous species-level con­fir­ma­tion, you can­not know for sure what is grow­ing on your plate.

With that said, the American Society for Microbiology’s blog sug­gested that were I suc­cess­ful in cul­tur­ing BC30, I would see growth of translu­cent white colonies on MRS agar plates af­ter 48 hours of in­cu­ba­tion at 30 – 33oC in the pres­ence of oxy­gen.

First, I needed to make tea.

I wanted the con­di­tions of the ex­per­i­ment to rep­re­sent a range of re­al­is­tic tea-drink­ing sce­nar­ios, from in­tended use to fla­grant im­pro­vi­sa­tion, and set up three steeps:

The Rule Follower — Brewed as di­rected for 4 min­utes in boil­ing wa­ter.

The Rule Follower — Brewed as di­rected for 4 min­utes in boil­ing wa­ter.

I for­got I made tea” — We’ve all been there. 15 min­utes, boil­ing wa­ter.

I for­got I made tea” — We’ve all been there. 15 min­utes, boil­ing wa­ter.

Cold brew an­ar­chist — Self-explanatory.

Cold brew an­ar­chist — Self-explanatory.

It was at this point I re­al­ized I needed a ster­ile-ish way to trans­port the steeped tea and tea bags from my house to the lab. Luckily, I had re­cently run a blind­folded vol­ume pour­ing ac­cu­racy com­pe­ti­tion at our de­part­men­tal re­treat and had left­over Falcon tubes still in their orig­i­nal pack­age. While the tea was def­i­nitely not ster­ile, I rea­soned that a lit­tle ex­tra asep­tic tech­nique would­n’t hurt. I poured the tea into the tubes over my kitchen stove, us­ing the open flame as a makeshift Bunsen burner.

In re­al­ity, lab came first. I had to make MRS agar plates be­fore I steeped the tea. Our lab does not use MRS broth very of­ten, and when I first looked for some all I found was a 10-year-old so­lid­i­fied block of MRS pow­der in our stock cab­i­net that was grow­ing large green spots in­side of its glass con­tainer. Behind it was one that looked mer­ci­fully nor­mal.

I mixed broth pow­der, agar and wa­ter in a glass bot­tle, loosely capped it, put it in a wa­ter bath and took it to the au­to­clave. Autoclave” is a nice word for gi­ant pres­sure cooker. Ours is made by the afore­men­tioned Steris. It rat­tled and hissed as its jaws opened to ac­cept my tray of cul­ture me­dia, which it then heated to 121oC for 45 min­utes, ster­il­iz­ing the liq­uid.

Back at my lab bench, when the molten MRS agar had cooled enough to han­dle, I lit a Bunsen burner next to a stack of empty plas­tic petri dishes and poured a layer of agar into each one. Left overnight, the plates so­lid­i­fied into nu­tri­ent-dense beds for BC30.

The next day, I took my tubes of tea to lab. I pipet­ted 400 mi­cro­liters (0.4mL) of each liq­uid tea con­di­tion onto a plate next to the Bunsen burner. I spread the liq­uid evenly across the plate with a hockey stick-shaped plas­tic spread­er5 and left the lids on the plates cracked open to dry near the flame.

But to an­swer my ques­tion, I needed one more test. If there were bac­te­ria in the tea bag ini­tially, but they died when boiled, then I might see bac­te­r­ial growth by plat­ing the dry in­gre­di­ents of an un­steeped tea bag, or the tea bag steeped in cold wa­ter. I cut open the tea bags and shook some of their con­tents onto the agar. I put my full set of plates, in­clud­ing a plain MRS plate to check its steril­ity, into the in­cu­ba­tor at 37oC — a stan­dard growth tem­per­a­ture, but a lit­tle warmer than rec­om­mended. I was skep­ti­cal that any­thing would grow. For the next two days, all I could do was wait and see.

The first thing I no­ticed when I took the plates out of the in­cu­ba­tor was the smell. I was in dis­be­lief when I saw lit­tle white colonies dot­ting al­most all the plates and opened one to get a closer look. A sickly sweet, gin­gery aroma wafted from the plate as I in­spected the translu­cent colonies — a byprod­uct of the bac­te­ria me­tab­o­liz­ing the sug­ars in the MRS plate. By all ac­counts, I was look­ing at BC30.

Colonies grew on all of the tea and tea bag plates, while my ster­ile con­trol plate re­mained bac­te­ria-free. The colonies from the boil­ing-wa­ter steeps and the tea bags were a va­ri­ety of sizes, in­clud­ing some that were sig­nif­i­cantly larger than oth­ers, while the colonies from the cold-wa­ter tea were uni­formly small.

Because I knew the vol­ume of tea I had put on each plate, I could cal­cu­late a stan­dard mea­sure­ment of bac­te­r­ial den­sity: colony-form­ing units (CFUs) per mil­li­liter. Contradictory to my ex­pec­ta­tions, I saw a five-fold in­crease in colonies from the tea steeped in boil­ing wa­ter rel­a­tive to the tea steeped in cold wa­ter for the four-minute con­di­tion. I saw the same pat­tern in the fif­teen-minute con­di­tion, with a nearly four-fold in­crease in boil­ing vs cold.

Before I could draw any con­clu­sions, I needed to know, for sure, that these colonies were Bacillus co­ag­u­lans. The most ro­bust way to check is by se­quenc­ing their DNA, but se­quenc­ing is ex­pen­sive. A sim­pler, cheaper way to check is with PCR, which am­pli­fies small re­gions of DNA unique to a species. I down­loaded the BC30 genome and se­lected two re­gions of its genome that did­n’t match other species in the NCBI data­base. Using Primer3, I gen­er­ated two pairs of PCR primers, short stretches of DNA to bind to ei­ther side of my re­gion of in­ter­est.

I picked the largest colony I could see from each plate (seven to­tal), sus­pended the cells in a small vol­ume of wa­ter, and set up stan­dard colony PCR re­ac­tions. The heat dur­ing the re­ac­tion bursts the cells, mak­ing the DNA avail­able for am­pli­fi­ca­tion. The com­pleted re­ac­tion was run through a porous gel with an elec­tric cur­rent and vi­su­al­ized with UV. If I saw bands on the gel for both primer sets, from to­tally dif­fer­ent parts of the BC30 genome, I could be con­fi­dent that this was, in fact, BC30.

I loaded the gel into the im­ager and hit run. There, in black re­lief against the grey back­ground of the gel, were my bands.

Bigelow knew some­thing I did­n’t6. It turns out that BC30, like many of its rel­a­tives, is a spore-form­ing bac­terium. When starved of nu­tri­ents, Bacillus co­ag­u­lans di­vides asym­met­ri­cally, pack­ing its ba­sic cel­lu­lar in­for­ma­tion into a spore with a thick pro­tec­tive coat. These spores are re­sis­tant to dessi­ca­tion, nu­tri­ent star­va­tion, ra­di­a­tion, chem­i­cal dis­in­fec­tants and ex­treme heat. It was these spores that were in the tea bag — spores that are per­fectly com­fort­able be­ing steeped in boil­ing wa­ter.

Like the seeds of plants, when the spores find them­selves in fa­vor­able con­di­tions for growth — say, on an MRS agar plate at a balmy 37oC — they ger­mi­nate back into ac­tively grow­ing cells. This is the premise of their abil­ity to func­tion as a pro­bi­otic. The spores are dor­mant and shelf-sta­ble in a tea bag, or any of the thou­sand prod­ucts ad­ver­tised to con­tain BC30, and will, in the­ory, ger­mi­nate upon ar­rival in the GI tract, where they can ex­ert some sort of ef­fect on the host that con­sumed them.

To pro­duce spores at scale, man­u­fac­tur­ers grow bac­te­ria in vats of nu­tri­ent-rich broth. If the nu­tri­ents are not re­plen­ished, the bac­te­ria even­tu­ally start to starve, trig­ger­ing the sporu­la­tion process. Around 24 hours later, the bac­te­r­ial broth is treated with en­zymes to kill any re­main­ing, non-sporu­lated cells. The mix­ture is con­cen­trated, washed with wa­ter, and fi­nally, in a fan­tas­tic twist of irony, pas­teur­ized.

The pitch for BC30 is that it im­proves digestive health” and protein ab­sorp­tion.” The re­ported end­points for di­ges­tive health on BC30s web­site are re­duc­tions in bowel move­ment fre­quency, ab­dom­i­nal pain and ab­dom­i­nal bloat­ing in adults with IBS. In the study pro­moted on the site, the base­line for the placebo group for ab­dom­i­nal pain and bloat­ing is, mys­te­ri­ously and re­spec­tively, 12.5% and 30% higher than the base­line for the BC30 treat­ment group. The placebo group ex­pe­ri­enced no change in sever­ity scores over the sub­se­quent course of treat­ment, while the BC30 group dropped to placebo lev­els af­ter a week and sta­bi­lized.

For one of the pro­tein ab­sorp­tion stud­ies, there is a small but sta­tis­ti­cally sig­nif­i­cant dif­fer­ence in amino acid lev­els in the blood, in­clud­ing when BC30 is paired with an­other one of its par­ent com­pa­ny’s prod­ucts, a nutritional milk pro­tein con­cen­trate” called Ultranor.

If these re­sults hold, they beg the ques­tion — could a prod­uct like Bigelow’s pro­bi­otic tea be able to pro­duce these ben­e­fi­cial ef­fects? Most of the clin­i­cal tri­als I could find, in­clud­ing the IBS study above, dosed peo­ple daily over the course of one to eight weeks with 1 bil­lion CFUs (spores) of BC30. Per my cal­cu­la­tions, a prop­erly steeped cup of pro­bi­otic tea yields around 30,000 CFUs: 0.003% of the clin­i­cally tested dose.

Granted, I am one per­son and this is one ex­per­i­ment. But there are in­de­pen­dent, con­flict­ing re­ports on whether BC30 sur­vives the GI tract at all. One study sug­gests that Bacillus pro­bi­otics don’t make it, while an­other re­ports about half of the ini­tial dose of spores sur­viv­ing tran­sit through an ar­ti­fi­cial hu­man gut sys­tem. Other re­search sug­gests that the ef­fect of the pro­bi­otic is not even due to the cells com­ing back to life, but due to an im­mune re­sponse against the dor­mant or veg­e­ta­tive cells.

These ob­ser­va­tions have con­se­quences for con­sumers be­ing parted from their money by un­sub­stan­ti­ated health claims. But they are in­ter­est­ing ob­ser­va­tions in their own right. Sporulating or­gan­isms’ im­per­vi­ous­ness to heat, while use­ful for com­mer­cial biotech ap­pli­ca­tions, causes prob­lems for the food in­dus­try. The food-poi­son­ing agent Bacillus cereus is a species nor­mally found in the soil. It can re­lease heat-re­sis­tant tox­ins if it mul­ti­plies in food, and live bac­te­ria can pro­duce tox­ins when they reach the small in­tes­tine. Even pas­teur­ized milk spoils even­tu­ally as heat-re­sis­tant spores, mostly soil Bacillus, be­gin to mul­ti­ply.

Bacillus co­ag­u­lans is also a soil bac­terium by na­ture7, and not a typ­i­cal res­i­dent of the com­mu­nity of mi­crobes in our gut (called the gut mi­cro­biome). It was dis­cov­ered in 1915, in canned milk that had spoiled and co­ag­u­lated. In spite of its ori­gins, BC30 seems in­ert as a pathogen, and re­ports of it caus­ing in­fec­tion are van­ish­ingly rare. BC30s safety track record is re­mark­able.

But safety is only the first step on the quest to use pro­bi­otics for good. New pro­bi­otic com­pa­nies like Pendulum and Seed mar­ket the fact that they are backed by clin­i­cal trial data — Pendulum for blood sugar con­trol in Type II di­a­betes, and Seed for gas, bloat­ing and reg­u­lar­ity in healthy adults. The back­bones of Pendulum and Seed’s prod­ucts are or­gan­isms found more com­monly in the gut, and, in­ter­est­ingly, both com­pa­nies fo­cus on multi-species prod­ucts, dos­ing pa­tients with minia­ture mi­cro­bial com­mu­ni­ties.

One fo­cus of my PhD lab is on ab­nor­mal path­o­genic be­hav­ior of nor­mally harm­less bac­te­r­ial res­i­dents of the gut, most com­monly in peo­ple who are al­ready quite sick. While un­happy mi­cro­bio­mes can be un­happy in their own way, we do not have a con­sen­sus on what a healthy” gut mi­cro­biome looks like, ei­ther. In col­lab­o­ra­tion with a con­ti­nent-wide con­sor­tium in Africa, our lab helped cat­a­logue the species found in healthy adult women across the con­ti­nent. We found over 1,000 new species rel­a­tive to what had been pre­vi­ously de­scribed in stud­ies fo­cused on Western coun­tries.

Companies try­ing to in­tro­duce tar­geted com­bi­na­tions of mi­crobes into the gut are thus for­ever shoot­ing at a mov­ing tar­get. Outside of spe­cific in­di­ca­tions for GI in­fec­tions, de­ter­min­ing whether to give (or take) a pro­bi­otic is a grey area. And for the com­mon GI com­plaints fo­cused on by the pro­bi­otic mar­ket, tar­get­ing the mi­cro­biome with ad­di­tional or­gan­isms may not be the an­swer at all. Rather, by un­der­stand­ing how bac­te­ria work to­gether in the mi­cro­biome, so­lu­tions may fa­vor chang­ing the meta­bolic en­vi­ron­ment of the gut to drive the for­ma­tion of species-ag­nos­tic guilds” that per­form spe­cific func­tions, likely via di­etary in­ter­ven­tions.

I never get tired of grow­ing bac­te­ria. For a colony to be vis­i­ble on a plate, it con­sists of at least a mil­lion, of­ten closer to a bil­lion, in­di­vid­ual cells. Learning how bac­te­ria grow and adapt does not di­min­ish the sense of won­der I feel when I ob­serve them — it only en­hances it. I like to think this same sense of won­der an­i­mated the sci­en­tist who first cul­tured Bacillus co­ag­u­lans out of canned milk that had spoiled. And I have to imag­ine some mix­ture of won­der, awe and hor­ror kept the Random Hall Milk alive for twenty-seven years.

The next Louis Pasteur could be a lac­tose-in­tol­er­ant un­der­grad, or a pro­cras­ti­nat­ing PhD stu­dent. It could be you. Pausing to look a lit­tle longer, to ask why the world is the way it is — this is how we up­end as­sump­tions of what is valu­able. What is worth look­ing at. Because in the end, trash is in the eye of the be­holder.

1

You can read the Milk’s ap­pli­ca­tion here.

2

As demon­strated by the Milk, even pas­teur­ized bev­er­ages spoil, a process sped up by ex­po­sure to the mi­crobes in the air but which will pro­ceed within an un­opened con­tainer any­way. How is this pos­si­ble?

It’s be­cause pas­teur­ized milk is not the same as ster­il­ized milk. Pasteurization heats to ~60C for a few min­utes. While the mi­crobes that we worry about caus­ing in­fec­tion can’t sur­vive this, some heat-tol­er­ant bac­te­ria and pro­teins can — those are what will even­tu­ally break down the milk, even if it is­n’t opened to the air. Heating to 140C, on the other hand, makes milk ef­fec­tively ster­ile, killing even the heat-tol­er­ant bac­te­ria. But this process, used to pro­duce ultra-high tem­per­a­ture” or UHT shelf-sta­ble milk, does some odd things to the pro­teins that sub­tly change the fla­vor, color, and tex­ture of the milk.

3

Author cor­rec­tion 7/28/26: I re­ported this ini­tially as 7 in 10 peo­ple were pre­scribed an­tibi­otics,” but a com­menter on an­other site pointed out that the orig­i­nal re­port data likely re­flects a smaller num­ber of peo­ple get­ting re­peat pre­scrip­tions, so I have re­ported the raw data from the CDC re­port in­stead.

4

If a pro­bi­otic is mar­keted as a food or di­etary sup­ple­ment, as most are, then it is not re­quired to un­dergo a clin­i­cal trial in the U.S. but in­stead to sub­mit a GRAS no­ti­fi­ca­tion. The sec­ond most im­por­tant thing to know about the GRAS sys­tem is that it does not re­quire proof of ef­fi­cacy — only safety. The most im­por­tant thing to know is that the proof of safety is pro­vided by the com­pany re­quest­ing the GRAS des­ig­na­tion. This proof is then re­viewed by the FDA, to de­ter­mine whether the no­tice pro­vides a sufficient ba­sis for a GRAS de­ter­mi­na­tion” and whether information in the no­tice or oth­er­wise avail­able to FDA raises any safety con­cerns.

5

This is one of the most po­lar­iz­ing choices one can make as a bio­med­ical re­search sci­en­tist. The al­ter­na­tive to the hockey stick is to use glass beads that you au­to­clave then sprin­kle on the plate and roll around. People are very pas­sion­ate about their cho­sen method and will at­tempt to con­vert you.

6

Saw this weirdly ag­gres­sive Bigelow com­mer­cial at the gym. I don’t think they’re go­ing to spon­sor me af­ter this ar­ti­cle.

7

As our un­der­stand­ing of mi­croor­gan­isms evolves, so do our nam­ing and clas­si­fi­ca­tion con­ven­tions. Bacillus co­ag­u­lans is a more dis­tant rel­a­tive of Bacillus cereus and sim­i­lar soil mi­crobes than pre­vi­ously thought and has been re-clas­si­fied into a new genus called Weizmannia. Its full nomen­cla­ture his­tory can be found on the LPSN.

Exclusive | Netflix exec goes ballistic after being fired for stunning 'trust exercise' confession at retreat: suit

nypost.com

A Netflix ex­ec­u­tive was fired from his $1.1 mil­lion a year job af­ter re­veal­ing dur­ing a trust ex­er­cise” at a work re­treat that he had taken med­ically pre­scribed ke­t­a­mine, a law­suit has claimed.

Kevin Baillie, who was vice pres­i­dent and head of cre­ative at Eyeline Studios, is su­ing the com­pany af­ter it launched an in­ves­ti­ga­tion into his com­ments that ul­ti­mately ended in his fir­ing, the pa­pers say.

Baillie, who’s been on the vi­sual ef­fects team for Pirates of the Caribbean” and the Harry Potter fran­chise, says he took the drug un­der med­ical su­per­vi­sion in October and November of 2022 at a Santa Barbara clinic.

He sought the treat­ment for clin­i­cal de­pres­sion af­ter the death of his mother, ac­cord­ing to the suit.

During what’s called a Vulnerability-Trust ex­er­cise” at a January 2026 re­treat at the ex­clu­sive Sendero Ranch, a Northern California prop­erty owned by Netflix, Baillie shared with his col­leagues that he had un­der­gone the treat­ment, the suit says.

Baillie claims he ex­plained the rea­son why he had taken the drug but was in­ves­ti­gated by Netflix. On March 18, 2026 a com­pany in­ves­ti­ga­tor brought the in­ci­dent up, in a man­ner sug­gest­ing sus­pi­cion of recre­ational drug use,” the suit reads.

The ex­ec­u­tive was fired in April with the com­pa­ny’s at­tor­ney con­firm­ing the ke­t­a­mine ther­apy is­sue has fac­tored into the ter­mi­na­tion,” the pa­pers say, and go on to sug­gest Baille was de­nied up to a year of sev­er­ance pay.

Baillie says in the law­suit the scope” of the in­ves­ti­ga­tion re­lated to al­leged pro­fan­ity and drink­ing. He had been warned dur­ing his per­for­mance re­view that he should drop one or two less f-bombs but don’t stop en­tirely.”

It goes on to say that at the same re­treat Baillie drank a Guinness stand­ing on his head, af­ter shar­ing that he had learned the trick from his for­mer fa­ther in law dur­ing a con­ver­sa­tion in­spired by the trust ses­sion.

His col­league im­me­di­ately asked for a demon­stra­tion, rather than with­hold the open­ness that the ses­sion had en­cour­aged, he per­formed the trick,” the pa­pers say.

Sign up for the California Morning Report newslet­ter

California’s top news, sports and en­ter­tain­ment de­liv­ered to your in­box every day.

Thanks for sign­ing up!

Baille also paints a pic­ture of an al­legedly al­co­hol-fu­eled com­pany en­vi­ron­ment en­cour­aged by the Eyeline Studios CEO Jeff Shapiro.

The law­suit al­leged alcohol con­sump­tion was com­pany-spon­sored, lead­er­ship-mod­eled and con­doned”, with mul­ti­ple ex­am­ples of how Shapiro set the cul­tural tone con­cern­ing al­co­hol at the ex­ec­u­tive level”.

The doc­u­ments claimed Shapiro on one oc­ca­sion pur­chased beer at a cor­ner store and brought it to a com­pany car ride to the Visual Effects Society Awards for staff to share.

Baillie claims he also wit­nessed the CEO con­sume al­co­hol with Netflix and Eyeline em­ploy­ees at his own wel­come din­ner in September 2024, the Netflix Annual Business Review events in March 2025 and even a Lakers game at­tended by Netflix ex­ecs in­clud­ing Shapiro’s di­rect su­per­vi­sor in February 2026.

All up, Baillie’s at­tor­neys pro­vided over half a dozen ex­am­ples of the CEO be­ing pre­sent at a work event with a drink in his hand, ac­cord­ing to court pa­pers.

Download The California Post App, fol­low us on so­cial, and sub­scribe to our newslet­ters

California Post News: Facebook, Instagram, TikTok, X, YouTube, WhatsApp, LinkedInCalifornia Post Sports Facebook, Instagram, TikTok, YouTube, XCalifornia Post Opinion California Post Newsletters: Sign up here!Cal­i­for­nia Post App: Download here!Home de­liv­ery: Sign up here!Page Six Hollywood: Sign up here!

In ad­di­tion to host­ing mul­ti­ple par­ties, Shapiro also had a per­sonal bar in his of­fice from which he served al­co­hol (to Baille) in­clud­ing af­ter a suc­cess­ful meet­ing with Netflix’s CEO Ted Sarandos,” ac­cord­ing to the pa­pers.

Baillie is ask­ing for a jury trial, com­pen­satory dam­ages, lost wages, dam­ages for emo­tional dis­tress, and puni­tive dam­ages.

Netflix and Eyeline were con­tacted for com­ment.

inc.com

www.inc.com

Please en­able JS and dis­able any ad blocker

Substack writers, you need a website!

elizabethtai.com

But I al­ready have a web­site on Substack,” you ar­gue.

No, no, Substack is just a dis­tri­b­u­tion tool to am­plify your web­site. It should not be your dig­i­tal home.

In the last few years, I’ve no­ticed a pat­tern of writ­ers leav­ing their web­sites to make Substack their dig­i­tal home.

Now, it’s kinda okay if they have bought a do­main and linked it to Substack. (Meaning, it’s bet­ter than noth­ing.)

Rachel from Conscious Living is a good ex­am­ple. This way, Substack more or less func­tions like a con­tent man­age­ment sys­tem (CMS) for you.

However, com­pared to other CMS it’s very lim­ited, such as the abil­ity to man­age your SEO and cus­tomize your pages to add more fea­tures, but I di­gress. If you just want a fuss-free plat­form, this is one way to get it and Substack’s con­di­tions for do­mains are very rea­son­able and cost-ef­fi­cient. As I will ex­plain later, this could change on a dime with­out warn­ing.

However, there are some writ­ers who are say­ing: Hey read­ers, I’m now writ­ing on Substack, so head on over there (and ig­nore my web­site)!”

Some writ­ers do have a web­site, but link to their Substacks, call­ing them their blogs”. If your Substack has a do­main name they own, it’s okay, but if it’s xx.sub­stack.com, Substack is say­ing All your con­tent are be­long to us”.

In con­clu­sion: Writers, don’t do this. It’s short-sighted and un­wise and can de­rail your long-term vis­i­bil­ity on the Internet.

The siren call of con­ve­nience

Every few years, the in­ter­net con­vinces writ­ers that a new dig­i­tal par­adise has ar­rived. First, it was so­cial me­dia like Facebook. Then blog­ging net­works like Tumblr. Then it was Medium. More re­cently, it’s been Substack.

Platforms promise us an ea­ger au­di­ence, built-in mon­e­ti­za­tion, a smooth user in­ter­face, and a sup­port­ive com­mu­nity. As a writer who just wants to fo­cus on writ­ing, it’s in­cred­i­bly tempt­ing to hand over the keys to our cre­ative king­doms and let these por­tals han­dle every­thing. (Believe me, I gave in at one point. For years, I just stopped blog­ging al­to­gether and even gave up a do­main that had high traf­fic! But I got back in 2012 and never left.)

However, this is the truth that has not changed since the dawn of the Internet: When you build your au­di­ence en­tirely on some­one else’s plat­form, you aren’t a home­owner. You are a ten­ant. Or worse, a dig­i­tal share­crop­per.

And cor­po­rate land­lords al­ways change the rules even­tu­ally. It’s not per­sonal, it’s just busi­ness.

The il­lu­sion of the safe space

It’s easy to feel se­cure when a plat­form is in its golden era. But we’ve watched the down­falls of Twitter, the pol­icy shifts of Reddit, and the chang­ing tides of al­go­rith­mic net­works. Relying blindly on a cen­tral­ized por­tal not owned by you means your life’s work can al­ter overnight based en­tirely on a cor­po­rate board­room de­ci­sion.

When I looked at how frag­ile our dig­i­tal ecosys­tems re­ally are, I re­al­ized I needed a space that would­n’t go poof” be­cause a com­pany needed to please its in­vestors or share­hold­ers. This re­al­iza­tion com­pletely changed my ap­proach, push­ing me to pro­tect my con­tent by learn­ing to blog the IndieWeb way.

Your writ­ing needs a per­ma­nent home­base—a do­main that you own and con­trol. Full stop.

Moving from renting” to syn­di­cat­ing

The biggest push­back I hear from writ­ers is: But my web­site does­n’t have an au­di­ence! Substack does.”

But you don’t have to com­pletely aban­don so­cial me­dia or plat­forms like Substack to pro­tect your au­ton­omy (personally, I pre­fer the word sov­er­eignty but it does sound a tad dra­matic).

You just need to change the or­der of op­er­a­tions. Instead of pub­lish­ing di­rectly to a por­tal, you can shift your mind­set to POSSE: Publish (on your) Own Site, Syndicate Elsewhere. (I ex­plain the POSSE/PESOS method in an older post.)

By treat­ing your web­site as the de­fin­i­tive source of truth and us­ing plat­forms sim­ply as dis­tri­b­u­tion pipes, you get the best of both worlds. I dug deep into this shift when I com­mit­ted to be­ing an im­per­fect gar­dener of my dig­i­tal gar­den, ex­plor­ing how a less mar­ket-y way of pre­sent­ing my con­tent on­line let me share my wild gar­den of thoughts with­out danc­ing to the al­go­rithm.

A re­al­ity check on plat­form hype

If you are still hold­ing out hope that Substack is different” from the so­cial plat­forms that came be­fore it, let’s look at the num­bers and be­hav­iors be­hind the mar­ket­ing copy.

After spend­ing a sig­nif­i­cant amount of time ob­serv­ing the plat­form ecosys­tem first­hand, I wrote a bru­tally hon­est take­away in What I learned from one year of Substack. The net­work ef­fects are real, but so is the pres­sure to con­form to what the plat­for­m’s ecosys­tem fa­vors.

This post, by the way, des­per­ately needs to be up­dated be­cause things have got­ten much, much worse since I wrote it.

When you hand your con­tent over to a plat­form, you have to con­form to their rules and their lo­cal­ized bi­ases. For those of us writ­ing from out­side the dom­i­nant US-centric echo cham­bers, plat­form al­go­rithms heav­ily pri­or­i­tize spe­cific west­ern nar­ra­tives, mak­ing it in­cred­i­bly tough for lo­cal­ized or mi­nor­ity voices to be seen un­less they con­form.

I wrote about this ex­act frus­tra­tion re­cently in Linkblog: Dwelling on the Internet, high­light­ing how al­go­rith­mic com­pla­cency forces us into ho­mog­e­nized bub­bles.

The flip side — the writ­ers who re­fused to leave their web­sites

Each time there’s a new drama on some plat­form, and writ­ers are shak­ing their sabers and de­clar­ing that they will leave for yet an­other so­cial me­dia plat­form they don’t con­trol, I think about writ­ers like John Scalzi.

As of date, John scalzi has been blog­ging on https://​what­ever.scalzi.com/ for 28 years!

This sci-fi nov­el­ist has main­tained a sin­gle in­de­pen­dent web­site con­tin­u­ously for nearly three decades; this makes him one of the longest-run­ning, most con­sis­tent orig­i­nal blog­gers on the in­ter­net. Imagine the amount of dig­i­tal foot­print on that web­site! Unbroken by time or plat­forms.

(Specifically, he uses word­press.com like I do, as we both don’t want to bother with the pain of set­ting up your own self-hosted word­press web­site and just want the folks at Automattic to do it.)

He blogs in the clas­sic Indieweb way, though I doubt he is even aware he’s do­ing it. He treats his so­cial me­dia chan­nels such as X or Bluesky as a way to am­plify his web­site. All roads lead back to https://​what­ever.scalzi.com/, and this is some­thing I wish every sin­gle writer would do.

He wrote re­cently in Various & Sundry, 6/3/26:

this site acts as my own in­sti­tu­tional mem­ory, if I post some­thing about it here it con­sti­tutes an of­fi­cial record. I mean, all the posts I ever placed on the for­mer Twitter are now en­tirely lost to time, since I have gone in and purged my en­tire time­line there. This site, how­ever, en­dures. — John Scalzi

this site acts as my own in­sti­tu­tional mem­ory, if I post some­thing about it here it con­sti­tutes an of­fi­cial record. I mean, all the posts I ever placed on the for­mer Twitter are now en­tirely lost to time, since I have gone in and purged my en­tire time­line there. This site, how­ever, en­dures. — John Scalzi

Breaking free from plat­form blues

Trying to adapt your pres­ence across var­i­ous plat­forms in an ever-shift­ing dig­i­tal land­scape is ex­haust­ing. One minute a plat­form is a writer’s dar­ling; the next, it’s be­ing boy­cotted. Railing against a plat­for­m’s fo­cus shift or the pres­ence of (long sigh) Nazis is a use­less en­deavor.

As I noted in Linkblog March 12, 2026: Platform blues, chas­ing plat­form pu­rity is an il­lu­sion. Tech will change, cor­po­rate al­go­rithms will con­tinue to pri­or­i­tize profit over hu­man con­nec­tion, and plat­forms will con­tinue to cy­cle through hype and de­cline.

The an­ti­dote to this ex­haus­tion is­n’t mov­ing to the next shiny new app. It’s an­chor­ing your work on an in­de­pen­dent web­site with open dis­tri­b­u­tion chan­nels like RSS. It also means ruth­lessly us­ing plat­forms as dis­tri­b­u­tion chan­nels. When one col­lapses or you pre­fer to just move, it’s easy to just change strate­gies be­cause your dig­i­tal home re­mains un­changed.

Use plat­forms to find your read­ers, but bring them back to your house. It’s time to stop dig­i­tal share­crop­ping on rented land.

Featured photo is by vivek vk on Unsplash

Using an open model feels surprisingly good

matthewsaltz.com

July 27, 2026

I’ve been us­ing Claude and ChatGPT like the next guy for prob­a­bly two years now. I’ve never been a huge open soft­ware” nerd or any­thing like that. But just now, I got open­code work­ing on my own in­fer­ence end­point and… it felt sur­pris­ingly good. It feels… free­ing, some­how. I own the end­point, and my data just goes from my lap­top to there and back. It feels like it’s mine. It’s re­ally nice.

The mo­ti­va­tion for this was that I just got home and wanted to start on a lit­tle side pro­ject, but I don’t have the best Claude or ChatGPT plan for my per­sonal ac­count. I work at Modal, and to­day we just launched Kimi K3 on man­aged end­points, and I know Kimi K3 is sup­posed to be pretty solid, so in­stead of up­grad­ing my Claude plan, I wanted to give it a try. (I did­n’t di­rectly con­tribute to this fea­ture, so I haven’t got­ten to play with it yet.)

Within about 5 min­utes, I had open­code pointed at my own Modal end­point and run­ning. Spinning up open­code, I just felt a nice, empty blank­ness. The best way I can de­scribe it is like open­ing vim af­ter spin­ning a bunch of time in a big fancy ed­i­tor. Or maybe like the ten­drils ty­ing me to other providers had been cut, and I could breathe freely. I’m be­ing some­what dra­matic here but not ex­ag­ger­at­ing that much. It’s weird, lol, and un­ex­pected to me, which is why I wanted to write about it.

The Rise of Intelligence Ownership

fermisense.com

Cost vs. qual­ity on cat­a­log in­tegrity

Part I

Everyone asked the same ques­tion

Since ChatGPT launched in 2022, busi­ness lead­ers have been ask­ing the same ques­tion: what can AI do for us? The an­swer be­gan with low-risk tasks: sum­ma­riz­ing doc­u­ments, draft­ing emails, pro­duc­ing first drafts that a hu­man would edit.

It quickly moved into higher-value cog­ni­tive work, such as soft­ware de­vel­op­ment and con­tent gen­er­a­tion, and grew into more am­bi­tious pro­jects, like at­tempts to build an AI com­pany brain, a sys­tem con­nected to in­ter­nal knowl­edge, data, and tools that could co­or­di­nate work and even­tu­ally op­er­ate parts of the busi­ness au­tonomously.

While a lot of time, en­ergy and to­kens have been in­vested in AI adop­tion, mea­sur­able out­comes have barely been achieved at scale. However, some com­pa­nies em­braced be­ing AI-first and saw enor­mous gains in pro­duc­tiv­ity, rev­enue, and cost, while oth­ers lagged be­hind or failed to change their or­ga­ni­za­tions enough to reach high ROI.

Recent data from cor­po­rate ex­pense man­age­ment plat­form Ramp re­veals a stark con­trast in per­for­mance: the top quar­tile of com­pa­nies in­vest­ing in AI saw their rev­enue more than dou­ble be­tween November 2022 and December 2025, while busi­nesses with zero AI ex­pen­di­ture ex­pe­ri­enced a mere 15% in­crease.

Top AI adopters out­per­form across in­dus­tries

top quar­tile of AI spenders busi­nesses with zero AI spend

There are many rea­sons why AI has done won­ders for some com­pa­nies while oth­ers have strug­gled to see the re­turn on their in­vest­ment, but re­search pri­mar­ily points in five di­rec­tions.

01

Redesign the process, not just the task

Becoming AI-first means re­think­ing how the work is struc­tured, not drop­ping a model into a work­flow built around peo­ple: what gets ap­proved, who re­views what, and which hand­offs still need a hu­man. Where the process stays un­touched, legacy bot­tle­necks ab­sorb the pro­duc­tiv­ity gains be­fore they reach the P&L. In McKinsey’s 2025 sur­vey of or­ga­ni­za­tions us­ing gen AI, work­flow re­design was the at­tribute most cor­re­lated with EBIT im­pact, and only 21% of them had re­designed any work­flow at all.

02

Incentivize ex­per­i­men­ta­tion

Models, tool­ing and best prac­tices change weekly, so last quar­ter’s setup is rarely still the right one. That only gets picked up if peo­ple are re­warded for try­ing things and re­port­ing what failed, not just for ship­ping. Technical teams are the nat­ural place to start, since they see the same prob­lems re­cur across func­tions and can tell which of them a model can ac­tu­ally take over.

03

Provide tai­lored busi­ness con­text

Prompt en­gi­neer­ing and re­trieval can in­ject busi­ness con­text at call time, but do­ing it well is its own en­gi­neer­ing pro­gram: get­ting to the data, en­forc­ing ac­cess con­trols on what each re­quest may see, build­ing re­trieval that sur­faces the right ev­i­dence, and man­ag­ing a con­text win­dow that mod­els use un­evenly as it grows.

04

Measure us­age and im­pact

Every AI line item even­tu­ally meets the CFO ques­tion: what did this change, and was it worth it? In most de­ploy­ments, no­body can an­swer it: there is no in­fra­struc­ture to track the mod­el’s per­for­mance, de­ci­sion costs, or im­pact on ef­fi­ciency, and self-re­ported time sav­ings are of­ten in­ac­cu­rate. Without a scored eval­u­a­tion on your own data, a vibe eval­u­a­tion” is the ceil­ing of what you can claim, and a hard bud­get to de­fend.

05

In this ar­ti­cle we give a de­tailed overview of the de­ploy­ment tech­nique the win­ning group keeps con­verg­ing on: fine-tun­ing open-source mod­els with re­in­force­ment learn­ing. We cover how it ad­dresses the last three chal­lenges above, and how it turns knowl­edge only your or­ga­ni­za­tion has (namely data, tools, and processes) into a model no ven­dor API can match at a frac­tion of the cost.

TL;DR

2.2×

Revenue growth of the top quar­tile of AI spenders be­tween November 2022 and December 2025 in Ramp’s data. Companies with zero AI spend grew about 15% over the same three years, in the same econ­omy: the heavy adopters grew eight times as much.

1

Playbook the win­ners con­verge on: an open-source model, pro­pri­etary task data, and re­in­force­ment learn­ing against a scored copy of the work­flow. Bridgewater’s trained model makes ~30% fewer mis­takes than the best fron­tier model, Harvey’s le­gal agent beats GPT-5.5 and Claude Opus 4.8 on its own rubrics, and Intercom’s Fin Apex re­solves more sup­port is­sues at lower cost.

87.3%

Share of the max­i­mum achiev­able score our GRPO-trained 9B open-source model reached on cat­a­log re­view, vs 76.9% for the best fron­tier con­fig­u­ra­tion: a 13.5% rel­a­tive im­prove­ment over the fron­tier, and 36% over its own un­trained base (64.2%). The five fron­tier mod­els, even with op­ti­mized prompts, plateaued within a tenth of a point of each other; the trained spe­cial­ist cleared that ceil­ing.

68×

Cost ad­van­tage per re­viewed list­ing: $0.50 per 1,000 with the spe­cial­ist vs $34 with the strongest fron­tier model, and still 40× cheaper than the least ex­pen­sive fron­tier op­tion. At roughly 40 mil­lion de­ci­sions a day, that is about $7M a year in­stead of $500M, a 98% cost re­duc­tion.

Part II

What the win­ners do dif­fer­ently

Most of the com­pa­nies pulling ahead in the AI race made the same dis­cov­ery: own­ing your in­tel­li­gence wins on both per­for­mance and cost. A model trained to com­plete your spe­cific work­flows in your spe­cific en­vi­ron­ment is very likely to out­per­form a gen­eral-pur­pose model that has never seen in­side your com­pany. Additionally, since you do not need to pack as many gen­eral-pur­pose ca­pa­bil­i­ties into a model that is meant to op­er­ate in a spe­cific en­vi­ron­ment, you can of­ten get away with a smaller model that is or­ders of mag­ni­tude cheaper to run.

Owning your in­tel­li­gence does not mean can­celling the ChatGPT or Claude sub­scrip­tion. Most work­flow au­toma­tion still starts with fron­tier mod­els, and that is the right first move: it es­tab­lishes a base­line for what is tech­ni­cally pos­si­ble, and every call gen­er­ates the data (inputs, de­ci­sions, cor­rec­tions) that a spe­cial­ist model later trains on. Once the au­toma­tion leaves the pro­to­typ­ing stage, the pri­or­ity flips to cost and per­for­mance at vol­ume, and that is where fine-tun­ing open-source mod­els with re­in­force­ment learn­ing comes in.

In ad­di­tion to that, your Fable 5 or ChatGPT model can call the spe­cial­ist model to han­dle the parts of the work­flow that re­quire your in­ter­nal knowl­edge, and the spe­cial­ist model can call the fron­tier model for tasks that re­quire high gen­eral abil­ity.

How a model learns to op­er­ate in your en­vi­ron­ment

Over the past two years this has hard­ened into a play­book: an open-source model, pro­pri­etary task data, and a re­in­force­ment-learn­ing stage against a scored ver­sion of the work­flow. Below, we dis­cuss three sce­nar­ios where this ap­proach has been ap­plied to real-world tasks.

Bridgewater Associates is one of the largest hedge funds in the world. Its an­a­lysts sift a con­stant stream of ar­ti­cles, fil­ings, and emails, judg­ing which doc­u­ments are rel­e­vant to the fir­m’s in­vest­ment the­sis and where boil­er­plate con­tent be­gins. The catch is that rel­e­vant means rel­e­vant by Bridgewater’s in­ter­nal judg­ment, and no amount of prompt­ing got fron­tier mod­els to ab­sorb that judg­ment re­li­ably. So, the com­pany de­cided to train an open-source model on la­bels from its own ex­pert in­vestors. The trained model makes roughly 30% fewer mis­takes than the best fron­tier model, at a frac­tion of the in­fer­ence cost.

Harvey builds AI agents for law firms. Its hard­est work­loads are long-hori­zon: trans­ac­tion due dili­gence and le­gal memo draft­ing, where the agent nav­i­gates large doc­u­ment sets, er­rors com­pound across steps, and even the best fron­tier mod­els at max­i­mum rea­son­ing ef­fort kept falling short of the qual­ity bar. Harvey ran re­in­force­ment learn­ing on an open-weight model and got a le­gal agent that out­per­forms both GPT-5.5 and Claude Opus 4.8 on its rubrics.

Intercom is a cus­tomer-ser­vice plat­form whose AI agent, Fin, re­solves al­most two mil­lion cus­tomer is­sues a week. At that vol­ume the prob­lem is unit eco­nom­ics: fron­tier per-call pric­ing adds up fast, and every point of res­o­lu­tion rate mat­ters. So Intercom’s AI group post-trained its own ver­ti­cal sup­port model, Fin Apex, on bil­lions of cus­tomer-ser­vice in­ter­ac­tions. Intercom re­ports that it re­solves more is­sues than the best fron­tier mod­els while be­ing cheaper to run.

The same shape re­peats well be­yond these three. The ap­pen­dix col­lects eight more de­ploy­ments, with what each model was trained to do and what changed once it shipped.

One de­ploy­ment pat­tern re­peats across these cases: re­in­force­ment learn­ing pushes an open-source model past the fron­tier on a spe­cific set of work­flows, at a fixed and dra­mat­i­cally lower cost per call. The de­ploy­ment typ­i­cally starts by us­ing prompt and con­text en­gi­neered fron­tier mod­els to es­tab­lish the strongest de­fault base­line. Those fron­tier traces and learn­ings are then reused to pack the fo­cused ca­pa­bil­ity into a com­pact model the com­pany owns.

Part III

Case Study: Catalogue Integrity Agent

One of the cen­tral chal­lenges of run­ning an e-com­merce plat­form is keep­ing the prod­uct cat­a­log trust­wor­thy. Every list­ing must land in the cor­rect cat­e­gory of the plat­for­m’s tax­on­omy, and the at­trib­utes that power search, fil­ters, rec­om­men­da­tions, and down­stream op­er­a­tions must be ac­cu­rately ex­tracted from its im­ages and de­scrip­tion.

+15%our fine-tuned 9B open-source model re­views e-com­merce prod­uct list­ings more ac­cu­rately than the best fron­tier model we tested

68×cheaper per list­ing when our fine-tuned model does the re­view­ing in­stead of that fron­tier model

Platforms staff this work with teams of cat­a­log-in­tegrity an­a­lysts, which grow to­gether with the cat­a­log: more list­ings mean more cat­e­gories to know, more at­trib­utes to check, and more am­bigu­ous edge cases to judge. Inconsistent de­ci­sions prop­a­gate: prod­ucts be­come harder to find, rec­om­men­da­tions de­te­ri­o­rate, and pol­icy vi­o­la­tions slip through. Miss a coun­ter­feit and you ex­pose cus­tomers and brands to fraud; over-flag and you build an ex­pen­sive re­view queue that frus­trates le­git­i­mate sell­ers.

Let’s start with the scale. eBay alone car­ries about 2.5 billion live list­ings, and Shopify’s cat­a­log ab­sorbs more than 10 million prod­uct up­dates a day. Walmart has said that do­ing its AI-assisted cat­a­log work with peo­ple alone would have taken roughly 100× the head­count.

Getting it wrong costs rev­enue: 71% of shop­pers say they have re­turned a prod­uct be­cause it did­n’t match its list­ing. But the re­view­ing it­self is also ex­pen­sive. A mid-size mar­ket­place gen­er­ates on the or­der of 10 million list­ing cre­ates and ed­its a day; re­viewed with a fron­tier model, that work­load costs roughly half a bil­lion dol­lars a year, while a fine-tuned spe­cial­ist does the same job for around $10M:

One work­load, two price tags

fron­tier API per call fine-tuned spe­cial­ist, self-hosted

The com­pa­nies ac­tu­ally run­ning this work­flow con­firm the math. Shopify clas­si­fies prod­ucts with fine-tuned open mod­els at roughly 40 million in­fer­ences a day, a vol­ume it says com­mer­cial APIs can’t eco­nom­i­cally serve. Inspired by the in­dus­try lead­ers, we ran the play­book our­selves: an agent that ex­am­ines a list­ing, searches the prod­uct tax­on­omy, checks the brand, re­trieves the at­tribute schema, and com­mits a struc­tured de­ci­sion, es­ca­lat­ing to hu­man re­view when the ev­i­dence is thin or the risk too high.

One episode, end to end

ListingWork gloves, PU-coated, black, size 10

Claimed bran­dAma­zon­Ba­sics

RegionUS

1search_taxonomy → Safety Work Gloves

2lookup_brand → reg­is­tered, not pro­tected

3get_attribute_schema → brand, color, ma­te­r­ial, size

✓Commits cat­e­gory, at­trib­utes, and ver­dict al­lowed

Part IV

A dig­i­tal twin the model can prac­tice in

A model can only prac­tice a work­flow it can ac­tu­ally per­form, fail at, and retry. So we re­built the cat­a­log work­flow as a dig­i­tal twin of an e-com­merce plat­form: a sim­u­lated en­vi­ron­ment with the same list­ings, the same tools, and the same stakes as the real thing. The raw ma­te­r­ial is the Amazon Berkeley Objects dataset of real prod­uct im­ages and list­ing data, which we turned into 177,767 re­view episodes: one list­ing each, with im­age, ti­tle, de­scrip­tion, claimed brand, and re­gion. Into that stream we planted con­trolled pol­icy cases, mis­matched im­ages, con­flict­ing brand claims, and de­lib­er­ately le­git­i­mate claims as hard neg­a­tives, so every episode has a known cor­rect an­swer to score against.

Inside the twin, the model works the way an an­a­lyst would. It can search a tax­on­omy of roughly 13,000 cat­e­gories, check whether a brand is reg­is­tered and pro­tected, and re­trieve the re­quired at­trib­utes for a cho­sen cat­e­gory; then it must com­mit a cat­e­gory, the at­trib­utes, and a pol­icy de­ci­sion. A scorer grades every episode, re­ward­ing cor­rect out­puts and pe­nal­iz­ing missed vi­o­la­tions, un­sup­ported at­trib­utes, in­valid cat­e­gories, and wasted tool calls. The penal­ties en­code busi­ness pri­or­i­ties di­rectly: a missed vi­o­la­tion costs more than a false alarm.

The dig­i­tal twin, end to end

TOOLS

search_­tax­on­omy() lookup_brand() get_at­tribut­e_schema()

list­ing agent de­ci­sion score

re­ward

0.3·category + 0.3·attributes + 0.4·policy − tool_over­age

Benchmarking the fron­tier mod­els

Before any train­ing run, we mea­sured how far the fron­tier could get on its own. We bench­marked five fron­tier mod­els (GPT-5.5, GPT-5.6-sol, Gemini 3.1 Pro, Claude Opus 4.8, and Claude Fable 5) on 200 strat­i­fied val­i­da­tion episodes with iden­ti­cal tools, im­ages, scorer, and turn bud­get, both with a plain prompt and with op­ti­mized prompt in­struc­tions: 2,800 char­ac­ters of ex­trac­tion con­ven­tions, lookup pro­ce­dure, and worked ex­am­ples, tuned the way a prompt en­gi­neer would tune a pro­duc­tion sys­tem. The chart be­low shows both con­fig­u­ra­tions, with our fi­nal GRPO-trained model added for com­par­i­son.

The best fron­tier con­fig­u­ra­tion reached 76.9% of the achiev­able score; the trained 9B reached 87.3%. The gap is not in­tel­li­gence. A fron­tier model starts every episode from zero: it has never seen this store’s tax­on­omy, does­n’t know its in­ven­tory con­ven­tions, which at­tribute val­ues count as sup­ported, or how the plat­form wants cor­ner cases re­solved, and it has to re­con­struct all of that on the fly, from what­ever fits in the prompt, on every sin­gle call. Optimized in­struc­tions can com­press some of that knowl­edge, but the cor­ner cases that de­cide the score are ex­actly the ones no in­struc­tion prompt can enu­mer­ate.

Model bench­marks

Part V

Training de­tails

The train­ing setup was mod­est. We rented two RTX PRO 6000 GPUs, one gen­er­at­ing roll­outs and one ap­ply­ing gra­di­ent up­dates, with the open-source prime-rl frame­work run­ning the in­fra­struc­ture. The full run was 1,000 op­ti­mizer steps, took about three and a half days, and cost roughly $500 in GPU time.

Most of that was not needed to catch the fron­tier. The model crossed the fron­tier band af­ter roughly 250 steps, about a day of train­ing; the re­main­ing steps were spent squeez­ing out max­i­mum per­for­mance. The fi­nal model scores 0.626 on the bench­mark, 87.3% of the achiev­able ceil­ing and about ten points above the best fron­tier con­fig­u­ra­tion.

The gain comes from teach­ing the model the se­man­tics of its en­vi­ron­ment: over thou­sands of scored episodes it learns how the store’s tax­on­omy, tools, and poli­cies re­late to the re­ward, and con­verges on the op­ti­mal dis­tri­b­u­tion of ac­tions over the en­vi­ron­men­t’s ac­tion space. A gen­eral model spreads its ca­pac­ity across every­thing; the spe­cial­ist spends all of it on the spe­cific task at the ex­pense of gen­eral abil­ity.

The re­sult also is­n’t a dead end. When a more ca­pa­ble open-source base model ships, the recipe trans­fers: as a rule of thumb, a stronger base yields a stronger post-trained spe­cial­ist. And the de­ployed model gen­er­ates its own train­ing data as it works: logged de­ci­sions can be dis­tilled into su­per­vised fine-tun­ing sets, mak­ing each re-train­ing cheaper and bet­ter-in­formed than the last.

Training run

So far, we have dis­cussed how re­in­force­ment learn­ing helps a lan­guage model nav­i­gate your com­pany en­vi­ron­ment more ac­cu­rately. But one of the most im­por­tant as­pects of this par­a­digm is that the ac­cu­racy also comes cheaper than run­ning fron­tier mod­els. The cost-vs-qual­ity chart that opens this ar­ti­cle puts both on one pic­ture, task score against cost per thou­sand list­ings, and the take­away for a bud­get owner is sim­ple: with every­thing else on the chart you are choos­ing be­tween qual­ity and cost. The trained spe­cial­ist ends that trade-off, de­liv­er­ing the best score we mea­sured at close to the low­est price.

Fine-tuning adds ~23 points over the base 9B at the same ~$0.50 per 1,000 list­ings: 40× cheaper than the least ex­pen­sive fron­tier con­fig­u­ra­tion (Gemini, $19/1k) and ~340× cheaper than the most ex­pen­sive (GPT-5.5-pro, $172/1k). The 2,800 char­ac­ters of prompt in­struc­tions also raised GPT-5.5′s mea­sured cost by a third, a prompt tax paid on every fu­ture call. The spe­cial­ist’s in­struc­tions live in its weights. At Shopify-scale vol­ume of roughly 40 mil­lion de­ci­sions a day, the gap be­tween $34 and $0.50 per thou­sand is about $500M a year ver­sus $7M.

Wondering where these curves sit for your work­flow? Book a 30-minute au­dit →

Part VI

What can own­ing AI do for you?

Here is the whole ex­per­i­ment in three sen­tences. We took one high-vol­ume e-com­merce work­flow, re­view­ing prod­uct list­ings, and re­built it as a dig­i­tal twin. We let a small open-source model prac­tice in­side it for a few days and roughly $500 of GPU time. The trained spe­cial­ist ended above every fron­tier con­fig­u­ra­tion we mea­sured, at a frac­tion of their price per de­ci­sion.

The recipe trans­fers to any work your busi­ness re­peats at vol­ume. Look for the places where peo­ple or API calls turn in­for­ma­tion into de­ci­sions all day: rout­ing tick­ets, ex­tract­ing fields from doc­u­ments, check­ing sub­mis­sions against pol­icy, clas­si­fy­ing prod­ucts, ap­prov­ing or flag­ging trans­ac­tions. What unites them is that every de­ci­sion can be checked: there is a rule, a schema, or an ex­pert who can say whether it was right. That check is the whole trick. If a de­ci­sion can be scored, a model can prac­tice it; if it can only be de­bated, it can­not.

It is just as im­por­tant to know where this is the wrong tool. Two ques­tions set­tle most cases: how fre­quently the task runs, and whether its out­come is ver­i­fi­able.

Right tool for the right-shaped prob­lem

prompt-op­ti­mized fron­tier model fine-tuned spe­cial­ist owned, prac­ticed, cheap at scale fron­tier model fron­tier model + hu­man re­view

ver­i­fi­able non-ver­i­fi­able

rare task fre­quency fre­quent

Your work­flow is a fine-tun­ing can­di­date if at least one point ap­plies…

It hap­pens at high vol­ume, of­ten enough that per-de­ci­sion cost and er­rors com­pound into real money

Every out­come can be checked by a rule, test, or rubric, with no per­son in the loop

You Could Have Come Up With Kimi Delta Attention | Doubleword

blog.doubleword.ai

A note on no­ta­tion: this ar­ti­cle de­faults to bra-ket no­ta­tion be­cause (in my quan­tum-in­spired opin­ion) it makes the shapes in this de­riva­tion very clear. The Math no­ta­tion switch above rewrites every equa­tion us­ing con­ven­tional bold vec­tors and ex­plicit trans­poses in­stead. In bra-ket mode, ∣q⟩\lvert q\ran­gle is a col­umn vec­tor, ⟨k∣\langle k\rvert is a row vec­tor, ⟨k∣q⟩\langle k\rvert q\ran­gle is a num­ber, and ∣v⟩⟨k∣\lvert v\ran­gle\lan­gle k\rvert is a ma­trix. Vectors face right by de­fault, while keys face left when writ­ten into the lin­ear-at­ten­tion state. We work with one causal at­ten­tion head and real-val­ued vec­tors, as­sume DeltaNet’s keys are nor­mal­ized, and let the state map from key space to value space.

Modern lin­ear at­ten­tion vari­ants are com­plex, and a upon first glance it is not so easy to see what they are de­signed to achieve. For ref­er­ence here is the state up­date equa­tion for Kimi Delta Attention (KDA):

S~t=St−1Diag⁡(αt)\widetilde S_t = S_{t-1}\operatorname{Diag}(\alpha_t) ∣v^t⟩=S~t∣kt⟩\lvert\widehat v_t\ran­gle = \widetilde S_t\lvert k_t\ran­gle ∣et⟩=βt(∣vt⟩−∣v^t⟩)\lvert e_t\ran­gle = \beta_t \left( \lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle \right) St=S~t+∣et⟩⟨kt∣S_t = \widetilde S_t+\lvert e_t\ran­gle\lan­gle k_t\rvert ∣ot⟩=St(dk−1/2∣qt⟩)\lvert o_t\ran­gle = S_t\left(d_k^{-1/2}\lvert q_t\ran­gle\right)

The rea­son they are so dif­fi­cult to un­der­stand is that this is the lat­est in a fam­ily of lin­ear at­ten­tion vari­ants that have been de­vel­oped over the last few years and the com­plex­ity of them has in­evitably bal­looned such that from the out­side the lat­est vari­ants ap­pear in­ac­ces­si­ble.

In this post we are go­ing to walk through the DeltaNet fam­ily of lin­ear at­ten­tion vari­ants, two of which are used by the lat­est Qwen and Kimi model fam­i­lies, and show how you might have ar­rived at the same equa­tions by as­sert­ing sim­ple things about your hid­den state.

That is the route we will take:

soft­max at­ten­tion → lin­ear at­ten­tion → DeltaNet → Gated DeltaNet → KDA

Only af­ter de­riv­ing KDA will we turn to the re­cur­rent and chunk­wise Triton pro­grams that ex­e­cute it.

1. Begin with qua­dratic at­ten­tion

For a query at to­ken tt, or­di­nary causal soft­max at­ten­tion is

ati=exp⁡ ⁣(s⟨ki∣qt⟩)∑j≤texp⁡ ⁣(s⟨kj∣qt⟩),s=dk−1/​2,∣ot⟩=∑i≤tati∣vi⟩.\be­gin{aligned} a_{ti} &= \frac{ \exp\!\left(s\langle k_i\rvert q_t\ran­gle\right) }{ \sum_{j\leq t} \exp\!\left(s\langle k_j\rvert q_t\ran­gle\right) }, \qquad s=d_k^{-1/​2},\\ \lvert o_t\ran­gle &= \sum_{i\leq t}a_{ti}\lvert v_i\ran­gle. \end{aligned}

Every at­ten­tion weight is a scalar. It mea­sures the sim­i­lar­ity be­tween one key and one query, then soft­max turns all of the scores for that query into a dis­tri­b­u­tion. The out­put is a weighted sum of value vec­tors.

Over a se­quence of length TT, there are T2T^2 key-query pairs. During au­tore­gres­sive in­fer­ence we can cache the keys and val­ues in­stead of re­com­put­ing them, but the cache still grows with the se­quence and every new query still has to in­spect the en­tire his­tory.

The ob­sta­cle to re­ar­rang­ing this com­pu­ta­tion is the soft­max. Its de­nom­i­na­tor de­pends jointly on the cur­rent query and every ear­lier key. So, for the mo­ment, re­move it.

1.1 Remove the soft­max

For clar­ity, ab­sorb the con­stant scale ss into the query. The de­lib­er­ately bare ver­sion of at­ten­tion is then

∣ot⟩=∑i≤t⟨ki∣qt⟩∣vi⟩.\lvert o_t\ran­gle = \sum_{i\leq t} \langle k_i\rvert q_t\ran­gle \lvert v_i\ran­gle.

The scalar in­ner prod­uct can move to the right:

∣ot⟩=∑i≤t∣vi⟩⟨ki∣qt⟩=(∑i≤t∣vi⟩⟨ki∣)∣qt⟩.\begin{aligned} \lvert o_t\ran­gle &= \sum_{i\leq t} \lvert v_i\ran­gle \langle k_i\rvert q_t\ran­gle\\ &= \left( \sum_{i\leq t} \lvert v_i\ran­gle\lan­gle k_i\rvert \right) \lvert q_t\ran­gle. \end{aligned}

Everything that de­pends on the past can now be col­lected into one ma­trix of a fixed size V×KV \times K:

St=∑i≤t∣vi⟩⟨ki∣\boxed{ S_t = \sum_{i\leq t} \lvert v_i\ran­gle\lan­gle k_i\rvert }

and at­ten­tion be­comes a re­cur­rent write fol­lowed by a read:

St=St−1+∣vt⟩⟨kt∣,∣ot⟩=St∣qt⟩.\boxed{ \begin{aligned} S_t &= S_{t-1} + \lvert v_t\ran­gle\lan­gle k_t\rvert,\\ \lvert o_t\ran­gle &= S_t\lvert q_t\ran­gle. \end{aligned} }

The iden­tity

(∣v⟩⟨k∣)∣q⟩=⟨k∣q⟩∣v⟩\left(\lvert v\ran­gle\lan­gle k\rvert\right)\lvert q\ran­gle = \langle k\rvert q\ran­gle\lvert v\ran­gle

is the whole trick. The outer prod­uct is a ma­trix; the in­ner prod­uct is a num­ber. We no longer store every past key and value. We store their summed outer prod­ucts in the fixed-size state StS_t.

This is lin­ear in se­quence length rather than qua­dratic: scan the to­kens once, up­dat­ing the same dv×dkd_v\times d_k state at every step. We have paid for that ef­fi­ciency by dis­card­ing soft­max’s nor­mal­iza­tion and se­lec­tiv­ity. More so­phis­ti­cated lin­ear-at­ten­tion meth­ods use fea­ture maps and nor­mal­iz­ers, but this un­adorned form ex­poses the mem­ory prob­lem that mo­ti­vates DeltaNet.

1.2 Addition is not as­sign­ment

Suppose we write a pair ∣vt⟩⟨kt∣\lvert v_t\ran­gle\lan­gle k_t\rvert and im­me­di­ately query the new state with that same key:

St∣kt⟩=(St−1+∣vt⟩⟨kt∣)∣kt⟩=St−1∣kt⟩+∣vt⟩⟨kt∣kt⟩⏟1=St−1∣kt⟩+∣vt⟩.\begin{aligned} S_t\lvert k_t\ran­gle &= \left( S_{t-1} + \lvert v_t\ran­gle\lan­gle k_t\rvert \right) \lvert k_t\ran­gle\\ &= S_{t-1}\lvert k_t\ran­gle + \lvert v_t\ran­gle \underbrace{\langle k_t\rvert k_t\ran­gle}_{1}\\ &= S_{t-1}\lvert k_t\ran­gle+\lvert v_t\ran­gle. \end{aligned}

The write does not make the mem­ory re­turn ∣vt⟩\lvert v_t\ran­gle. It adds ∣vt⟩\lvert v_t\ran­gle to what­ever the mem­ory al­ready re­turned.

If the old state al­ready pro­duced the cor­rect value, the ad­di­tive write makes the new state pro­duce twice that value. More gen­er­ally, keys are not mu­tu­ally or­thog­o­nal, so every write can in­ter­fere with pre­vi­ous writes. Linear at­ten­tion has given us a com­pact as­so­cia­tive mem­ory, but its up­date be­haves like += when what we want is closer to =.

2. DeltaNet: write the er­ror, not the value

DeltaNet re­places the un­con­di­tional lin­ear-at­ten­tion write with a delta-rule cor­rec­tion. There are two use­ful ways to de­rive it.

2.1 Derivation one: de­mand that the write can be read back

Before writ­ing to­ken tt, ask the mem­ory what it cur­rently as­so­ci­ates with the new key:

∣v^t⟩=St−1∣kt⟩.\lvert\widehat v_t\ran­gle = S_{t-1}\lvert k_t\ran­gle.

If we want the mem­ory to re­turn ∣vt⟩\lvert v_t\ran­gle, we should not add the whole value. We should add only the dif­fer­ence:

∣vt⟩−∣v^t⟩.\lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle.

Introduce a learned write strength βt∈[0,1]\be­ta_t\in[0,1] and de­fine

∣et⟩=βt(∣vt⟩−St−1∣kt⟩).\lvert e_t\ran­gle = \beta_t \left( \lvert v_t\ran­gle - S_{t-1}\lvert k_t\ran­gle \right).

Then write this er­ror at the cur­rent key:

St=St−1+∣et⟩⟨kt∣.\boxed{ S_t = S_{t-1} + \lvert e_t\ran­gle\lan­gle k_t\rvert. }

Now im­me­di­ately read the same key:

St∣kt⟩=St−1∣kt⟩+∣et⟩⟨kt∣kt⟩=(1−βt)St−1∣kt⟩+βt∣vt⟩.\begin{aligned} S_t\lvert k_t\ran­gle &= S_{t-1}\lvert k_t\ran­gle + \lvert e_t\ran­gle \langle k_t\rvert k_t\ran­gle\\ &= (1-\beta_t)S_{t-1}\lvert k_t\ran­gle + \beta_t\lvert v_t\ran­gle. \end{aligned}

When βt=1\be­ta_t=1, the re­sult is ex­actly ∣vt⟩\lvert v_t\ran­gle. Smaller βt\be­ta_t moves the old pre­dic­tion part­way to­wards the tar­get.

The cor­rec­tion is also lo­cal in key space. For any query ∣x⟩\lvert x\ran­gle or­thog­o­nal to the cur­rent key,

⟨kt∣x⟩=0⟹(St−St−1)∣x⟩=∣et⟩⟨kt∣x⟩⏟0=0.\langle k_t\rvert x\ran­gle=0 \quad\Longrightarrow\quad (S_t-S_{t-1})\lvert x\ran­gle = \lvert e_t\ran­gle \underbrace{\langle k_t\rvert x\ran­gle}_{0} =0.

So the rank-one write changes the re­sponse in the se­lected key di­rec­tion while leav­ing every or­thog­o­nal di­rec­tion alone.

2.2 Derivation two: take one step on re­con­struc­tion loss

The same up­date falls out of an on­line learn­ing ob­jec­tive. Treat the cur­rent key-value pair as one train­ing ex­am­ple for the lin­ear map SS:

Lt(S)=12∥S∣kt⟩−∣vt⟩∥22.\mathcal L_t(S) = \frac12 \left\| S\lvert k_t\ran­gle-\lvert v_t\ran­gle \right\|_2^2.

Its gra­di­ent with re­spect to the state is

∇SLt(S)=(S∣kt⟩−∣vt⟩)⟨kt∣.\nabla_S\mathcal L_t(S) = \left( S\lvert k_t\ran­gle-\lvert v_t\ran­gle \right) \langle k_t\rvert.

This is vis­i­bly an outer prod­uct: a value-space pre­dic­tion er­ror times the key bra at which that er­ror was ob­served. Take one gra­di­ent-de­scent step of size βt\be­ta_t from St−1S_{t-1}:

St=St−1−βt∇SLt(St−1)=St−1−βt(St−1∣kt⟩−∣vt⟩)⟨kt∣=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣.\begin{aligned} S_t &= S_{t-1} - \beta_t\nabla_S\mathcal L_t(S_{t-1})\\ &= S_{t-1} - \beta_t \left( S_{t-1}\lvert k_t\ran­gle-\lvert v_t\ran­gle \right) \langle k_t\rvert\\ &= S_{t-1} + \beta_t \left( \lvert v_t\ran­gle-S_{t-1}\lvert k_t\ran­gle \right) \langle k_t\rvert. \end{aligned}

This is ex­actly the up­date we got by re­quir­ing im­me­di­ate re­con­struc­tion. The two in­ter­pre­ta­tions are the same:

as a mem­ory op­er­a­tion, βt\be­ta_t con­trols how strongly to re­place the old as­so­ci­a­tion;

as on­line learn­ing, βt\be­ta_t is the step size;

as lin­ear al­ge­bra, the change is a rank-one outer prod­uct.

2.3 The DeltaNet state tran­si­tion

Expanding the er­ror ex­poses DeltaNet as a struc­tured state tran­si­tion plus a new in­put:

St=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣=St−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1} + \beta_t \left( \lvert v_t\ran­gle-S_{t-1}\lvert k_t\ran­gle \right) \langle k_t\rvert\\ &= S_{t-1} \left( I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert \right) + \beta_t\lvert v_t\ran­gle\lan­gle k_t\rvert. \end{aligned}

For a unit key, I−βt∣kt⟩⟨kt∣I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert has eigen­value 1−βt1-\beta_t in the cur­rent key di­rec­tion and eigen­value 11 in every or­thog­o­nal di­rec­tion. It re­moves the old as­so­ci­a­tion along the cur­rent key be­fore adding the new one.

DeltaNet fixes the write. It does not yet fix the life­time of the state.

3. Gated DeltaNet: some­times old in­for­ma­tion should dis­ap­pear

The lin­ear state com­presses the whole his­tory into one ma­trix. A read

St∣q⟩=∑i≤t⟨ki∣q⟩∣vi⟩S_t\lvert q\ran­gle = \sum_{i\leq t} \langle k_i\rvert q\ran­gle\lvert v_i\ran­gle

can­not choose to skip an in­di­vid­ual old to­ken af­ter that to­ken has been folded into StS_t. Every stored di­rec­tion that over­laps the query con­tributes. The delta rule can cor­rect the state around the cur­rent key, but stale in­for­ma­tion in other di­rec­tions re­mains avail­able and can dis­tort fu­ture reads.

We there­fore need a way to for­get the old state be­fore us­ing it. Let αt∈[0,1]\al­pha_t\in[0,1] be a learned scalar re­ten­tion gate:

S~t=αtSt−1.\widetilde S_t = \alpha_t S_{t-1}.

Run the same delta rule against this gated state:

S~t=αtSt−1,forget,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write.\boxed{ \begin{aligned} \widetilde S_t &= \alpha_tS_{t-1}, &&\text{forget},\\ \lvert\widehat v_t\ran­gle &= \widetilde S_t\lvert k_t\ran­gle, &&\text{predict},\\ \lvert e_t\ran­gle &= \beta_t \left( \lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\ran­gle\lan­gle k_t\rvert, &&\text{write}. \end{aligned} }

This is Gated DeltaNet. The or­der mat­ters: for­get first, pre­dict from the re­tained state, then cor­rect that pre­dic­tion. If we pre­dicted be­fore for­get­ting, the er­ror would de­scribe a dif­fer­ent mem­ory from the one we up­date.

Expanding the re­cur­rence gives

St=αtSt−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.S_t = \alpha_tS_{t-1} \left( I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert \right) + \beta_t\lvert v_t\ran­gle\lan­gle k_t\rvert.

The delta rule gives tar­geted re­place­ment; the scalar gate gives global era­sure. They solve dif­fer­ent prob­lems and are com­ple­men­tary.

But αt\al­pha_t still makes one de­ci­sion for the en­tire ma­trix. The model must re­tain or for­get every key chan­nel at the same rate.

4. Kimi Delta Attention: for­get each chan­nel in­de­pen­dently

Kimi Delta Attention re­places Gated DeltaNet’s scalar re­ten­tion with a vec­tor αt∈[0,1]dk\al­pha_t\in[0,1]^{d_k}. Put the vec­tor on the di­ag­o­nal:

Dt=Diag⁡(αt)∈Rdk×dk.D_t = \operatorname{Diag}(\alpha_t) \in\mathbb R^{d_k\times d_k}.

Our state maps keys to val­ues, so the key chan­nels are the columns of SS. Right-multiplication ap­plies a dif­fer­ent re­ten­tion fac­tor to every one:

S~t=St−1Dt.\widetilde S_t = S_{t-1}D_t.

Everything else is the delta rule we have al­ready de­rived:

S~t=St−1Dt,forget each key channel,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write,∣ot⟩=St(s∣qt⟩),s=dk−1/2,read.\boxed{ \begin{aligned} \widetilde S_t &= S_{t-1}D_t, &&\text{forget each key chan­nel},\\ \lvert\widehat v_t\ran­gle &= \widetilde S_t\lvert k_t\ran­gle, &&\text{predict},\\ \lvert e_t\ran­gle &= \beta_t \left( \lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\ran­gle\lan­gle k_t\rvert, &&\text{write},\\ \lvert o_t\ran­gle &= S_t(s\lvert q_t\ran­gle), \qquad s=d_k^{-1/​2}, &&\text{read}. \end{aligned} }

That is KDA. Compared with Gated DeltaNet, the con­cep­tual change is only the pro­mo­tion

αt⟶Dt=Diag⁡(αt).\al­pha_t \quad\longrightarrow\quad D_t=\operatorname{Diag}(\alpha_t).

The ef­fect is sub­stan­tial: one chan­nel can be cleared while an­other is re­tained.

4.1 Why the tran­si­tion is di­ag­o­nal-plus-low-rank

Expand the KDA cor­rec­tion:

St=St−1Dt+βt(∣vt⟩−St−1Dt∣kt⟩)⟨kt∣=St−1Dt(I−βt∣kt⟩⟨kt∣)⏟At+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1}D_t + \beta_t \left( \lvert v_t\ran­gle - S_{t-1}D_t\lvert k_t\ran­gle \right) \langle k_t\rvert\\ &= S_{t-1} \underbrace{ D_t \left( I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert \right) }_{A_t} + \beta_t\lvert v_t\ran­gle\lan­gle k_t\rvert. \end{aligned}

The key-space tran­si­tion is

At=Dt−βtDt∣kt⟩⟨kt∣=Dt−∣bt⟩⟨at∣,\begin{aligned} A_t &= D_t-\beta_tD_t\lvert k_t\ran­gle\lan­gle k_t\rvert\\ &= D_t-\lvert b_t\ran­gle\lan­gle a_t\rvert, \end{aligned}

where

∣bt⟩=Dt∣kt⟩,⟨at∣=βt⟨kt∣.\lvert b_t\ran­gle=D_t\lvert k_t\ran­gle, \qquad \langle a_t\rvert=\be­ta_t\lan­gle k_t\rvert.

So AtA_t is a di­ag­o­nal ma­trix mi­nus a rank-one ma­trix: a di­ag­o­nal-plus-low-rank, or DPLR, tran­si­tion. DPLR de­scribes the dk×dkd_k\times d_k tran­si­tion act­ing on key space. The mem­ory state it­self is still the dv×dkd_v\times d_k ma­trix StS_t.

The full jour­ney can now be sum­ma­rized com­pactly:

The im­ple­men­ta­tion usu­ally stores gt=log⁡αt­g_t=\log\al­pha_t with gt≤0g_t\leq0, then ob­tains the re­ten­tion fac­tors as exp⁡(gt)\exp(g_t). In the trans­posed dk×dvd_k\times d_v lay­out used by the ref­er­ence code, the re­cur­rence is only five lines:

state = state * g_t.exp().un­squeeze(-1) pre­dic­tion = ein­sum(“bhkv,bhk->bhv”, state, k_t) resid­ual = be­ta_t.un­squeeze(-1) * (v_t - pre­dic­tion) state = state + ein­sum(“bhk,bhv->bhkv”, k_t, resid­ual) out­put = ein­sum(“bhk,bhkv->bhv”, q_t * scale, state)

See the of­fi­cial naive_re­cur­ren­t_kda ref­er­ence.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.