topへ
10 interesting stories served every morning and every evening.
10 interesting stories served every morning and every evening.
topへ
Highlights:
Scientists at La Jolla Institute for Immunology, Scripps Research have developed an HIV vaccine that trains immune cells to see past HIV’s defenses.
This HIV vaccine works by prompting the body’s immune system to make substantial numbers of rarely seen “broadly neutralizing” antibodies.
In this new study, this vaccine resulted in the best HIV-fighting antibody response ever seen in primates. Human trials have now started.
LA JOLLA, CA—A new HIV vaccine developed by La Jolla Institute for Immunology (LJI), Scripps Research scientists, and IAVI has the potential to protect humans from developing HIV infection and AIDS. This HIV vaccine is the first to generate a high number of “broadly neutralizing,” virus-fighting antibodies in primates.
“This feels like a huge success,” says LJI Professor and Chief Scientific Officer Shane Crotty, Ph.D., who co-led the research with Scripps Research Professor William Schief, Ph.D. “We constructed a successful vaccine from the ground up, which required a deep understanding of the immune system.”
This groundbreaking research, published in Nature, is the result of 14 years of collaboration between La Jolla Institute for Immunology and Scripps Research, as part of the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD). “This has been one of those Apollo moon mission-type projects, where there is an exceptional goal and the team has to accomplish a myriad of discoveries and inventions along the way,” says Crotty.
Outsmarting HIV
The new vaccine works by intervening in a process called B cell maturation. B cells make antibodies. Like many immune cells, B cells have an early “naive” stage before they are ready to make antibodies. B cells start to mature once they get the signal that a pathogen, such as a virus, is trying to attack. B cells see pieces of that pathogen’s molecular structure and start producing antibodies that can bind to that structure and halt infection.
It can take a little while for B cells to find the right “bullseye” on a pathogen. But B cells keep trying. As they mature, B cells tweak their antibody production, refining antibody structures to bind to a pathogen in just the right, vulnerable spots.
Scientists describe B cell development as a training process or bootcamp. In most cases, the body is left with a well-honed B cell army.
HIV is hard to beat because it doesn’t give B cells a chance to develop effective antibodies. The first problem is that HIV disguises itself from the immune system. The virus is wrapped in an ever-shifting cloak of sugar molecules, called glycans. This lets HIV sneak undetected past human cells, which are also covered in glycans.
The second big problem is that HIV mutates very quickly. “The worldwide diversity of HIV mutations is extraordinary. Even the diversity within one individual person living with HIV is dramatic,” says LJI Instructor Patrick Madden, Ph.D., who served as study co-first author with Jon Steichen, Ph.D., an institute investigator at Scripps Research.
The third problem is that HIV changes its shape when it infects human cells. Even if B cells get a glimpse of its viral structure—snap!—the structure changes.
Taken together, these problems rarely give B cells a chance to hone their antibody responses against HIV. Even if a B cell manages to make neutralizing antibodies, the virus can mutate or change its shape, rendering those antibodies useless.
The LJI and Scripps Research teams spent years hunting for “broadly neutralizing” antibodies that can actually bind to HIV and recognize key viral structures, even if the rest of the virus mutates. These antibodies are very, very rare, but they can be found in blood samples from a small number of people living with HIV.
An effective HIV vaccine would need to prompt the immune system to make these same broadly neutralizing antibodies. “How could we flip the whole immune response on its head so the rare responses become the common responses? That was a critical challenge we faced,” says Crotty.
Testing the new vaccine
It was time to go back to B cell bootcamp. The scientists studied what made the HIV-fighting B cells special. Then they reversed the process to see exactly how those B cells matured. By looking back at the maturation process, the researchers could track how the B cells changed when they saw specific pieces of the HIV structure.
The team discovered that B cells matured to make broadly neutralizing antibodies after they got an early look at parts of HIV’s outer “envelope” protein. Because these viral sites sparked an immune response, scientists would call them “antigens.”
An effective HIV vaccine would likely need to include models of these antigens. The antigens would work like mugshots of America’s most wanted. If B cells saw those antigens early and often, they would get really good at recognizing and even neutralizing HIV. “We were trying to mimic the progression of those neutralizing antibodies,” says Madden.
In a feat of molecular engineering, the Schief Lab developed vaccine molecules that resembled the real HIV antigens. The scientists then worked with Emory National Primate Research Center, to test this potential HIV vaccine in a non-human primate species called rhesus macaques.
The researchers first administered a “priming” vaccine meant to activate each animal’s naive B cells. The animals then received a series of “shepherding” booster shots to help their B cells develop along the right path.
“This series of vaccinations will guide, or ‘walk’, a B cell from its naive state to its broadly neutralizing state,” says Madden.
This new type of vaccine approach is called “germline targeting” because it targets naive B cells in their “germline” or naive form, before they begin their training process.
The scientists found that around 44 percent of the animals went on to produce broadly neutralizing antibodies against HIV in their blood. These antibodies were impressively abundant.
“We succeeded in taking ultra-rare antibody responses and turning them into common responses by the end of the vaccination process,” adds Crotty. In other research recently published, they reported a new strategy to accelerate related vaccine antibody responses [See Nature Immunology paper].
The team didn’t test whether these antibodies could prevent infection, but it’s significant that these antibodies could be found in the blood, where they could encounter and potentially block HIV.
Bringing the HIV vaccine to humans
The Crotty Lab plans to investigate how they might change the booster shot regimen to make the HIV vaccine even more effective. “It was incredible to get those results, but of course we’d like to see a response in 100 percent of the animals,” says Madden.
Importantly, the antibodies found in the animal subjects resembled the exact kinds of broadly neutralizing antibodies seen in those rare humans who made their own neutralizing antibodies. It’s clear that our immune systems can make these powerful antibodies, given the right training.
“We believe this vaccine approach is even more likely to succeed in humans, because of the immunogenetics,” Crotty says.
The priming immunogen used in this study was evaluated in humans in the HVTN 144 trial and is currently being tested in the Phase 1 trial IAVI G004. IAVI, Scripps Research, the HIV Vaccine Trials Network, and partners are now advancing plans to further evaluate the full immunization regimen in a future human clinical study.
Additional authors of the study, “Vaccination elicits HIV broadly neutralizing antibodies in primates,” include Claudia T. Flynn, Swastik Phulera, Monolina Shil, Oleksandr Kalyuzhniy, Alessia Liguori, Carolyne Kifude, Leigh M. Sewall, Christopher A. Cottrell, Krystal M. Ma, Sabyasachi Baboo, Jolene K. Diedrich, Katherine McKenney, Allan C. deCamp, Diane G. Carnathan, Ivy Phung, Parham Ramezani-Rad, Ester Marina-Zárate, Brian Freeman, Zhenfei Xie, Jeong Hyun Lee, Troy Sincomb, Nicole Phelps, Danny Lu, Diana Goodwin, Ryan Tingle, Yumiko Adachi, Nushin Alavi, Jenny Tran, Andy S. Tran, Alyne Nascimento, Catherine Sovie, Daniel L. V. Bader, Hannah Voic, Xiaoya Zhou, Grace Pixton, Agnes Walsh, Mariane B. Melo, Torben Schiffner, Facundo D. Batista, Dennis R. Burton, Darrell J. Irvine, James C. Paulson, John R. Yates III, Gabriel Ozorowski, Andrew B. Ward, Guido Silvestri.
This work was supported by National Institute of Allergy and Infectious Diseases (NIAID), of the National Institutes of Health, through grant UM1 Al100663 to the Scripps Center for HIV/AIDS Vaccine Immunology and Immunogen Discovery (CHAVI-ID), grant UM1 AI144462 to the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD), P51 OD011132 to Emory National Primate Research Center, and R01 AI113867; by the Gates Foundation under the Collaboration for AIDS Vaccine Discovery (NAC INV-007522, INV-008813, INV-034657, and INV-064772), via the IAVI Neutralizing Antibody Center (NAC); and by the National Institute of Health grant S10OD025052.
European Citizens’ Initiative
The story of MIT’s most notorious milk carton begins, as many good stories do, in a college dorm.
The milk in question was purchased in 1994 and rediscovered in 1995 by an undergrad named Justin Cave. By that point, the reportedly lactose-intolerant Cave had even less use for the milk he had abandoned in his fridge ten months earlier. For reasons lost to history, he did not throw the milk away. He threw it a birthday party.
The Milk lived the rest of its life unrefrigerated, stored in a tall, single-walled jar. For twenty-seven years, the residents of Random Hall dorm gathered faithfully to celebrate its birthday. At age 20, the Milk applied to, and was rejected from, MIT1. The jar was periodically “burped” to release the gas pressure inside, until the Milk reached its stable final form — a cloudy brown liquid. When asked why the Milk was never thrown away, one resident of Random Hall replied: “Why throw something away when you can tell a story about it?”
“Stuff I learned from things that normally get thrown away” could be the title of many scientists’ memoirs, including Louis Pasteur’s. Winemaking produces, well, wine, but it also produces acidic crystals on the walls of the vats. These byproducts were not discarded — they were studied by Pasteur and his contemporaries. Pasteur’s observations both revolutionized our understanding of chemistry and led him to the phenomenon that would define his career and lay the foundation of modern food safety: fermentation.
The microorganisms responsible for fermentation are visible to our naked senses only through the textures, colors and smells resulting from their collective efforts. From Aristotle through the 1850s, it was assumed that some intrinsic property of a non-living starting substance (like grain or milk) enabled its spontaneous fermentation into something useful (like beer or yogurt) or its eventual spoilage.
It was Pasteur who proved that living organisms were required for the transformations that took place during fermentation. He heated up nutrient-rich broths in custom flasks that let gases, but not microbes, flow in and out of the flasks. Pasteur then broke the neck off of one of the flasks to expose the broth to the air. If the boiled broth could spontaneously transform, it would do so with or without exposure to microbes in the air and environment.
The flask with the neck broken off grew cloudy and fermented as bacteria bloomed, but the sterile one remained clear.
This finding was great news for Napoleon. The French were losing money, and perhaps more alarmingly, their reputation, exporting wine to the British — the wine would “spontaneously” go bad during shipping. The French government offered a prize for a scientist to solve the case of the spoiled wine. With the knowledge of the microorganism-driven process of fermentation in hand, Pasteur did to the wine what he did to the broth, just more gently — he heated the wine enough to kill microbes without damaging the wine’s flavor. Immortalized as pasteurization, this process was adapted shortly after its invention in 1865 to let us safely drink stored milk2.
Killing bacteria thus became a major preoccupation of modern life. Its most visible manifestation today might be taking antibiotics (first available in the 1940s): there were ~700 antibiotic prescriptions per 1000 people3 in 2024 according to CDC data. A close second might be the dizzying array of disinfectant products found in U.S. grocery stores.
What’s less visible is the sanitization infrastructure that makes things like grocery stores or medicine possible at all. The company Steris, one maker of high temperature, pressurized sterilization equipment and other medical instruments, is a $5 billion annual revenue company, with a $21 billion market cap. The U.S. pasteurizes around 50 billion liters of fluid milk every year. To package salad greens like spinach, the greens are washed in a dilute bleach solution to kill any lingering soil microbes. I could go on.
But our war on bacteria has its own warring industry. This industry has captured the imaginations of scientists, the food and beverage industry, pharma companies and doctors along with influencers, marketing gurus and opportunists of all flavors. This industry emphasizes that some microbes are friends, not foe, (true) and you should be eating them in large quantities, on purpose, all the time, and preferably paying more for products that contain them (dubious). This is the probiotics industry.
“Probiotic” is a bit of a misnomer — it means for life, or promoting life, but the formal definition of a probiotic is an actual living microorganism. In simple terms: taking a probiotic is just eating bacteria on purpose. I say on purpose because we consume microbes accidentally all the time from our environment, largely oblivious to their existence or effects. The bacteria we spend much of our time and energy trying to kill are outnumbered, at a species level, at least 1000 to 1 by a combination of harmless and beneficial bacteria living in and on our bodies. It’s this latter property of beneficialness that probiotics are trying to exploit.
I say exploit because of a recent trip I took to the grocery store. I had a cold and was in search of lemon ginger tea. I bought a box of Bigelow, went home, boiled some water, poured it over a tea bag, waited a bit, added honey, took a sip, and almost spit it out. The tea had its expected notes of ginger, a hint of lemon, and some powdery, alkaline aftertaste that I couldn’t place. Frankly, it tasted terrible. (Sorry, Bigelow).
I inspected the box again. In my congested state, I had unwittingly purchased a new offering from the tea company — Bigelow Lemon Ginger, with probiotics. What made this tea different from all the other teas I happily sipped on was that in addition to nice-sounding things like lemongrass and cinnamon, it contained bacteria. Bacteria which I had just boiled, at a temperature 40oC hotter than pasteurization.
Did the tea taste bad because I was drinking dead bacteria water? And if that was the ultimate outcome of the normal brewing process, why bother putting bacteria in the tea at all?
I was at a lab happy hour when I mentioned this to my PhD thesis advisor. “I know, right?” she said, suddenly animated. “Probiotic teas taste SO BAD.” I was thrilled to have another witness. “Doesn’t it seem crazy to add in bacteria that you’re just going to boil and kill anyway?” I asked. “Is it all a scam?” She was already nodding. “You have to wonder whether the bacteria in the tea make it to the gut at all, and whether they do anything helpful once they get there,” she said.
I’d be lying if I said I remembered exactly what happened next, or who suggested what. All I remember is an idea. An idea to test this seemingly paradoxical marketing tactic like the microbiologists we are. The idea was simple: What if we tried to grow the bacteria from the tea bag, in the lab?
I went home. I stared at the box of bacteria tea.
Why throw something away when you can tell a story about it?
BC30™, the bacterial strain in the tea, is short for Bacillus coagulans GBI-30, 6086®. It received the FDA’s GRAS (Generally Recognized as Safe4) designation in 2012 and is found in over a thousand “leading food, beverage and pet food products worldwide” according to the probiotic’s website.
To coax these bacteria to grow out of steeped tea, I needed to know three things:
Is this species safe to grow in the lab?
Is this species safe to grow in the lab?
What does it like to eat?
What does it like to eat?
What are its preferred growth conditions?
What are its preferred growth conditions?
In general, I try not to ingest the bacteria I grow in the lab — even ones with the lowest safety designation, BSL-1. By nature of it being a commercial probiotic, BC30 is both BSL-1 (safe to grow under normal lab precautions) and edible.
But, I still wouldn’t try this at home or eat bacteria off of a culture plate. Why? BC30’s preferred food source is not that different from the preferred food source of many other microorganisms: a sugar- and amino acid-rich nutrient medium called MRS (De Man, Rogosa and Sharpe) agar.
While MRS agar has some adjustments to make it preferentially appetizing to BC30 and its relatives, those relatives also include Streptococcus pyogenes (causes strep throat) and Bacillus cereus (causes food poisoning). Without sterile technique and rigorous species-level confirmation, you cannot know for sure what is growing on your plate.
With that said, the American Society for Microbiology’s blog suggested that were I successful in culturing BC30, I would see growth of translucent white colonies on MRS agar plates after 48 hours of incubation at 30 – 33oC in the presence of oxygen.
First, I needed to make tea.
I wanted the conditions of the experiment to represent a range of realistic tea-drinking scenarios, from intended use to flagrant improvisation, and set up three steeps:
The Rule Follower — Brewed as directed for 4 minutes in boiling water.
The Rule Follower — Brewed as directed for 4 minutes in boiling water.
“I forgot I made tea” — We’ve all been there. 15 minutes, boiling water.
“I forgot I made tea” — We’ve all been there. 15 minutes, boiling water.
Cold brew anarchist — Self-explanatory.
Cold brew anarchist — Self-explanatory.
It was at this point I realized I needed a sterile-ish way to transport the steeped tea and tea bags from my house to the lab. Luckily, I had recently run a blindfolded volume pouring accuracy competition at our departmental retreat and had leftover Falcon tubes still in their original package. While the tea was definitely not sterile, I reasoned that a little extra aseptic technique wouldn’t hurt. I poured the tea into the tubes over my kitchen stove, using the open flame as a makeshift Bunsen burner.
In reality, lab came first. I had to make MRS agar plates before I steeped the tea. Our lab does not use MRS broth very often, and when I first looked for some all I found was a 10-year-old solidified block of MRS powder in our stock cabinet that was growing large green spots inside of its glass container. Behind it was one that looked mercifully normal.
I mixed broth powder, agar and water in a glass bottle, loosely capped it, put it in a water bath and took it to the autoclave. “Autoclave” is a nice word for giant pressure cooker. Ours is made by the aforementioned Steris. It rattled and hissed as its jaws opened to accept my tray of culture media, which it then heated to 121oC for 45 minutes, sterilizing the liquid.
Back at my lab bench, when the molten MRS agar had cooled enough to handle, I lit a Bunsen burner next to a stack of empty plastic petri dishes and poured a layer of agar into each one. Left overnight, the plates solidified into nutrient-dense beds for BC30.
The next day, I took my tubes of tea to lab. I pipetted 400 microliters (0.4mL) of each liquid tea condition onto a plate next to the Bunsen burner. I spread the liquid evenly across the plate with a hockey stick-shaped plastic spreader5 and left the lids on the plates cracked open to dry near the flame.
But to answer my question, I needed one more test. If there were bacteria in the tea bag initially, but they died when boiled, then I might see bacterial growth by plating the dry ingredients of an unsteeped tea bag, or the tea bag steeped in cold water. I cut open the tea bags and shook some of their contents onto the agar. I put my full set of plates, including a plain MRS plate to check its sterility, into the incubator at 37oC — a standard growth temperature, but a little warmer than recommended. I was skeptical that anything would grow. For the next two days, all I could do was wait and see.
The first thing I noticed when I took the plates out of the incubator was the smell. I was in disbelief when I saw little white colonies dotting almost all the plates and opened one to get a closer look. A sickly sweet, gingery aroma wafted from the plate as I inspected the translucent colonies — a byproduct of the bacteria metabolizing the sugars in the MRS plate. By all accounts, I was looking at BC30.
Colonies grew on all of the tea and tea bag plates, while my sterile control plate remained bacteria-free. The colonies from the boiling-water steeps and the tea bags were a variety of sizes, including some that were significantly larger than others, while the colonies from the cold-water tea were uniformly small.
Because I knew the volume of tea I had put on each plate, I could calculate a standard measurement of bacterial density: colony-forming units (CFUs) per milliliter. Contradictory to my expectations, I saw a five-fold increase in colonies from the tea steeped in boiling water relative to the tea steeped in cold water for the four-minute condition. I saw the same pattern in the fifteen-minute condition, with a nearly four-fold increase in boiling vs cold.
Before I could draw any conclusions, I needed to know, for sure, that these colonies were Bacillus coagulans. The most robust way to check is by sequencing their DNA, but sequencing is expensive. A simpler, cheaper way to check is with PCR, which amplifies small regions of DNA unique to a species. I downloaded the BC30 genome and selected two regions of its genome that didn’t match other species in the NCBI database. Using Primer3, I generated two pairs of PCR primers, short stretches of DNA to bind to either side of my region of interest.
I picked the largest colony I could see from each plate (seven total), suspended the cells in a small volume of water, and set up standard colony PCR reactions. The heat during the reaction bursts the cells, making the DNA available for amplification. The completed reaction was run through a porous gel with an electric current and visualized with UV. If I saw bands on the gel for both primer sets, from totally different parts of the BC30 genome, I could be confident that this was, in fact, BC30.
I loaded the gel into the imager and hit run. There, in black relief against the grey background of the gel, were my bands.
Bigelow knew something I didn’t6. It turns out that BC30, like many of its relatives, is a spore-forming bacterium. When starved of nutrients, Bacillus coagulans divides asymmetrically, packing its basic cellular information into a spore with a thick protective coat. These spores are resistant to dessication, nutrient starvation, radiation, chemical disinfectants and extreme heat. It was these spores that were in the tea bag — spores that are perfectly comfortable being steeped in boiling water.
Like the seeds of plants, when the spores find themselves in favorable conditions for growth — say, on an MRS agar plate at a balmy 37oC — they germinate back into actively growing cells. This is the premise of their ability to function as a probiotic. The spores are dormant and shelf-stable in a tea bag, or any of the thousand products advertised to contain BC30, and will, in theory, germinate upon arrival in the GI tract, where they can exert some sort of effect on the host that consumed them.
To produce spores at scale, manufacturers grow bacteria in vats of nutrient-rich broth. If the nutrients are not replenished, the bacteria eventually start to starve, triggering the sporulation process. Around 24 hours later, the bacterial broth is treated with enzymes to kill any remaining, non-sporulated cells. The mixture is concentrated, washed with water, and finally, in a fantastic twist of irony, pasteurized.
The pitch for BC30 is that it improves “digestive health” and “protein absorption.” The reported endpoints for digestive health on BC30’s website are reductions in bowel movement frequency, abdominal pain and abdominal bloating in adults with IBS. In the study promoted on the site, the baseline for the placebo group for abdominal pain and bloating is, mysteriously and respectively, 12.5% and 30% higher than the baseline for the BC30 treatment group. The placebo group experienced no change in severity scores over the subsequent course of treatment, while the BC30 group dropped to placebo levels after a week and stabilized.
For one of the protein absorption studies, there is a small but statistically significant difference in amino acid levels in the blood, including when BC30 is paired with another one of its parent company’s products, a “nutritional milk protein concentrate” called Ultranor.
If these results hold, they beg the question — could a product like Bigelow’s probiotic tea be able to produce these beneficial effects? Most of the clinical trials I could find, including the IBS study above, dosed people daily over the course of one to eight weeks with 1 billion CFUs (spores) of BC30. Per my calculations, a properly steeped cup of probiotic tea yields around 30,000 CFUs: 0.003% of the clinically tested dose.
Granted, I am one person and this is one experiment. But there are independent, conflicting reports on whether BC30 survives the GI tract at all. One study suggests that Bacillus probiotics don’t make it, while another reports about half of the initial dose of spores surviving transit through an artificial human gut system. Other research suggests that the effect of the probiotic is not even due to the cells coming back to life, but due to an immune response against the dormant or vegetative cells.
These observations have consequences for consumers being parted from their money by unsubstantiated health claims. But they are interesting observations in their own right. Sporulating organisms’ imperviousness to heat, while useful for commercial biotech applications, causes problems for the food industry. The food-poisoning agent Bacillus cereus is a species normally found in the soil. It can release heat-resistant toxins if it multiplies in food, and live bacteria can produce toxins when they reach the small intestine. Even pasteurized milk spoils eventually as heat-resistant spores, mostly soil Bacillus, begin to multiply.
Bacillus coagulans is also a soil bacterium by nature7, and not a typical resident of the community of microbes in our gut (called the gut microbiome). It was discovered in 1915, in canned milk that had spoiled and coagulated. In spite of its origins, BC30 seems inert as a pathogen, and reports of it causing infection are vanishingly rare. BC30’s safety track record is remarkable.
But safety is only the first step on the quest to use probiotics for good. New probiotic companies like Pendulum and Seed market the fact that they are backed by clinical trial data — Pendulum for blood sugar control in Type II diabetes, and Seed for gas, bloating and regularity in healthy adults. The backbones of Pendulum and Seed’s products are organisms found more commonly in the gut, and, interestingly, both companies focus on multi-species products, dosing patients with miniature microbial communities.
One focus of my PhD lab is on abnormal pathogenic behavior of normally harmless bacterial residents of the gut, most commonly in people who are already quite sick. While unhappy microbiomes can be unhappy in their own way, we do not have a consensus on what a “healthy” gut microbiome looks like, either. In collaboration with a continent-wide consortium in Africa, our lab helped catalogue the species found in healthy adult women across the continent. We found over 1,000 new species relative to what had been previously described in studies focused on Western countries.
Companies trying to introduce targeted combinations of microbes into the gut are thus forever shooting at a moving target. Outside of specific indications for GI infections, determining whether to give (or take) a probiotic is a grey area. And for the common GI complaints focused on by the probiotic market, targeting the microbiome with additional organisms may not be the answer at all. Rather, by understanding how bacteria work together in the microbiome, solutions may favor changing the metabolic environment of the gut to drive the formation of species-agnostic “guilds” that perform specific functions, likely via dietary interventions.
I never get tired of growing bacteria. For a colony to be visible on a plate, it consists of at least a million, often closer to a billion, individual cells. Learning how bacteria grow and adapt does not diminish the sense of wonder I feel when I observe them — it only enhances it. I like to think this same sense of wonder animated the scientist who first cultured Bacillus coagulans out of canned milk that had spoiled. And I have to imagine some mixture of wonder, awe and horror kept the Random Hall Milk alive for twenty-seven years.
The next Louis Pasteur could be a lactose-intolerant undergrad, or a procrastinating PhD student. It could be you. Pausing to look a little longer, to ask why the world is the way it is — this is how we upend assumptions of what is valuable. What is worth looking at. Because in the end, trash is in the eye of the beholder.
1
You can read the Milk’s application here.
2
As demonstrated by the Milk, even pasteurized beverages spoil, a process sped up by exposure to the microbes in the air but which will proceed within an unopened container anyway. How is this possible?
It’s because pasteurized milk is not the same as sterilized milk. Pasteurization heats to ~60C for a few minutes. While the microbes that we worry about causing infection can’t survive this, some heat-tolerant bacteria and proteins can — those are what will eventually break down the milk, even if it isn’t opened to the air. Heating to 140C, on the other hand, makes milk effectively sterile, killing even the heat-tolerant bacteria. But this process, used to produce “ultra-high temperature” or UHT shelf-stable milk, does some odd things to the proteins that subtly change the flavor, color, and texture of the milk.
3
Author correction 7/28/26: I reported this initially as “7 in 10 people were prescribed antibiotics,” but a commenter on another site pointed out that the original report data likely reflects a smaller number of people getting repeat prescriptions, so I have reported the raw data from the CDC report instead.
4
If a probiotic is marketed as a food or dietary supplement, as most are, then it is not required to undergo a clinical trial in the U.S. but instead to submit a GRAS notification. The second most important thing to know about the GRAS system is that it does not require proof of efficacy — only safety. The most important thing to know is that the proof of safety is provided by the company requesting the GRAS designation. This proof is then reviewed by the FDA, to determine whether the notice provides a “sufficient basis for a GRAS determination” and whether “information in the notice or otherwise available to FDA” raises any safety concerns.
5
This is one of the most polarizing choices one can make as a biomedical research scientist. The alternative to the hockey stick is to use glass beads that you autoclave then sprinkle on the plate and roll around. People are very passionate about their chosen method and will attempt to convert you.
6
Saw this weirdly aggressive Bigelow commercial at the gym. I don’t think they’re going to sponsor me after this article.
7
As our understanding of microorganisms evolves, so do our naming and classification conventions. Bacillus coagulans is a more distant relative of Bacillus cereus and similar soil microbes than previously thought and has been re-classified into a new genus called Weizmannia. Its full nomenclature history can be found on the LPSN.
A Netflix executive was fired from his $1.1 million a year job after revealing during a “trust exercise” at a work retreat that he had taken medically prescribed ketamine, a lawsuit has claimed.
Kevin Baillie, who was vice president and head of creative at Eyeline Studios, is suing the company after it launched an investigation into his comments that ultimately ended in his firing, the papers say.
Baillie, who’s been on the visual effects team for “Pirates of the Caribbean” and the Harry Potter franchise, says he took the drug under medical supervision in October and November of 2022 at a Santa Barbara clinic.
He sought the treatment for clinical depression after the death of his mother, according to the suit.
During what’s called a “Vulnerability-Trust exercise” at a January 2026 retreat at the exclusive Sendero Ranch, a Northern California property owned by Netflix, Baillie shared with his colleagues that he had undergone the treatment, the suit says.
Baillie claims he explained the reason why he had taken the drug but was investigated by Netflix. On March 18, 2026 a company investigator brought the incident up, “in a manner suggesting suspicion of recreational drug use,” the suit reads.
The executive was fired in April with the company’s attorney confirming “the ketamine therapy issue has factored into the termination,” the papers say, and go on to suggest Baille was denied up to a year of severance pay.
Baillie says in the lawsuit the “scope” of the investigation related to alleged profanity and drinking. He had been warned during his performance review that he should “drop one or two less f-bombs but don’t stop entirely.”
It goes on to say that at the same retreat Baillie drank a Guinness standing on his head, after sharing that he had learned the trick from his former father in law during a conversation inspired by the trust session.
“His colleague immediately asked for a demonstration, rather than withhold the openness that the session had encouraged, he performed the trick,” the papers say.
Sign up for the California Morning Report newsletter
California’s top news, sports and entertainment delivered to your inbox every day.
Thanks for signing up!
Baille also paints a picture of an allegedly alcohol-fueled company environment encouraged by the Eyeline Studios CEO Jeff Shapiro.
The lawsuit alleged “alcohol consumption was company-sponsored, leadership-modeled and condoned”, with multiple examples of how Shapiro “set the cultural tone concerning alcohol at the executive level”.
The documents claimed Shapiro on one occasion purchased beer at a corner store and brought it to a company car ride to the Visual Effects Society Awards for staff to share.
Baillie claims he also witnessed the CEO consume alcohol with Netflix and Eyeline employees at his own welcome dinner in September 2024, the Netflix Annual Business Review events in March 2025 and even a Lakers game attended by Netflix execs including Shapiro’s direct supervisor in February 2026.
All up, Baillie’s attorneys provided over half a dozen examples of the CEO being present at a work event with a drink in his hand, according to court papers.
Download The California Post App, follow us on social, and subscribe to our newsletters
California Post News: Facebook, Instagram, TikTok, X, YouTube, WhatsApp, LinkedInCalifornia Post Sports Facebook, Instagram, TikTok, YouTube, XCalifornia Post Opinion California Post Newsletters: Sign up here!California Post App: Download here!Home delivery: Sign up here!Page Six Hollywood: Sign up here!
In addition to hosting multiple parties, Shapiro also had a personal bar in his office “from which he served alcohol (to Baille) including after a successful meeting with Netflix’s CEO Ted Sarandos,” according to the papers.
Baillie is asking for a jury trial, compensatory damages, lost wages, damages for emotional distress, and punitive damages.
Netflix and Eyeline were contacted for comment.
Please enable JS and disable any ad blocker
“But I already have a website on Substack,” you argue.
No, no, Substack is just a distribution tool to amplify your website. It should not be your digital home.
In the last few years, I’ve noticed a pattern of writers leaving their websites to make Substack their digital home.
Now, it’s kinda okay if they have bought a domain and linked it to Substack. (Meaning, it’s better than nothing.)
Rachel from Conscious Living is a good example. This way, Substack more or less functions like a content management system (CMS) for you.
However, compared to other CMS it’s very limited, such as the ability to manage your SEO and customize your pages to add more features, but I digress. If you just want a fuss-free platform, this is one way to get it and Substack’s conditions for domains are very reasonable and cost-efficient. As I will explain later, this could change on a dime without warning.
However, there are some writers who are saying: “Hey readers, I’m now writing on Substack, so head on over there (and ignore my website)!”
Some writers do have a website, but link to their Substacks, calling them their “blogs”. If your Substack has a domain name they own, it’s okay, but if it’s xx.substack.com, Substack is saying “All your content are belong to us”.
In conclusion: Writers, don’t do this. It’s short-sighted and unwise and can derail your long-term visibility on the Internet.
The siren call of convenience
Every few years, the internet convinces writers that a new digital paradise has arrived. First, it was social media like Facebook. Then blogging networks like Tumblr. Then it was Medium. More recently, it’s been Substack.
Platforms promise us an eager audience, built-in monetization, a smooth user interface, and a supportive community. As a writer who just wants to focus on writing, it’s incredibly tempting to hand over the keys to our creative kingdoms and let these portals handle everything. (Believe me, I gave in at one point. For years, I just stopped blogging altogether and even gave up a domain that had high traffic! But I got back in 2012 and never left.)
However, this is the truth that has not changed since the dawn of the Internet: When you build your audience entirely on someone else’s platform, you aren’t a homeowner. You are a tenant. Or worse, a digital sharecropper.
And corporate landlords always change the rules eventually. It’s not personal, it’s just business.
The illusion of the safe space
It’s easy to feel secure when a platform is in its golden era. But we’ve watched the downfalls of Twitter, the policy shifts of Reddit, and the changing tides of algorithmic networks. Relying blindly on a centralized portal not owned by you means your life’s work can alter overnight based entirely on a corporate boardroom decision.
When I looked at how fragile our digital ecosystems really are, I realized I needed a space that wouldn’t go “poof” because a company needed to please its investors or shareholders. This realization completely changed my approach, pushing me to protect my content by learning to blog the IndieWeb way.
Your writing needs a permanent homebase—a domain that you own and control. Full stop.
Moving from “renting” to syndicating
The biggest pushback I hear from writers is: “But my website doesn’t have an audience! Substack does.”
But you don’t have to completely abandon social media or platforms like Substack to protect your autonomy (personally, I prefer the word sovereignty but it does sound a tad dramatic).
You just need to change the order of operations. Instead of publishing directly to a portal, you can shift your mindset to POSSE: Publish (on your) Own Site, Syndicate Elsewhere. (I explain the POSSE/PESOS method in an older post.)
By treating your website as the definitive source of truth and using platforms simply as distribution pipes, you get the best of both worlds. I dug deep into this shift when I committed to being an imperfect gardener of my digital garden, exploring how a less market-y way of presenting my content online let me share my wild garden of thoughts without dancing to the algorithm.
A reality check on platform hype
If you are still holding out hope that Substack is “different” from the social platforms that came before it, let’s look at the numbers and behaviors behind the marketing copy.
After spending a significant amount of time observing the platform ecosystem firsthand, I wrote a brutally honest takeaway in What I learned from one year of Substack. The network effects are real, but so is the pressure to conform to what the platform’s ecosystem favors.
This post, by the way, desperately needs to be updated because things have gotten much, much worse since I wrote it.
When you hand your content over to a platform, you have to conform to their rules and their localized biases. For those of us writing from outside the dominant US-centric echo chambers, platform algorithms heavily prioritize specific western narratives, making it incredibly tough for localized or minority voices to be seen unless they conform.
I wrote about this exact frustration recently in Linkblog: Dwelling on the Internet, highlighting how algorithmic complacency forces us into homogenized bubbles.
The flip side — the writers who refused to leave their websites
Each time there’s a new drama on some platform, and writers are shaking their sabers and declaring that they will leave for yet another social media platform they don’t control, I think about writers like John Scalzi.
As of date, John scalzi has been blogging on https://whatever.scalzi.com/ for 28 years!
This sci-fi novelist has maintained a single independent website continuously for nearly three decades; this makes him one of the longest-running, most consistent original bloggers on the internet. Imagine the amount of digital footprint on that website! Unbroken by time or platforms.
(Specifically, he uses wordpress.com like I do, as we both don’t want to bother with the pain of setting up your own self-hosted wordpress website and just want the folks at Automattic to do it.)
He blogs in the classic Indieweb way, though I doubt he is even aware he’s doing it. He treats his social media channels such as X or Bluesky as a way to amplify his website. All roads lead back to https://whatever.scalzi.com/, and this is something I wish every single writer would do.
He wrote recently in Various & Sundry, 6/3/26:
this site acts as my own institutional memory, if I post something about it here it constitutes an official record. I mean, all the posts I ever placed on the former Twitter are now entirely lost to time, since I have gone in and purged my entire timeline there. This site, however, endures. — John Scalzi
this site acts as my own institutional memory, if I post something about it here it constitutes an official record. I mean, all the posts I ever placed on the former Twitter are now entirely lost to time, since I have gone in and purged my entire timeline there. This site, however, endures. — John Scalzi
Breaking free from platform blues
Trying to adapt your presence across various platforms in an ever-shifting digital landscape is exhausting. One minute a platform is a writer’s darling; the next, it’s being boycotted. Railing against a platform’s focus shift or the presence of (long sigh) Nazis is a useless endeavor.
As I noted in Linkblog March 12, 2026: Platform blues, chasing platform purity is an illusion. Tech will change, corporate algorithms will continue to prioritize profit over human connection, and platforms will continue to cycle through hype and decline.
The antidote to this exhaustion isn’t moving to the next shiny new app. It’s anchoring your work on an independent website with open distribution channels like RSS. It also means ruthlessly using platforms as distribution channels. When one collapses or you prefer to just move, it’s easy to just change strategies because your digital home remains unchanged.
Use platforms to find your readers, but bring them back to your house. It’s time to stop digital sharecropping on rented land.
Featured photo is by vivek vk on Unsplash
July 27, 2026
I’ve been using Claude and ChatGPT like the next guy for probably two years now. I’ve never been a huge “open software” nerd or anything like that. But just now, I got opencode working on my own inference endpoint and… it felt surprisingly good. It feels… freeing, somehow. I own the endpoint, and my data just goes from my laptop to there and back. It feels like it’s mine. It’s really nice.
The motivation for this was that I just got home and wanted to start on a little side project, but I don’t have the best Claude or ChatGPT plan for my personal account. I work at Modal, and today we just launched Kimi K3 on managed endpoints, and I know Kimi K3 is supposed to be pretty solid, so instead of upgrading my Claude plan, I wanted to give it a try. (I didn’t directly contribute to this feature, so I haven’t gotten to play with it yet.)
Within about 5 minutes, I had opencode pointed at my own Modal endpoint and running. Spinning up opencode, I just felt a nice, empty blankness. The best way I can describe it is like opening vim after spinning a bunch of time in a big fancy editor. Or maybe like the tendrils tying me to other providers had been cut, and I could breathe freely. I’m being somewhat dramatic here but not exaggerating that much. It’s weird, lol, and unexpected to me, which is why I wanted to write about it.
Cost vs. quality on catalog integrity
Part I
Everyone asked the same question
Since ChatGPT launched in 2022, business leaders have been asking the same question: what can AI do for us? The answer began with low-risk tasks: summarizing documents, drafting emails, producing first drafts that a human would edit.
It quickly moved into higher-value cognitive work, such as software development and content generation, and grew into more ambitious projects, like attempts to build an AI company brain, a system connected to internal knowledge, data, and tools that could coordinate work and eventually operate parts of the business autonomously.
While a lot of time, energy and tokens have been invested in AI adoption, measurable outcomes have barely been achieved at scale. However, some companies embraced being AI-first and saw enormous gains in productivity, revenue, and cost, while others lagged behind or failed to change their organizations enough to reach high ROI.
Recent data from corporate expense management platform Ramp reveals a stark contrast in performance: the top quartile of companies investing in AI saw their revenue more than double between November 2022 and December 2025, while businesses with zero AI expenditure experienced a mere 15% increase.
Top AI adopters outperform across industries
top quartile of AI spenders businesses with zero AI spend
There are many reasons why AI has done wonders for some companies while others have struggled to see the return on their investment, but research primarily points in five directions.
01
Redesign the process, not just the task
Becoming AI-first means rethinking how the work is structured, not dropping a model into a workflow built around people: what gets approved, who reviews what, and which handoffs still need a human. Where the process stays untouched, legacy bottlenecks absorb the productivity gains before they reach the P&L. In McKinsey’s 2025 survey of organizations using gen AI, workflow redesign was the attribute most correlated with EBIT impact, and only 21% of them had redesigned any workflow at all.
02
Incentivize experimentation
Models, tooling and best practices change weekly, so last quarter’s setup is rarely still the right one. That only gets picked up if people are rewarded for trying things and reporting what failed, not just for shipping. Technical teams are the natural place to start, since they see the same problems recur across functions and can tell which of them a model can actually take over.
03
Provide tailored business context
Prompt engineering and retrieval can inject business context at call time, but doing it well is its own engineering program: getting to the data, enforcing access controls on what each request may see, building retrieval that surfaces the right evidence, and managing a context window that models use unevenly as it grows.
04
Measure usage and impact
Every AI line item eventually meets the CFO question: what did this change, and was it worth it? In most deployments, nobody can answer it: there is no infrastructure to track the model’s performance, decision costs, or impact on efficiency, and self-reported time savings are often inaccurate. Without a scored evaluation on your own data, a “vibe evaluation” is the ceiling of what you can claim, and a hard budget to defend.
05
In this article we give a detailed overview of the deployment technique the winning group keeps converging on: fine-tuning open-source models with reinforcement learning. We cover how it addresses the last three challenges above, and how it turns knowledge only your organization has (namely data, tools, and processes) into a model no vendor API can match at a fraction of the cost.
TL;DR
2.2×
Revenue growth of the top quartile of AI spenders between November 2022 and December 2025 in Ramp’s data. Companies with zero AI spend grew about 15% over the same three years, in the same economy: the heavy adopters grew eight times as much.
1
Playbook the winners converge on: an open-source model, proprietary task data, and reinforcement learning against a scored copy of the workflow. Bridgewater’s trained model makes ~30% fewer mistakes than the best frontier model, Harvey’s legal agent beats GPT-5.5 and Claude Opus 4.8 on its own rubrics, and Intercom’s Fin Apex resolves more support issues at lower cost.
87.3%
Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling.
68×
Cost advantage per reviewed listing: $0.50 per 1,000 with the specialist vs $34 with the strongest frontier model, and still 40× cheaper than the least expensive frontier option. At roughly 40 million decisions a day, that is about $7M a year instead of $500M, a 98% cost reduction.
Part II
What the winners do differently
Most of the companies pulling ahead in the AI race made the same discovery: owning your intelligence wins on both performance and cost. A model trained to complete your specific workflows in your specific environment is very likely to outperform a general-purpose model that has never seen inside your company. Additionally, since you do not need to pack as many general-purpose capabilities into a model that is meant to operate in a specific environment, you can often get away with a smaller model that is orders of magnitude cheaper to run.
Owning your intelligence does not mean cancelling the ChatGPT or Claude subscription. Most workflow automation still starts with frontier models, and that is the right first move: it establishes a baseline for what is technically possible, and every call generates the data (inputs, decisions, corrections) that a specialist model later trains on. Once the automation leaves the prototyping stage, the priority flips to cost and performance at volume, and that is where fine-tuning open-source models with reinforcement learning comes in.
In addition to that, your Fable 5 or ChatGPT model can call the specialist model to handle the parts of the workflow that require your internal knowledge, and the specialist model can call the frontier model for tasks that require high general ability.
How a model learns to operate in your environment
Over the past two years this has hardened into a playbook: an open-source model, proprietary task data, and a reinforcement-learning stage against a scored version of the workflow. Below, we discuss three scenarios where this approach has been applied to real-world tasks.
Bridgewater Associates is one of the largest hedge funds in the world. Its analysts sift a constant stream of articles, filings, and emails, judging which documents are relevant to the firm’s investment thesis and where boilerplate content begins. The catch is that relevant means relevant by Bridgewater’s internal judgment, and no amount of prompting got frontier models to absorb that judgment reliably. So, the company decided to train an open-source model on labels from its own expert investors. The trained model makes roughly 30% fewer mistakes than the best frontier model, at a fraction of the inference cost.
Harvey builds AI agents for law firms. Its hardest workloads are long-horizon: transaction due diligence and legal memo drafting, where the agent navigates large document sets, errors compound across steps, and even the best frontier models at maximum reasoning effort kept falling short of the quality bar. Harvey ran reinforcement learning on an open-weight model and got a legal agent that outperforms both GPT-5.5 and Claude Opus 4.8 on its rubrics.
Intercom is a customer-service platform whose AI agent, Fin, resolves almost two million customer issues a week. At that volume the problem is unit economics: frontier per-call pricing adds up fast, and every point of resolution rate matters. So Intercom’s AI group post-trained its own vertical support model, Fin Apex, on billions of customer-service interactions. Intercom reports that it resolves more issues than the best frontier models while being cheaper to run.
The same shape repeats well beyond these three. The appendix collects eight more deployments, with what each model was trained to do and what changed once it shipped.
One deployment pattern repeats across these cases: reinforcement learning pushes an open-source model past the frontier on a specific set of workflows, at a fixed and dramatically lower cost per call. The deployment typically starts by using prompt and context engineered frontier models to establish the strongest default baseline. Those frontier traces and learnings are then reused to pack the focused capability into a compact model the company owns.
Part III
Case Study: Catalogue Integrity Agent
One of the central challenges of running an e-commerce platform is keeping the product catalog trustworthy. Every listing must land in the correct category of the platform’s taxonomy, and the attributes that power search, filters, recommendations, and downstream operations must be accurately extracted from its images and description.
+15%our fine-tuned 9B open-source model reviews e-commerce product listings more accurately than the best frontier model we tested
68×cheaper per listing when our fine-tuned model does the reviewing instead of that frontier model
Platforms staff this work with teams of catalog-integrity analysts, which grow together with the catalog: more listings mean more categories to know, more attributes to check, and more ambiguous edge cases to judge. Inconsistent decisions propagate: products become harder to find, recommendations deteriorate, and policy violations slip through. Miss a counterfeit and you expose customers and brands to fraud; over-flag and you build an expensive review queue that frustrates legitimate sellers.
Let’s start with the scale. eBay alone carries about 2.5 billion live listings, and Shopify’s catalog absorbs more than 10 million product updates a day. Walmart has said that doing its AI-assisted catalog work with people alone would have taken roughly 100× the headcount.
Getting it wrong costs revenue: 71% of shoppers say they have returned a product because it didn’t match its listing. But the reviewing itself is also expensive. A mid-size marketplace generates on the order of 10 million listing creates and edits a day; reviewed with a frontier model, that workload costs roughly half a billion dollars a year, while a fine-tuned specialist does the same job for around $10M:
One workload, two price tags
frontier API per call fine-tuned specialist, self-hosted
The companies actually running this workflow confirm the math. Shopify classifies products with fine-tuned open models at roughly 40 million inferences a day, a volume it says commercial APIs can’t economically serve. Inspired by the industry leaders, we ran the playbook ourselves: an agent that examines a listing, searches the product taxonomy, checks the brand, retrieves the attribute schema, and commits a structured decision, escalating to human review when the evidence is thin or the risk too high.
One episode, end to end
ListingWork gloves, PU-coated, black, size 10
Claimed brandAmazonBasics
RegionUS
1search_taxonomy → Safety Work Gloves
2lookup_brand → registered, not protected
3get_attribute_schema → brand, color, material, size
✓Commits category, attributes, and verdict allowed
Part IV
A digital twin the model can practice in
A model can only practice a workflow it can actually perform, fail at, and retry. So we rebuilt the catalog workflow as a digital twin of an e-commerce platform: a simulated environment with the same listings, the same tools, and the same stakes as the real thing. The raw material is the Amazon Berkeley Objects dataset of real product images and listing data, which we turned into 177,767 review episodes: one listing each, with image, title, description, claimed brand, and region. Into that stream we planted controlled policy cases, mismatched images, conflicting brand claims, and deliberately legitimate claims as hard negatives, so every episode has a known correct answer to score against.
Inside the twin, the model works the way an analyst would. It can search a taxonomy of roughly 13,000 categories, check whether a brand is registered and protected, and retrieve the required attributes for a chosen category; then it must commit a category, the attributes, and a policy decision. A scorer grades every episode, rewarding correct outputs and penalizing missed violations, unsupported attributes, invalid categories, and wasted tool calls. The penalties encode business priorities directly: a missed violation costs 7× more than a false alarm.
The digital twin, end to end
TOOLS
search_taxonomy() lookup_brand() get_attribute_schema()
listing agent decision score
reward
0.3·category + 0.3·attributes + 0.4·policy − tool_overage
Benchmarking the frontier models
Before any training run, we measured how far the frontier could get on its own. We benchmarked five frontier models (GPT-5.5, GPT-5.6-sol, Gemini 3.1 Pro, Claude Opus 4.8, and Claude Fable 5) on 200 stratified validation episodes with identical tools, images, scorer, and turn budget, both with a plain prompt and with optimized prompt instructions: 2,800 characters of extraction conventions, lookup procedure, and worked examples, tuned the way a prompt engineer would tune a production system. The chart below shows both configurations, with our final GRPO-trained model added for comparison.
The best frontier configuration reached 76.9% of the achievable score; the trained 9B reached 87.3%. The gap is not intelligence. A frontier model starts every episode from zero: it has never seen this store’s taxonomy, doesn’t know its inventory conventions, which attribute values count as supported, or how the platform wants corner cases resolved, and it has to reconstruct all of that on the fly, from whatever fits in the prompt, on every single call. Optimized instructions can compress some of that knowledge, but the corner cases that decide the score are exactly the ones no instruction prompt can enumerate.
Model benchmarks
Part V
Training details
The training setup was modest. We rented two RTX PRO 6000 GPUs, one generating rollouts and one applying gradient updates, with the open-source prime-rl framework running the infrastructure. The full run was 1,000 optimizer steps, took about three and a half days, and cost roughly $500 in GPU time.
Most of that was not needed to catch the frontier. The model crossed the frontier band after roughly 250 steps, about a day of training; the remaining steps were spent squeezing out maximum performance. The final model scores 0.626 on the benchmark, 87.3% of the achievable ceiling and about ten points above the best frontier configuration.
The gain comes from teaching the model the semantics of its environment: over thousands of scored episodes it learns how the store’s taxonomy, tools, and policies relate to the reward, and converges on the optimal distribution of actions over the environment’s action space. A general model spreads its capacity across everything; the specialist spends all of it on the specific task at the expense of general ability.
The result also isn’t a dead end. When a more capable open-source base model ships, the recipe transfers: as a rule of thumb, a stronger base yields a stronger post-trained specialist. And the deployed model generates its own training data as it works: logged decisions can be distilled into supervised fine-tuning sets, making each re-training cheaper and better-informed than the last.
Training run
So far, we have discussed how reinforcement learning helps a language model navigate your company environment more accurately. But one of the most important aspects of this paradigm is that the accuracy also comes cheaper than running frontier models. The cost-vs-quality chart that opens this article puts both on one picture, task score against cost per thousand listings, and the takeaway for a budget owner is simple: with everything else on the chart you are choosing between quality and cost. The trained specialist ends that trade-off, delivering the best score we measured at close to the lowest price.
Fine-tuning adds ~23 points over the base 9B at the same ~$0.50 per 1,000 listings: 40× cheaper than the least expensive frontier configuration (Gemini, $19/1k) and ~340× cheaper than the most expensive (GPT-5.5-pro, $172/1k). The 2,800 characters of prompt instructions also raised GPT-5.5′s measured cost by a third, a prompt tax paid on every future call. The specialist’s instructions live in its weights. At Shopify-scale volume of roughly 40 million decisions a day, the gap between $34 and $0.50 per thousand is about $500M a year versus $7M.
Wondering where these curves sit for your workflow? Book a 30-minute audit →
Part VI
What can owning AI do for you?
Here is the whole experiment in three sentences. We took one high-volume e-commerce workflow, reviewing product listings, and rebuilt it as a digital twin. We let a small open-source model practice inside it for a few days and roughly $500 of GPU time. The trained specialist ended above every frontier configuration we measured, at a fraction of their price per decision.
The recipe transfers to any work your business repeats at volume. Look for the places where people or API calls turn information into decisions all day: routing tickets, extracting fields from documents, checking submissions against policy, classifying products, approving or flagging transactions. What unites them is that every decision can be checked: there is a rule, a schema, or an expert who can say whether it was right. That check is the whole trick. If a decision can be scored, a model can practice it; if it can only be debated, it cannot.
It is just as important to know where this is the wrong tool. Two questions settle most cases: how frequently the task runs, and whether its outcome is verifiable.
Right tool for the right-shaped problem
prompt-optimized frontier model fine-tuned specialist owned, practiced, cheap at scale frontier model frontier model + human review
verifiable non-verifiable
rare task frequency frequent
Your workflow is a fine-tuning candidate if at least one point applies…
It happens at high volume, often enough that per-decision cost and errors compound into real money
Every outcome can be checked by a rule, test, or rubric, with no person in the loop
A note on notation: this article defaults to bra-ket notation because (in my quantum-inspired opinion) it makes the shapes in this derivation very clear. The Math notation switch above rewrites every equation using conventional bold vectors and explicit transposes instead. In bra-ket mode, ∣q⟩\lvert q\rangle is a column vector, ⟨k∣\langle k\rvert is a row vector, ⟨k∣q⟩\langle k\rvert q\rangle is a number, and ∣v⟩⟨k∣\lvert v\rangle\langle k\rvert is a matrix. Vectors face right by default, while keys face left when written into the linear-attention state. We work with one causal attention head and real-valued vectors, assume DeltaNet’s keys are normalized, and let the state map from key space to value space.
Modern linear attention variants are complex, and a upon first glance it is not so easy to see what they are designed to achieve. For reference here is the state update equation for Kimi Delta Attention (KDA):
S~t=St−1Diag(αt)\widetilde S_t = S_{t-1}\operatorname{Diag}(\alpha_t) ∣v^t⟩=S~t∣kt⟩\lvert\widehat v_t\rangle = \widetilde S_t\lvert k_t\rangle ∣et⟩=βt(∣vt⟩−∣v^t⟩)\lvert e_t\rangle = \beta_t \left( \lvert v_t\rangle-\lvert\widehat v_t\rangle \right) St=S~t+∣et⟩⟨kt∣S_t = \widetilde S_t+\lvert e_t\rangle\langle k_t\rvert ∣ot⟩=St(dk−1/2∣qt⟩)\lvert o_t\rangle = S_t\left(d_k^{-1/2}\lvert q_t\rangle\right)
The reason they are so difficult to understand is that this is the latest in a family of linear attention variants that have been developed over the last few years and the complexity of them has inevitably ballooned such that from the outside the latest variants appear inaccessible.
In this post we are going to walk through the DeltaNet family of linear attention variants, two of which are used by the latest Qwen and Kimi model families, and show how you might have arrived at the same equations by asserting simple things about your hidden state.
That is the route we will take:
softmax attention → linear attention → DeltaNet → Gated DeltaNet → KDA
Only after deriving KDA will we turn to the recurrent and chunkwise Triton programs that execute it.
1. Begin with quadratic attention
For a query at token tt, ordinary causal softmax attention is
ati=exp (s⟨ki∣qt⟩)∑j≤texp (s⟨kj∣qt⟩),s=dk−1/2,∣ot⟩=∑i≤tati∣vi⟩.\begin{aligned} a_{ti} &= \frac{ \exp\!\left(s\langle k_i\rvert q_t\rangle\right) }{ \sum_{j\leq t} \exp\!\left(s\langle k_j\rvert q_t\rangle\right) }, \qquad s=d_k^{-1/2},\\ \lvert o_t\rangle &= \sum_{i\leq t}a_{ti}\lvert v_i\rangle. \end{aligned}
Every attention weight is a scalar. It measures the similarity between one key and one query, then softmax turns all of the scores for that query into a distribution. The output is a weighted sum of value vectors.
Over a sequence of length TT, there are T2T^2 key-query pairs. During autoregressive inference we can cache the keys and values instead of recomputing them, but the cache still grows with the sequence and every new query still has to inspect the entire history.
The obstacle to rearranging this computation is the softmax. Its denominator depends jointly on the current query and every earlier key. So, for the moment, remove it.
1.1 Remove the softmax
For clarity, absorb the constant scale ss into the query. The deliberately bare version of attention is then
∣ot⟩=∑i≤t⟨ki∣qt⟩∣vi⟩.\lvert o_t\rangle = \sum_{i\leq t} \langle k_i\rvert q_t\rangle \lvert v_i\rangle.
The scalar inner product can move to the right:
∣ot⟩=∑i≤t∣vi⟩⟨ki∣qt⟩=(∑i≤t∣vi⟩⟨ki∣)∣qt⟩.\begin{aligned} \lvert o_t\rangle &= \sum_{i\leq t} \lvert v_i\rangle \langle k_i\rvert q_t\rangle\\ &= \left( \sum_{i\leq t} \lvert v_i\rangle\langle k_i\rvert \right) \lvert q_t\rangle. \end{aligned}
Everything that depends on the past can now be collected into one matrix of a fixed size V×KV \times K:
St=∑i≤t∣vi⟩⟨ki∣\boxed{ S_t = \sum_{i\leq t} \lvert v_i\rangle\langle k_i\rvert }
and attention becomes a recurrent write followed by a read:
St=St−1+∣vt⟩⟨kt∣,∣ot⟩=St∣qt⟩.\boxed{ \begin{aligned} S_t &= S_{t-1} + \lvert v_t\rangle\langle k_t\rvert,\\ \lvert o_t\rangle &= S_t\lvert q_t\rangle. \end{aligned} }
The identity
(∣v⟩⟨k∣)∣q⟩=⟨k∣q⟩∣v⟩\left(\lvert v\rangle\langle k\rvert\right)\lvert q\rangle = \langle k\rvert q\rangle\lvert v\rangle
is the whole trick. The outer product is a matrix; the inner product is a number. We no longer store every past key and value. We store their summed outer products in the fixed-size state StS_t.
This is linear in sequence length rather than quadratic: scan the tokens once, updating the same dv×dkd_v\times d_k state at every step. We have paid for that efficiency by discarding softmax’s normalization and selectivity. More sophisticated linear-attention methods use feature maps and normalizers, but this unadorned form exposes the memory problem that motivates DeltaNet.
1.2 Addition is not assignment
Suppose we write a pair ∣vt⟩⟨kt∣\lvert v_t\rangle\langle k_t\rvert and immediately query the new state with that same key:
St∣kt⟩=(St−1+∣vt⟩⟨kt∣)∣kt⟩=St−1∣kt⟩+∣vt⟩⟨kt∣kt⟩⏟1=St−1∣kt⟩+∣vt⟩.\begin{aligned} S_t\lvert k_t\rangle &= \left( S_{t-1} + \lvert v_t\rangle\langle k_t\rvert \right) \lvert k_t\rangle\\ &= S_{t-1}\lvert k_t\rangle + \lvert v_t\rangle \underbrace{\langle k_t\rvert k_t\rangle}_{1}\\ &= S_{t-1}\lvert k_t\rangle+\lvert v_t\rangle. \end{aligned}
The write does not make the memory return ∣vt⟩\lvert v_t\rangle. It adds ∣vt⟩\lvert v_t\rangle to whatever the memory already returned.
If the old state already produced the correct value, the additive write makes the new state produce twice that value. More generally, keys are not mutually orthogonal, so every write can interfere with previous writes. Linear attention has given us a compact associative memory, but its update behaves like += when what we want is closer to =.
2. DeltaNet: write the error, not the value
DeltaNet replaces the unconditional linear-attention write with a delta-rule correction. There are two useful ways to derive it.
2.1 Derivation one: demand that the write can be read back
Before writing token tt, ask the memory what it currently associates with the new key:
∣v^t⟩=St−1∣kt⟩.\lvert\widehat v_t\rangle = S_{t-1}\lvert k_t\rangle.
If we want the memory to return ∣vt⟩\lvert v_t\rangle, we should not add the whole value. We should add only the difference:
∣vt⟩−∣v^t⟩.\lvert v_t\rangle-\lvert\widehat v_t\rangle.
Introduce a learned write strength βt∈[0,1]\beta_t\in[0,1] and define
∣et⟩=βt(∣vt⟩−St−1∣kt⟩).\lvert e_t\rangle = \beta_t \left( \lvert v_t\rangle - S_{t-1}\lvert k_t\rangle \right).
Then write this error at the current key:
St=St−1+∣et⟩⟨kt∣.\boxed{ S_t = S_{t-1} + \lvert e_t\rangle\langle k_t\rvert. }
Now immediately read the same key:
St∣kt⟩=St−1∣kt⟩+∣et⟩⟨kt∣kt⟩=(1−βt)St−1∣kt⟩+βt∣vt⟩.\begin{aligned} S_t\lvert k_t\rangle &= S_{t-1}\lvert k_t\rangle + \lvert e_t\rangle \langle k_t\rvert k_t\rangle\\ &= (1-\beta_t)S_{t-1}\lvert k_t\rangle + \beta_t\lvert v_t\rangle. \end{aligned}
When βt=1\beta_t=1, the result is exactly ∣vt⟩\lvert v_t\rangle. Smaller βt\beta_t moves the old prediction partway towards the target.
The correction is also local in key space. For any query ∣x⟩\lvert x\rangle orthogonal to the current key,
⟨kt∣x⟩=0⟹(St−St−1)∣x⟩=∣et⟩⟨kt∣x⟩⏟0=0.\langle k_t\rvert x\rangle=0 \quad\Longrightarrow\quad (S_t-S_{t-1})\lvert x\rangle = \lvert e_t\rangle \underbrace{\langle k_t\rvert x\rangle}_{0} =0.
So the rank-one write changes the response in the selected key direction while leaving every orthogonal direction alone.
2.2 Derivation two: take one step on reconstruction loss
The same update falls out of an online learning objective. Treat the current key-value pair as one training example for the linear map SS:
Lt(S)=12∥S∣kt⟩−∣vt⟩∥22.\mathcal L_t(S) = \frac12 \left\| S\lvert k_t\rangle-\lvert v_t\rangle \right\|_2^2.
Its gradient with respect to the state is
∇SLt(S)=(S∣kt⟩−∣vt⟩)⟨kt∣.\nabla_S\mathcal L_t(S) = \left( S\lvert k_t\rangle-\lvert v_t\rangle \right) \langle k_t\rvert.
This is visibly an outer product: a value-space prediction error times the key bra at which that error was observed. Take one gradient-descent step of size βt\beta_t from St−1S_{t-1}:
St=St−1−βt∇SLt(St−1)=St−1−βt(St−1∣kt⟩−∣vt⟩)⟨kt∣=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣.\begin{aligned} S_t &= S_{t-1} - \beta_t\nabla_S\mathcal L_t(S_{t-1})\\ &= S_{t-1} - \beta_t \left( S_{t-1}\lvert k_t\rangle-\lvert v_t\rangle \right) \langle k_t\rvert\\ &= S_{t-1} + \beta_t \left( \lvert v_t\rangle-S_{t-1}\lvert k_t\rangle \right) \langle k_t\rvert. \end{aligned}
This is exactly the update we got by requiring immediate reconstruction. The two interpretations are the same:
as a memory operation, βt\beta_t controls how strongly to replace the old association;
as online learning, βt\beta_t is the step size;
as linear algebra, the change is a rank-one outer product.
2.3 The DeltaNet state transition
Expanding the error exposes DeltaNet as a structured state transition plus a new input:
St=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣=St−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1} + \beta_t \left( \lvert v_t\rangle-S_{t-1}\lvert k_t\rangle \right) \langle k_t\rvert\\ &= S_{t-1} \left( I-\beta_t\lvert k_t\rangle\langle k_t\rvert \right) + \beta_t\lvert v_t\rangle\langle k_t\rvert. \end{aligned}
For a unit key, I−βt∣kt⟩⟨kt∣I-\beta_t\lvert k_t\rangle\langle k_t\rvert has eigenvalue 1−βt1-\beta_t in the current key direction and eigenvalue 11 in every orthogonal direction. It removes the old association along the current key before adding the new one.
DeltaNet fixes the write. It does not yet fix the lifetime of the state.
3. Gated DeltaNet: sometimes old information should disappear
The linear state compresses the whole history into one matrix. A read
St∣q⟩=∑i≤t⟨ki∣q⟩∣vi⟩S_t\lvert q\rangle = \sum_{i\leq t} \langle k_i\rvert q\rangle\lvert v_i\rangle
cannot choose to skip an individual old token after that token has been folded into StS_t. Every stored direction that overlaps the query contributes. The delta rule can correct the state around the current key, but stale information in other directions remains available and can distort future reads.
We therefore need a way to forget the old state before using it. Let αt∈[0,1]\alpha_t\in[0,1] be a learned scalar retention gate:
S~t=αtSt−1.\widetilde S_t = \alpha_t S_{t-1}.
Run the same delta rule against this gated state:
S~t=αtSt−1,forget,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write.\boxed{ \begin{aligned} \widetilde S_t &= \alpha_tS_{t-1}, &&\text{forget},\\ \lvert\widehat v_t\rangle &= \widetilde S_t\lvert k_t\rangle, &&\text{predict},\\ \lvert e_t\rangle &= \beta_t \left( \lvert v_t\rangle-\lvert\widehat v_t\rangle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\rangle\langle k_t\rvert, &&\text{write}. \end{aligned} }
This is Gated DeltaNet. The order matters: forget first, predict from the retained state, then correct that prediction. If we predicted before forgetting, the error would describe a different memory from the one we update.
Expanding the recurrence gives
St=αtSt−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.S_t = \alpha_tS_{t-1} \left( I-\beta_t\lvert k_t\rangle\langle k_t\rvert \right) + \beta_t\lvert v_t\rangle\langle k_t\rvert.
The delta rule gives targeted replacement; the scalar gate gives global erasure. They solve different problems and are complementary.
But αt\alpha_t still makes one decision for the entire matrix. The model must retain or forget every key channel at the same rate.
4. Kimi Delta Attention: forget each channel independently
Kimi Delta Attention replaces Gated DeltaNet’s scalar retention with a vector αt∈[0,1]dk\alpha_t\in[0,1]^{d_k}. Put the vector on the diagonal:
Dt=Diag(αt)∈Rdk×dk.D_t = \operatorname{Diag}(\alpha_t) \in\mathbb R^{d_k\times d_k}.
Our state maps keys to values, so the key channels are the columns of SS. Right-multiplication applies a different retention factor to every one:
S~t=St−1Dt.\widetilde S_t = S_{t-1}D_t.
Everything else is the delta rule we have already derived:
S~t=St−1Dt,forget each key channel,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write,∣ot⟩=St(s∣qt⟩),s=dk−1/2,read.\boxed{ \begin{aligned} \widetilde S_t &= S_{t-1}D_t, &&\text{forget each key channel},\\ \lvert\widehat v_t\rangle &= \widetilde S_t\lvert k_t\rangle, &&\text{predict},\\ \lvert e_t\rangle &= \beta_t \left( \lvert v_t\rangle-\lvert\widehat v_t\rangle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\rangle\langle k_t\rvert, &&\text{write},\\ \lvert o_t\rangle &= S_t(s\lvert q_t\rangle), \qquad s=d_k^{-1/2}, &&\text{read}. \end{aligned} }
That is KDA. Compared with Gated DeltaNet, the conceptual change is only the promotion
αt⟶Dt=Diag(αt).\alpha_t \quad\longrightarrow\quad D_t=\operatorname{Diag}(\alpha_t).
The effect is substantial: one channel can be cleared while another is retained.
4.1 Why the transition is diagonal-plus-low-rank
Expand the KDA correction:
St=St−1Dt+βt(∣vt⟩−St−1Dt∣kt⟩)⟨kt∣=St−1Dt(I−βt∣kt⟩⟨kt∣)⏟At+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1}D_t + \beta_t \left( \lvert v_t\rangle - S_{t-1}D_t\lvert k_t\rangle \right) \langle k_t\rvert\\ &= S_{t-1} \underbrace{ D_t \left( I-\beta_t\lvert k_t\rangle\langle k_t\rvert \right) }_{A_t} + \beta_t\lvert v_t\rangle\langle k_t\rvert. \end{aligned}
The key-space transition is
At=Dt−βtDt∣kt⟩⟨kt∣=Dt−∣bt⟩⟨at∣,\begin{aligned} A_t &= D_t-\beta_tD_t\lvert k_t\rangle\langle k_t\rvert\\ &= D_t-\lvert b_t\rangle\langle a_t\rvert, \end{aligned}
where
∣bt⟩=Dt∣kt⟩,⟨at∣=βt⟨kt∣.\lvert b_t\rangle=D_t\lvert k_t\rangle, \qquad \langle a_t\rvert=\beta_t\langle k_t\rvert.
So AtA_t is a diagonal matrix minus a rank-one matrix: a diagonal-plus-low-rank, or DPLR, transition. “DPLR” describes the dk×dkd_k\times d_k transition acting on key space. The memory state itself is still the dv×dkd_v\times d_k matrix StS_t.
The full journey can now be summarized compactly:
The implementation usually stores gt=logαtg_t=\log\alpha_t with gt≤0g_t\leq0, then obtains the retention factors as exp(gt)\exp(g_t). In the transposed dk×dvd_k\times d_v layout used by the reference code, the recurrence is only five lines:
state = state * g_t.exp().unsqueeze(-1) prediction = einsum(“bhkv,bhk->bhv”, state, k_t) residual = beta_t.unsqueeze(-1) * (v_t - prediction) state = state + einsum(“bhk,bhv->bhkv”, k_t, residual) output = einsum(“bhk,bhkv->bhv”, q_t * scale, state)
See the official naive_recurrent_kda reference.
To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".
10HN is also available as an iOS App
If you visit 10HN only rarely, check out the the best articles from the past week.
Visit pancik.com for more.