topへ
10 interesting stories served every morning and every evening.
10 interesting stories served every morning and every evening.
topへ
Highlights:
Scientists at La Jolla Institute for Immunology, Scripps Research have developed an HIV vaccine that trains immune cells to see past HIV’s defenses.
This HIV vaccine works by prompting the body’s immune system to make substantial numbers of rarely seen “broadly neutralizing” antibodies.
In this new study, this vaccine resulted in the best HIV-fighting antibody response ever seen in primates. Human trials have now started.
LA JOLLA, CA—A new HIV vaccine developed by La Jolla Institute for Immunology (LJI), Scripps Research scientists, and IAVI has the potential to protect humans from developing HIV infection and AIDS. This HIV vaccine is the first to generate a high number of “broadly neutralizing,” virus-fighting antibodies in primates.
“This feels like a huge success,” says LJI Professor and Chief Scientific Officer Shane Crotty, Ph.D., who co-led the research with Scripps Research Professor William Schief, Ph.D. “We constructed a successful vaccine from the ground up, which required a deep understanding of the immune system.”
This groundbreaking research, published in Nature, is the result of 14 years of collaboration between La Jolla Institute for Immunology and Scripps Research, as part of the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD). “This has been one of those Apollo moon mission-type projects, where there is an exceptional goal and the team has to accomplish a myriad of discoveries and inventions along the way,” says Crotty.
Outsmarting HIV
The new vaccine works by intervening in a process called B cell maturation. B cells make antibodies. Like many immune cells, B cells have an early “naive” stage before they are ready to make antibodies. B cells start to mature once they get the signal that a pathogen, such as a virus, is trying to attack. B cells see pieces of that pathogen’s molecular structure and start producing antibodies that can bind to that structure and halt infection.
It can take a little while for B cells to find the right “bullseye” on a pathogen. But B cells keep trying. As they mature, B cells tweak their antibody production, refining antibody structures to bind to a pathogen in just the right, vulnerable spots.
Scientists describe B cell development as a training process or bootcamp. In most cases, the body is left with a well-honed B cell army.
HIV is hard to beat because it doesn’t give B cells a chance to develop effective antibodies. The first problem is that HIV disguises itself from the immune system. The virus is wrapped in an ever-shifting cloak of sugar molecules, called glycans. This lets HIV sneak undetected past human cells, which are also covered in glycans.
The second big problem is that HIV mutates very quickly. “The worldwide diversity of HIV mutations is extraordinary. Even the diversity within one individual person living with HIV is dramatic,” says LJI Instructor Patrick Madden, Ph.D., who served as study co-first author with Jon Steichen, Ph.D., an institute investigator at Scripps Research.
The third problem is that HIV changes its shape when it infects human cells. Even if B cells get a glimpse of its viral structure—snap!—the structure changes.
Taken together, these problems rarely give B cells a chance to hone their antibody responses against HIV. Even if a B cell manages to make neutralizing antibodies, the virus can mutate or change its shape, rendering those antibodies useless.
The LJI and Scripps Research teams spent years hunting for “broadly neutralizing” antibodies that can actually bind to HIV and recognize key viral structures, even if the rest of the virus mutates. These antibodies are very, very rare, but they can be found in blood samples from a small number of people living with HIV.
An effective HIV vaccine would need to prompt the immune system to make these same broadly neutralizing antibodies. “How could we flip the whole immune response on its head so the rare responses become the common responses? That was a critical challenge we faced,” says Crotty.
Testing the new vaccine
It was time to go back to B cell bootcamp. The scientists studied what made the HIV-fighting B cells special. Then they reversed the process to see exactly how those B cells matured. By looking back at the maturation process, the researchers could track how the B cells changed when they saw specific pieces of the HIV structure.
The team discovered that B cells matured to make broadly neutralizing antibodies after they got an early look at parts of HIV’s outer “envelope” protein. Because these viral sites sparked an immune response, scientists would call them “antigens.”
An effective HIV vaccine would likely need to include models of these antigens. The antigens would work like mugshots of America’s most wanted. If B cells saw those antigens early and often, they would get really good at recognizing and even neutralizing HIV. “We were trying to mimic the progression of those neutralizing antibodies,” says Madden.
In a feat of molecular engineering, the Schief Lab developed vaccine molecules that resembled the real HIV antigens. The scientists then worked with Emory National Primate Research Center, to test this potential HIV vaccine in a non-human primate species called rhesus macaques.
The researchers first administered a “priming” vaccine meant to activate each animal’s naive B cells. The animals then received a series of “shepherding” booster shots to help their B cells develop along the right path.
“This series of vaccinations will guide, or ‘walk’, a B cell from its naive state to its broadly neutralizing state,” says Madden.
This new type of vaccine approach is called “germline targeting” because it targets naive B cells in their “germline” or naive form, before they begin their training process.
The scientists found that around 44 percent of the animals went on to produce broadly neutralizing antibodies against HIV in their blood. These antibodies were impressively abundant.
“We succeeded in taking ultra-rare antibody responses and turning them into common responses by the end of the vaccination process,” adds Crotty. In other research recently published, they reported a new strategy to accelerate related vaccine antibody responses [See Nature Immunology paper].
The team didn’t test whether these antibodies could prevent infection, but it’s significant that these antibodies could be found in the blood, where they could encounter and potentially block HIV.
Bringing the HIV vaccine to humans
The Crotty Lab plans to investigate how they might change the booster shot regimen to make the HIV vaccine even more effective. “It was incredible to get those results, but of course we’d like to see a response in 100 percent of the animals,” says Madden.
Importantly, the antibodies found in the animal subjects resembled the exact kinds of broadly neutralizing antibodies seen in those rare humans who made their own neutralizing antibodies. It’s clear that our immune systems can make these powerful antibodies, given the right training.
“We believe this vaccine approach is even more likely to succeed in humans, because of the immunogenetics,” Crotty says.
The priming immunogen used in this study was evaluated in humans in the HVTN 144 trial and is currently being tested in the Phase 1 trial IAVI G004. IAVI, Scripps Research, the HIV Vaccine Trials Network, and partners are now advancing plans to further evaluate the full immunization regimen in a future human clinical study.
Additional authors of the study, “Vaccination elicits HIV broadly neutralizing antibodies in primates,” include Claudia T. Flynn, Swastik Phulera, Monolina Shil, Oleksandr Kalyuzhniy, Alessia Liguori, Carolyne Kifude, Leigh M. Sewall, Christopher A. Cottrell, Krystal M. Ma, Sabyasachi Baboo, Jolene K. Diedrich, Katherine McKenney, Allan C. deCamp, Diane G. Carnathan, Ivy Phung, Parham Ramezani-Rad, Ester Marina-Zárate, Brian Freeman, Zhenfei Xie, Jeong Hyun Lee, Troy Sincomb, Nicole Phelps, Danny Lu, Diana Goodwin, Ryan Tingle, Yumiko Adachi, Nushin Alavi, Jenny Tran, Andy S. Tran, Alyne Nascimento, Catherine Sovie, Daniel L. V. Bader, Hannah Voic, Xiaoya Zhou, Grace Pixton, Agnes Walsh, Mariane B. Melo, Torben Schiffner, Facundo D. Batista, Dennis R. Burton, Darrell J. Irvine, James C. Paulson, John R. Yates III, Gabriel Ozorowski, Andrew B. Ward, Guido Silvestri.
This work was supported by National Institute of Allergy and Infectious Diseases (NIAID), of the National Institutes of Health, through grant UM1 Al100663 to the Scripps Center for HIV/AIDS Vaccine Immunology and Immunogen Discovery (CHAVI-ID), grant UM1 AI144462 to the Scripps Consortium for HIV/AIDS Vaccine Development (CHAVD), P51 OD011132 to Emory National Primate Research Center, and R01 AI113867; by the Gates Foundation under the Collaboration for AIDS Vaccine Discovery (NAC INV-007522, INV-008813, INV-034657, and INV-064772), via the IAVI Neutralizing Antibody Center (NAC); and by the National Institute of Health grant S10OD025052.
European Citizens’ Initiative
“But I already have a website on Substack,” you argue.
No, no, Substack is just a distribution tool to amplify your website. It should not be your digital home.
In the last few years, I’ve noticed a pattern of writers leaving their websites to make Substack their digital home.
Now, it’s kinda okay if they have bought a domain and linked it to Substack. (Meaning, it’s better than nothing.)
Rachel from Conscious Living is a good example. This way, Substack more or less functions like a content management system (CMS) for you.
However, compared to other CMS it’s very limited, such as the ability to manage your SEO and customize your pages to add more features, but I digress. If you just want a fuss-free platform, this is one way to get it and Substack’s conditions for domains are very reasonable and cost-efficient. As I will explain later, this could change on a dime without warning.
However, there are some writers who are saying: “Hey readers, I’m now writing on Substack, so head on over there (and ignore my website)!”
Some writers do have a website, but link to their Substacks, calling them their “blogs”. If your Substack has a domain name they own, it’s okay, but if it’s xx.substack.com, Substack is saying “All your content are belong to us”.
In conclusion: Writers, don’t do this. It’s short-sighted and unwise and can derail your long-term visibility on the Internet.
The siren call of convenience
Every few years, the internet convinces writers that a new digital paradise has arrived. First, it was social media like Facebook. Then blogging networks like Tumblr. Then it was Medium. More recently, it’s been Substack.
Platforms promise us an eager audience, built-in monetization, a smooth user interface, and a supportive community. As a writer who just wants to focus on writing, it’s incredibly tempting to hand over the keys to our creative kingdoms and let these portals handle everything. (Believe me, I gave in at one point. For years, I just stopped blogging altogether and even gave up a domain that had high traffic! But I got back in 2012 and never left.)
However, this is the truth that has not changed since the dawn of the Internet: When you build your audience entirely on someone else’s platform, you aren’t a homeowner. You are a tenant. Or worse, a digital sharecropper.
And corporate landlords always change the rules eventually. It’s not personal, it’s just business.
The illusion of the safe space
It’s easy to feel secure when a platform is in its golden era. But we’ve watched the downfalls of Twitter, the policy shifts of Reddit, and the changing tides of algorithmic networks. Relying blindly on a centralized portal not owned by you means your life’s work can alter overnight based entirely on a corporate boardroom decision.
When I looked at how fragile our digital ecosystems really are, I realized I needed a space that wouldn’t go “poof” because a company needed to please its investors or shareholders. This realization completely changed my approach, pushing me to protect my content by learning to blog the IndieWeb way.
Your writing needs a permanent homebase—a domain that you own and control. Full stop.
Moving from “renting” to syndicating
The biggest pushback I hear from writers is: “But my website doesn’t have an audience! Substack does.”
But you don’t have to completely abandon social media or platforms like Substack to protect your autonomy (personally, I prefer the word sovereignty but it does sound a tad dramatic).
You just need to change the order of operations. Instead of publishing directly to a portal, you can shift your mindset to POSSE: Publish (on your) Own Site, Syndicate Elsewhere. (I explain the POSSE/PESOS method in an older post.)
By treating your website as the definitive source of truth and using platforms simply as distribution pipes, you get the best of both worlds. I dug deep into this shift when I committed to being an imperfect gardener of my digital garden, exploring how a less market-y way of presenting my content online let me share my wild garden of thoughts without dancing to the algorithm.
A reality check on platform hype
If you are still holding out hope that Substack is “different” from the social platforms that came before it, let’s look at the numbers and behaviors behind the marketing copy.
After spending a significant amount of time observing the platform ecosystem firsthand, I wrote a brutally honest takeaway in What I learned from one year of Substack. The network effects are real, but so is the pressure to conform to what the platform’s ecosystem favors.
This post, by the way, desperately needs to be updated because things have gotten much, much worse since I wrote it.
When you hand your content over to a platform, you have to conform to their rules and their localized biases. For those of us writing from outside the dominant US-centric echo chambers, platform algorithms heavily prioritize specific western narratives, making it incredibly tough for localized or minority voices to be seen unless they conform.
I wrote about this exact frustration recently in Linkblog: Dwelling on the Internet, highlighting how algorithmic complacency forces us into homogenized bubbles.
The flip side — the writers who refused to leave their websites
Each time there’s a new drama on some platform, and writers are shaking their sabers and declaring that they will leave for yet another social media platform they don’t control, I think about writers like John Scalzi.
As of date, John scalzi has been blogging on https://whatever.scalzi.com/ for 28 years!
This sci-fi novelist has maintained a single independent website continuously for nearly three decades; this makes him one of the longest-running, most consistent original bloggers on the internet. Imagine the amount of digital footprint on that website! Unbroken by time or platforms.
(Specifically, he uses wordpress.com like I do, as we both don’t want to bother with the pain of setting up your own self-hosted wordpress website and just want the folks at Automattic to do it.)
He blogs in the classic Indieweb way, though I doubt he is even aware he’s doing it. He treats his social media channels such as X or Bluesky as a way to amplify his website. All roads lead back to https://whatever.scalzi.com/, and this is something I wish every single writer would do.
He wrote recently in Various & Sundry, 6/3/26:
this site acts as my own institutional memory, if I post something about it here it constitutes an official record. I mean, all the posts I ever placed on the former Twitter are now entirely lost to time, since I have gone in and purged my entire timeline there. This site, however, endures. — John Scalzi
this site acts as my own institutional memory, if I post something about it here it constitutes an official record. I mean, all the posts I ever placed on the former Twitter are now entirely lost to time, since I have gone in and purged my entire timeline there. This site, however, endures. — John Scalzi
Breaking free from platform blues
Trying to adapt your presence across various platforms in an ever-shifting digital landscape is exhausting. One minute a platform is a writer’s darling; the next, it’s being boycotted. Railing against a platform’s focus shift or the presence of (long sigh) Nazis is a useless endeavor.
As I noted in Linkblog March 12, 2026: Platform blues, chasing platform purity is an illusion. Tech will change, corporate algorithms will continue to prioritize profit over human connection, and platforms will continue to cycle through hype and decline.
The antidote to this exhaustion isn’t moving to the next shiny new app. It’s anchoring your work on an independent website with open distribution channels like RSS. It also means ruthlessly using platforms as distribution channels. When one collapses or you prefer to just move, it’s easy to just change strategies because your digital home remains unchanged.
Use platforms to find your readers, but bring them back to your house. It’s time to stop digital sharecropping on rented land.
Featured photo is by vivek vk on Unsplash
The story of MIT’s most notorious milk carton begins, as many good stories do, in a college dorm.
The milk in question was purchased in 1994 and rediscovered in 1995 by an undergrad named Justin Cave. By that point, the reportedly lactose-intolerant Cave had even less use for the milk he had abandoned in his fridge ten months earlier. For reasons lost to history, he did not throw the milk away. He threw it a birthday party.
The Milk lived the rest of its life unrefrigerated, stored in a tall, single-walled jar. For twenty-seven years, the residents of Random Hall dorm gathered faithfully to celebrate its birthday. At age 20, the Milk applied to, and was rejected from, MIT1. The jar was periodically “burped” to release the gas pressure inside, until the Milk reached its stable final form — a cloudy brown liquid. When asked why the Milk was never thrown away, one resident of Random Hall replied: “Why throw something away when you can tell a story about it?”
“Stuff I learned from things that normally get thrown away” could be the title of many scientists’ memoirs, including Louis Pasteur’s. Winemaking produces, well, wine, but it also produces acidic crystals on the walls of the vats. These byproducts were not discarded — they were studied by Pasteur and his contemporaries. Pasteur’s observations both revolutionized our understanding of chemistry and led him to the phenomenon that would define his career and lay the foundation of modern food safety: fermentation.
The microorganisms responsible for fermentation are visible to our naked senses only through the textures, colors and smells resulting from their collective efforts. From Aristotle through the 1850s, it was assumed that some intrinsic property of a non-living starting substance (like grain or milk) enabled its spontaneous fermentation into something useful (like beer or yogurt) or its eventual spoilage.
It was Pasteur who proved that living organisms were required for the transformations that took place during fermentation. He heated up nutrient-rich broths in custom flasks that let gases, but not microbes, flow in and out of the flasks. Pasteur then broke the neck off of one of the flasks to expose the broth to the air. If the boiled broth could spontaneously transform, it would do so with or without exposure to microbes in the air and environment.
The flask with the neck broken off grew cloudy and fermented as bacteria bloomed, but the sterile one remained clear.
This finding was great news for Napoleon III2. The French were losing money, and perhaps more alarmingly, their reputation, exporting wine to the British — the wine would “spontaneously” go bad during shipping. The French government offered a prize for a scientist to solve the case of the spoiled wine. With the knowledge of the microorganism-driven process of fermentation in hand, Pasteur did to the wine what he did to the broth, just more gently — he heated the wine enough to kill microbes without damaging the wine’s flavor. Immortalized as pasteurization, this process was adapted shortly after its invention in 1865 to let us safely drink stored milk3.
Killing bacteria thus became a major preoccupation of modern life. Its most visible manifestation today might be taking antibiotics (first available in the 1940s): there were ~700 antibiotic prescriptions per 1000 people4 in 2024 according to CDC data. A close second might be the dizzying array of disinfectant products found in U.S. grocery stores.
What’s less visible is the sanitization infrastructure that makes things like grocery stores or medicine possible at all. The company Steris, one maker of high temperature, pressurized sterilization equipment and other medical instruments, is a $5 billion annual revenue company, with a $21 billion market cap. The U.S. pasteurizes around 50 billion liters of fluid milk every year. To package salad greens like spinach, the greens are washed in a dilute bleach solution to kill any lingering soil microbes. I could go on.
But our war on bacteria has its own warring industry. This industry has captured the imaginations of scientists, the food and beverage industry, pharma companies and doctors along with influencers, marketing gurus and opportunists of all flavors. This industry emphasizes that some microbes are friends, not foe, (true) and you should be eating them in large quantities, on purpose, all the time, and preferably paying more for products that contain them (dubious). This is the probiotics industry.
“Probiotic” is a bit of a misnomer — it means for life, or promoting life, but the formal definition of a probiotic is an actual living microorganism. In simple terms: taking a probiotic is just eating bacteria on purpose. I say on purpose because we consume microbes accidentally all the time from our environment, largely oblivious to their existence or effects. The bacteria we spend much of our time and energy trying to kill are outnumbered, at a species level, at least 1000 to 1 by a combination of harmless and beneficial bacteria living in and on our bodies. It’s this latter property of beneficialness that probiotics are trying to exploit.
I say exploit because of a recent trip I took to the grocery store. I had a cold and was in search of lemon ginger tea. I bought a box of Bigelow, went home, boiled some water, poured it over a tea bag, waited a bit, added honey, took a sip, and almost spit it out. The tea had its expected notes of ginger, a hint of lemon, and some powdery, alkaline aftertaste that I couldn’t place. Frankly, it tasted terrible. (Sorry, Bigelow).
I inspected the box again. In my congested state, I had unwittingly purchased a new offering from the tea company — Bigelow Lemon Ginger, with probiotics. What made this tea different from all the other teas I happily sipped on was that in addition to nice-sounding things like lemongrass and cinnamon, it contained bacteria. Bacteria which I had just boiled, at a temperature 40oC hotter than pasteurization.
Did the tea taste bad because I was drinking dead bacteria water? And if that was the ultimate outcome of the normal brewing process, why bother putting bacteria in the tea at all?
I was at a lab happy hour when I mentioned this to my PhD thesis advisor. “I know, right?” she said, suddenly animated. “Probiotic teas taste SO BAD.” I was thrilled to have another witness. “Doesn’t it seem crazy to add in bacteria that you’re just going to boil and kill anyway?” I asked. “Is it all a scam?” She was already nodding. “You have to wonder whether the bacteria in the tea make it to the gut at all, and whether they do anything helpful once they get there,” she said.
I’d be lying if I said I remembered exactly what happened next, or who suggested what. All I remember is an idea. An idea to test this seemingly paradoxical marketing tactic like the microbiologists we are. The idea was simple: What if we tried to grow the bacteria from the tea bag, in the lab?
I went home. I stared at the box of bacteria tea.
Why throw something away when you can tell a story about it?
BC30™, the bacterial strain in the tea, is short for Bacillus coagulans GBI-30, 6086®. It received the FDA’s GRAS (Generally Recognized as Safe5) designation in 2012 and is found in over a thousand “leading food, beverage and pet food products worldwide” according to the probiotic’s website.
To coax these bacteria to grow out of steeped tea, I needed to know three things:
Is this species safe to grow in the lab?
Is this species safe to grow in the lab?
What does it like to eat?
What does it like to eat?
What are its preferred growth conditions?
What are its preferred growth conditions?
In general, I try not to ingest the bacteria I grow in the lab — even ones with the lowest safety designation, BSL-1. By nature of it being a commercial probiotic, BC30 is both BSL-1 (safe to grow under normal lab precautions) and edible.
But, I still wouldn’t try this at home or eat bacteria off of a culture plate. Why? BC30’s preferred food source is not that different from the preferred food source of many other microorganisms: a sugar- and amino acid-rich nutrient medium called MRS (De Man, Rogosa and Sharpe) agar.
While MRS agar has some adjustments to make it preferentially appetizing to BC30 and its relatives, those relatives also include Streptococcus pyogenes (causes strep throat) and Bacillus cereus (causes food poisoning). Without sterile technique and rigorous species-level confirmation, you cannot know for sure what is growing on your plate.
With that said, the American Society for Microbiology’s blog suggested that were I successful in culturing BC30, I would see growth of translucent white colonies on MRS agar plates after 48 hours of incubation at 30 – 33oC in the presence of oxygen.
First, I needed to make tea.
I wanted the conditions of the experiment to represent a range of realistic tea-drinking scenarios, from intended use to flagrant improvisation, and set up three steeps:
The Rule Follower — Brewed as directed for 4 minutes in boiling water.
The Rule Follower — Brewed as directed for 4 minutes in boiling water.
“I forgot I made tea” — We’ve all been there. 15 minutes, boiling water.
“I forgot I made tea” — We’ve all been there. 15 minutes, boiling water.
Cold brew anarchist — Self-explanatory.
Cold brew anarchist — Self-explanatory.
It was at this point I realized I needed a sterile-ish way to transport the steeped tea and tea bags from my house to the lab. Luckily, I had recently run a blindfolded volume pouring accuracy competition at our departmental retreat and had leftover Falcon tubes still in their original package. While the tea was definitely not sterile, I reasoned that a little extra aseptic technique wouldn’t hurt. I poured the tea into the tubes over my kitchen stove, using the open flame as a makeshift Bunsen burner.
In reality, lab came first. I had to make MRS agar plates before I steeped the tea. Our lab does not use MRS broth very often, and when I first looked for some all I found was a 10-year-old solidified block of MRS powder in our stock cabinet that was growing large green spots inside of its glass container. Behind it was one that looked mercifully normal.
I mixed broth powder, agar and water in a glass bottle, loosely capped it, put it in a water bath and took it to the autoclave. “Autoclave” is a nice word for giant pressure cooker. Ours is made by the aforementioned Steris. It rattled and hissed as its jaws opened to accept my tray of culture media, which it then heated to 121oC for 45 minutes, sterilizing the liquid.
Back at my lab bench, when the molten MRS agar had cooled enough to handle, I lit a Bunsen burner next to a stack of empty plastic petri dishes and poured a layer of agar into each one. Left overnight, the plates solidified into nutrient-dense beds for BC30.
The next day, I took my tubes of tea to lab. I pipetted 400 microliters (0.4mL) of each liquid tea condition onto a plate next to the Bunsen burner. I spread the liquid evenly across the plate with a hockey stick-shaped plastic spreader6 and left the lids on the plates cracked open to dry near the flame.
But to answer my question, I needed one more test. If there were bacteria in the tea bag initially, but they died when boiled, then I might see bacterial growth by plating the dry ingredients of an unsteeped tea bag, or the tea bag steeped in cold water. I cut open the tea bags and shook some of their contents onto the agar. I put my full set of plates, including a plain MRS plate to check its sterility, into the incubator at 37oC — a standard growth temperature, but a little warmer than recommended. I was skeptical that anything would grow. For the next two days, all I could do was wait and see.
The first thing I noticed when I took the plates out of the incubator was the smell. I was in disbelief when I saw little white colonies dotting almost all the plates and opened one to get a closer look. A sickly sweet, gingery aroma wafted from the plate as I inspected the translucent colonies — a byproduct of the bacteria metabolizing the sugars in the MRS plate. By all accounts, I was looking at BC30.
Colonies grew on all of the tea and tea bag plates, while my sterile control plate remained bacteria-free. The colonies from the boiling-water steeps and the tea bags were a variety of sizes, including some that were significantly larger than others, while the colonies from the cold-water tea were uniformly small.
Because I knew the volume of tea I had put on each plate, I could calculate a standard measurement of bacterial density: colony-forming units (CFUs) per milliliter. Contradictory to my expectations, I saw a five-fold increase in colonies from the tea steeped in boiling water relative to the tea steeped in cold water for the four-minute condition. I saw the same pattern in the fifteen-minute condition, with a nearly four-fold increase in boiling vs cold.
Before I could draw any conclusions, I needed to know, for sure, that these colonies were Bacillus coagulans. The most robust way to check is by sequencing their DNA, but sequencing is expensive. A simpler, cheaper way to check is with PCR, which amplifies small regions of DNA unique to a species. I downloaded the BC30 genome and selected two regions of its genome that didn’t match other species in the NCBI database. Using Primer3, I generated two pairs of PCR primers, short stretches of DNA to bind to either side of my region of interest.
I picked the largest colony I could see from each plate (seven total), suspended the cells in a small volume of water, and set up standard colony PCR reactions. The heat during the reaction bursts the cells, making the DNA available for amplification. The completed reaction was run through a porous gel with an electric current and visualized with UV. If I saw bands on the gel for both primer sets, from totally different parts of the BC30 genome, I could be confident that this was, in fact, BC30.
I loaded the gel into the imager and hit run. There, in black relief against the grey background of the gel, were my bands.
Bigelow knew something I didn’t7. It turns out that BC30, like many of its relatives, is a spore-forming bacterium. When starved of nutrients, Bacillus coagulans divides asymmetrically, packing its basic cellular information into a spore with a thick protective coat. These spores are resistant to dessication, nutrient starvation, radiation, chemical disinfectants and extreme heat. It was these spores that were in the tea bag — spores that are perfectly comfortable being steeped in boiling water.
Like the seeds of plants, when the spores find themselves in favorable conditions for growth — say, on an MRS agar plate at a balmy 37oC — they germinate back into actively growing cells. This is the premise of their ability to function as a probiotic. The spores are dormant and shelf-stable in a tea bag, or any of the thousand products advertised to contain BC30, and will, in theory, germinate upon arrival in the GI tract, where they can exert some sort of effect on the host that consumed them.
To produce spores at scale, manufacturers grow bacteria in vats of nutrient-rich broth. If the nutrients are not replenished, the bacteria eventually start to starve, triggering the sporulation process. Around 24 hours later, the bacterial broth is treated with enzymes to kill any remaining, non-sporulated cells. The mixture is concentrated, washed with water, and finally, in a fantastic twist of irony, pasteurized.
The pitch for BC30 is that it improves “digestive health” and “protein absorption.” The reported endpoints for digestive health on BC30’s website are reductions in bowel movement frequency, abdominal pain and abdominal bloating in adults with IBS. In the study promoted on the site, the baseline for the placebo group for abdominal pain and bloating is, mysteriously and respectively, 12.5% and 30% higher than the baseline for the BC30 treatment group. The placebo group experienced no change in severity scores over the subsequent course of treatment, while the BC30 group dropped to placebo levels after a week and stabilized.
For one of the protein absorption studies, there is a small but statistically significant difference in amino acid levels in the blood, including when BC30 is paired with another one of its parent company’s products, a “nutritional milk protein concentrate” called Ultranor.
If these results hold, they beg the question — could a product like Bigelow’s probiotic tea be able to produce these beneficial effects? Most of the clinical trials I could find, including the IBS study above, dosed people daily over the course of one to eight weeks with 1 billion CFUs (spores) of BC30. Per my calculations, a properly steeped cup of probiotic tea yields around 30,000 CFUs: 0.003% of the clinically tested dose.
Granted, I am one person and this is one experiment. But there are independent, conflicting reports on whether BC30 survives the GI tract at all. One study suggests that Bacillus probiotics don’t make it, while another reports about half of the initial dose of spores surviving transit through an artificial human gut system. Other research suggests that the effect of the probiotic is not even due to the cells coming back to life, but due to an immune response against the dormant or vegetative cells.
These observations have consequences for consumers being parted from their money by unsubstantiated health claims. But they are interesting observations in their own right. Sporulating organisms’ imperviousness to heat, while useful for commercial biotech applications, causes problems for the food industry. The food-poisoning agent Bacillus cereus is a species normally found in the soil. It can release heat-resistant toxins if it multiplies in food, and live bacteria can produce toxins when they reach the small intestine. Even pasteurized milk spoils eventually as heat-resistant spores, mostly soil Bacillus, begin to multiply.
Bacillus coagulans is also a soil bacterium by nature8, and not a typical resident of the community of microbes in our gut (called the gut microbiome). It was discovered in 1915, in canned milk that had spoiled and coagulated. In spite of its origins, BC30 seems inert as a pathogen, and reports of it causing infection are vanishingly rare. BC30’s safety track record is remarkable.
But safety is only the first step on the quest to use probiotics for good. New probiotic companies like Pendulum and Seed market the fact that they are backed by clinical trial data — Pendulum for blood sugar control in Type II diabetes, and Seed for gas, bloating and regularity in healthy adults. The backbones of Pendulum and Seed’s products are organisms found more commonly in the gut, and, interestingly, both companies focus on multi-species products, dosing patients with miniature microbial communities.
One focus of my PhD lab is on abnormal pathogenic behavior of normally harmless bacterial residents of the gut, most commonly in people who are already quite sick. While unhappy microbiomes can be unhappy in their own way, we do not have a consensus on what a “healthy” gut microbiome looks like, either. In collaboration with a continent-wide consortium in Africa, our lab helped catalogue the species found in healthy adult women across the continent. We found over 1,000 new species relative to what had been previously described in studies focused on Western countries.
Companies trying to introduce targeted combinations of microbes into the gut are thus forever shooting at a moving target. Outside of specific indications for GI infections, determining whether to give (or take) a probiotic is a grey area. And for the common GI complaints focused on by the probiotic market, targeting the microbiome with additional organisms may not be the answer at all. Rather, by understanding how bacteria work together in the microbiome, solutions may favor changing the metabolic environment of the gut to drive the formation of species-agnostic “guilds” that perform specific functions, likely via dietary interventions.
I never get tired of growing bacteria. For a colony to be visible on a plate, it consists of at least a million, often closer to a billion, individual cells. Learning how bacteria grow and adapt does not diminish the sense of wonder I feel when I observe them — it only enhances it. I like to think this same sense of wonder animated the scientist who first cultured Bacillus coagulans out of canned milk that had spoiled. And I have to imagine some mixture of wonder, awe and horror kept the Random Hall Milk alive for twenty-seven years.
The next Louis Pasteur could be a lactose-intolerant undergrad, or a procrastinating PhD student. It could be you. Pausing to look a little longer, to ask why the world is the way it is — this is how we upend assumptions of what is valuable. What is worth looking at. Because in the end, trash is in the eye of the beholder.
1
You can read the Milk’s application here.
2
Author correction 7/28/26: Changed from “Napoleon” to “Napoleon III.” A commenter on another site correctly pointed out that Napoleon I was exiled before Pasteur was born. According to the Pasteur Institute, “in 1863, Napoleon III asked [Pasteur] to study diseases in wine.”
3
As demonstrated by the Milk, even pasteurized beverages spoil, a process sped up by exposure to the microbes in the air but which will proceed within an unopened container anyway. How is this possible?
It’s because pasteurized milk is not the same as sterilized milk. Pasteurization heats to ~60C for a few minutes. While the microbes that we worry about causing infection can’t survive this, some heat-tolerant bacteria and proteins can — those are what will eventually break down the milk, even if it isn’t opened to the air. Heating to 140C, on the other hand, makes milk effectively sterile, killing even the heat-tolerant bacteria. But this process, used to produce “ultra-high temperature” or UHT shelf-stable milk, does some odd things to the proteins that subtly change the flavor, color, and texture of the milk.
4
Author correction 7/28/26: I reported this initially as “7 in 10 people were prescribed antibiotics,” but a commenter on another site pointed out that the original report data likely reflects a smaller number of people getting repeat prescriptions, so I have reported the raw data from the CDC report instead.
5
If a probiotic is marketed as a food or dietary supplement, as most are, then it is not required to undergo a clinical trial in the U.S. but instead to submit a GRAS notification. The second most important thing to know about the GRAS system is that it does not require proof of efficacy — only safety. The most important thing to know is that the proof of safety is provided by the company requesting the GRAS designation. This proof is then reviewed by the FDA, to determine whether the notice provides a “sufficient basis for a GRAS determination” and whether “information in the notice or otherwise available to FDA” raises any safety concerns.
6
This is one of the most polarizing choices one can make as a biomedical research scientist. The alternative to the hockey stick is to use glass beads that you autoclave then sprinkle on the plate and roll around. People are very passionate about their chosen method and will attempt to convert you.
7
Saw this weirdly aggressive Bigelow commercial at the gym. I don’t think they’re going to sponsor me after this article.
8
As our understanding of microorganisms evolves, so do our naming and classification conventions. Bacillus coagulans is a more distant relative of Bacillus cereus and similar soil microbes than previously thought and has been re-classified into a new genus called Weizmannia. Its full nomenclature history can be found on the LPSN.
@openai/codex-security is a CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities in your code. Scan repositories, review changes, track findings over time, and run security checks in CI.
Documentation
Quick start
Requires Node.js 22 or later, Python 3.10 or later, and access to Codex Security.
npm install @openai/codex-security npx codex-security login npx codex-security scan .
For CI, set OPENAI_API_KEY instead of signing in.
If both a ChatGPT sign-in and an API key are available, interactive scans ask which credential to use. CI and other noninteractive scans keep the existing API-key precedence. Select a credential explicitly when needed:
npx codex-security scan . –auth chatgpt npx codex-security scan . –auth api-key
To make your ChatGPT sign-in the automatic default, unset any configured API keys:
unset OPENAI_API_KEY CODEX_API_KEY
Scan history is stored in the Codex Security workbench state directory. If that directory cannot be written, set CODEX_SECURITY_STATE_DIR to a writable directory outside the repository.
TypeScript SDK
import { CodexSecurity } from “@openai/codex-security”;
const security = new CodexSecurity(); const result = await security.run(”.“);
console.log(result.reportPath); await security.close();
For installation, authentication, scan options, and CI setup, see the official documentation.
The Kimi K3 architecture figure for yesterday’s big open-weight model release, along with some observations and thoughts.
Yes, it looks relatively complicated, but it’s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
Yes, it looks relatively complicated, but it’s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)
The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it’s already very crowded, but that’s essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it’s already very crowded, but that’s essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
Kimi K3’s overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
Kimi K3’s overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
Kimi K3 now also has native multimodal support, which is great!
Kimi K3 now also has native multimodal support, which is great!
There are several other interesting training tidbits in the technical report, but that’s it from the architecture front so far. A really great release overall.
Source: website version of my Substack note.
Read Next
A Few Notable Open-Weight Models This Week
Short note on the architectures of six new open-weight models, including Nanbeige 4.2, Laguna S 2.1, Motif-3-Beta, Solar Open 2, Antares 1B, and BTL-3.
Correction for Listing 6.5 in Build a Reasoning Model From Scratch
Short correction note for the random seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch.
Inkling: A New Open-Weight 975B MoE with a Few Surprises
Short note on Thinking Machines Lab’s 975B Inkling model, including benchmarks, sparse MoE design, short convolutions, RMSNorm, and position bias.
Authors:Kimi Team: Yu Zhang, Zongyu Lin, Xingcheng Yao, Jiaxi Hu, Fanqing Meng, Chengyin Liu, Xin Men, Songlin Yang, Zhiyuan Li, Wentao Li, Enzhe Lu, Weizhou Liu, Yanru Chen, Weixin Xu, Longhui Yu, Yejie Wang, Yu Fan, Longguang Zhong, Enming Yuan, Dehao Zhang, Yizhi Zhang, T.Y. Liu, Haiming Wang, Shengjun Fang, Weiran He, Shaowei Liu, Yiwei Li, Jianlin Su, Jiezhong Qiu, Bo Pang, Junjie Yan, Zhejun Jiang, Weixiao Huang, Bohong Yin, Jiacheng You, Chu Wei, Zhengtao Wang, Chao Hong, Yutian Chen, Guanduo Chen, Yucheng Wang, Huabin Zheng, Feng Wang, Yibo Liu, Mengnan Dong, Zheng Zhang, Siyuan Pan, Wenhao Wu, Yuhao Wu, Longyu Guan, Jiawen Tao, Guohong Fu, Xinran Xu, Yuzhi Wang, Guokun Lai, Yuxin Wu, Xinyu Zhou, Zhilin Yang, Yulun Du
View PDF
Abstract:We introduce Kimi Linear, a hybrid linear attention architecture that, for the first time, outperforms full attention under fair comparisons across various scenarios — including short-context, long-context, and reinforcement learning (RL) scaling regimes. At its core lies Kimi Delta Attention (KDA), an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating mechanism, enabling more effective use of limited finite-state RNN memory. Our bespoke chunkwise algorithm achieves high hardware efficiency through a specialized variant of the Diagonal-Plus-Low-Rank (DPLR) transition matrices, which substantially reduces computation compared to the general DPLR formulation while remaining more consistent with the classical delta rule. We pretrain a Kimi Linear model with 3B activated parameters and 48B total parameters, based on a layerwise hybrid of KDA and Multi-Head Latent Attention (MLA). Our experiments show that with an identical training recipe, Kimi Linear outperforms full MLA with a sizeable margin across all evaluated tasks, while reducing KV cache usage by up to 75% and achieving up to 6 times decoding throughput for a 1M context. These results demonstrate that Kimi Linear can be a drop-in replacement for full attention architectures with superior performance and efficiency, including tasks with longer input and output lengths. To support further research, we open-source the KDA kernel and vLLM implementations, and release the pre-trained and instruction-tuned model checkpoints.
arXiv-issued DOI via DataCite
Submission history
From: Yulun Du [view email] [v1] Thu, 30 Oct 2025 16:59:43 UTC (645 KB) [v2] Sat, 1 Nov 2025 12:05:18 UTC (691 KB)
A note on notation: this article defaults to bra-ket notation because (in my quantum-inspired opinion) it makes the shapes in this derivation very clear. The Math notation switch above rewrites every equation using conventional bold vectors and explicit transposes instead. In bra-ket mode, ∣q⟩\lvert q\rangle is a column vector, ⟨k∣\langle k\rvert is a row vector, ⟨k∣q⟩\langle k\rvert q\rangle is a number, and ∣v⟩⟨k∣\lvert v\rangle\langle k\rvert is a matrix. Vectors face right by default, while keys face left when written into the linear-attention state. We work with one causal attention head and real-valued vectors, assume DeltaNet’s keys are normalized, and let the state map from key space to value space.
Modern linear attention variants are complex, and a upon first glance it is not so easy to see what they are designed to achieve. For reference here is the state update equation for Kimi Delta Attention (KDA):
S~t=St−1Diag(αt)\widetilde S_t = S_{t-1}\operatorname{Diag}(\alpha_t) ∣v^t⟩=S~t∣kt⟩\lvert\widehat v_t\rangle = \widetilde S_t\lvert k_t\rangle ∣et⟩=βt(∣vt⟩−∣v^t⟩)\lvert e_t\rangle = \beta_t \left( \lvert v_t\rangle-\lvert\widehat v_t\rangle \right) St=S~t+∣et⟩⟨kt∣S_t = \widetilde S_t+\lvert e_t\rangle\langle k_t\rvert ∣ot⟩=St(dk−1/2∣qt⟩)\lvert o_t\rangle = S_t\left(d_k^{-1/2}\lvert q_t\rangle\right)
The reason they are so difficult to understand is that this is the latest in a family of linear attention variants that have been developed over the last few years and the complexity of them has inevitably ballooned such that from the outside the latest variants appear inaccessible.
In this post we are going to walk through the DeltaNet family of linear attention variants, two of which are used by the latest Qwen and Kimi model families, and show how you might have arrived at the same equations by asserting simple things about your hidden state.
That is the route we will take:
softmax attention → linear attention → DeltaNet → Gated DeltaNet → KDA
Only after deriving KDA will we turn to the recurrent and chunkwise Triton programs that execute it.
1. Begin with quadratic attention
For a query at token tt, ordinary causal softmax attention is
ati=exp (s⟨ki∣qt⟩)∑j≤texp (s⟨kj∣qt⟩),s=dk−1/2,∣ot⟩=∑i≤tati∣vi⟩.\begin{aligned} a_{ti} &= \frac{ \exp\!\left(s\langle k_i\rvert q_t\rangle\right) }{ \sum_{j\leq t} \exp\!\left(s\langle k_j\rvert q_t\rangle\right) }, \qquad s=d_k^{-1/2},\\ \lvert o_t\rangle &= \sum_{i\leq t}a_{ti}\lvert v_i\rangle. \end{aligned}
Every attention weight is a scalar. It measures the similarity between one key and one query, then softmax turns all of the scores for that query into a distribution. The output is a weighted sum of value vectors.
Over a sequence of length TT, there are T2T^2 key-query pairs. During autoregressive inference we can cache the keys and values instead of recomputing them, but the cache still grows with the sequence and every new query still has to inspect the entire history.
The obstacle to rearranging this computation is the softmax. Its denominator depends jointly on the current query and every earlier key. So, for the moment, remove it.
1.1 Remove the softmax
For clarity, absorb the constant scale ss into the query. The deliberately bare version of attention is then
∣ot⟩=∑i≤t⟨ki∣qt⟩∣vi⟩.\lvert o_t\rangle = \sum_{i\leq t} \langle k_i\rvert q_t\rangle \lvert v_i\rangle.
The scalar inner product can move to the right:
∣ot⟩=∑i≤t∣vi⟩⟨ki∣qt⟩=(∑i≤t∣vi⟩⟨ki∣)∣qt⟩.\begin{aligned} \lvert o_t\rangle &= \sum_{i\leq t} \lvert v_i\rangle \langle k_i\rvert q_t\rangle\\ &= \left( \sum_{i\leq t} \lvert v_i\rangle\langle k_i\rvert \right) \lvert q_t\rangle. \end{aligned}
Everything that depends on the past can now be collected into one matrix of a fixed size V×KV \times K:
St=∑i≤t∣vi⟩⟨ki∣\boxed{ S_t = \sum_{i\leq t} \lvert v_i\rangle\langle k_i\rvert }
and attention becomes a recurrent write followed by a read:
St=St−1+∣vt⟩⟨kt∣,∣ot⟩=St∣qt⟩.\boxed{ \begin{aligned} S_t &= S_{t-1} + \lvert v_t\rangle\langle k_t\rvert,\\ \lvert o_t\rangle &= S_t\lvert q_t\rangle. \end{aligned} }
The identity
(∣v⟩⟨k∣)∣q⟩=⟨k∣q⟩∣v⟩\left(\lvert v\rangle\langle k\rvert\right)\lvert q\rangle = \langle k\rvert q\rangle\lvert v\rangle
is the whole trick. The outer product is a matrix; the inner product is a number. We no longer store every past key and value. We store their summed outer products in the fixed-size state StS_t.
This is linear in sequence length rather than quadratic: scan the tokens once, updating the same dv×dkd_v\times d_k state at every step. We have paid for that efficiency by discarding softmax’s normalization and selectivity. More sophisticated linear-attention methods use feature maps and normalizers, but this unadorned form exposes the memory problem that motivates DeltaNet.
1.2 Addition is not assignment
Suppose we write a pair ∣vt⟩⟨kt∣\lvert v_t\rangle\langle k_t\rvert and immediately query the new state with that same key:
St∣kt⟩=(St−1+∣vt⟩⟨kt∣)∣kt⟩=St−1∣kt⟩+∣vt⟩⟨kt∣kt⟩⏟1=St−1∣kt⟩+∣vt⟩.\begin{aligned} S_t\lvert k_t\rangle &= \left( S_{t-1} + \lvert v_t\rangle\langle k_t\rvert \right) \lvert k_t\rangle\\ &= S_{t-1}\lvert k_t\rangle + \lvert v_t\rangle \underbrace{\langle k_t\rvert k_t\rangle}_{1}\\ &= S_{t-1}\lvert k_t\rangle+\lvert v_t\rangle. \end{aligned}
The write does not make the memory return ∣vt⟩\lvert v_t\rangle. It adds ∣vt⟩\lvert v_t\rangle to whatever the memory already returned.
If the old state already produced the correct value, the additive write makes the new state produce twice that value. More generally, keys are not mutually orthogonal, so every write can interfere with previous writes. Linear attention has given us a compact associative memory, but its update behaves like += when what we want is closer to =.
2. DeltaNet: write the error, not the value
DeltaNet replaces the unconditional linear-attention write with a delta-rule correction. There are two useful ways to derive it.
2.1 Derivation one: demand that the write can be read back
Before writing token tt, ask the memory what it currently associates with the new key:
∣v^t⟩=St−1∣kt⟩.\lvert\widehat v_t\rangle = S_{t-1}\lvert k_t\rangle.
If we want the memory to return ∣vt⟩\lvert v_t\rangle, we should not add the whole value. We should add only the difference:
∣vt⟩−∣v^t⟩.\lvert v_t\rangle-\lvert\widehat v_t\rangle.
Introduce a learned write strength βt∈[0,1]\beta_t\in[0,1] and define
∣et⟩=βt(∣vt⟩−St−1∣kt⟩).\lvert e_t\rangle = \beta_t \left( \lvert v_t\rangle - S_{t-1}\lvert k_t\rangle \right).
Then write this error at the current key:
St=St−1+∣et⟩⟨kt∣.\boxed{ S_t = S_{t-1} + \lvert e_t\rangle\langle k_t\rvert. }
Now immediately read the same key:
St∣kt⟩=St−1∣kt⟩+∣et⟩⟨kt∣kt⟩=(1−βt)St−1∣kt⟩+βt∣vt⟩.\begin{aligned} S_t\lvert k_t\rangle &= S_{t-1}\lvert k_t\rangle + \lvert e_t\rangle \langle k_t\rvert k_t\rangle\\ &= (1-\beta_t)S_{t-1}\lvert k_t\rangle + \beta_t\lvert v_t\rangle. \end{aligned}
When βt=1\beta_t=1, the result is exactly ∣vt⟩\lvert v_t\rangle. Smaller βt\beta_t moves the old prediction partway towards the target.
The correction is also local in key space. For any query ∣x⟩\lvert x\rangle orthogonal to the current key,
⟨kt∣x⟩=0⟹(St−St−1)∣x⟩=∣et⟩⟨kt∣x⟩⏟0=0.\langle k_t\rvert x\rangle=0 \quad\Longrightarrow\quad (S_t-S_{t-1})\lvert x\rangle = \lvert e_t\rangle \underbrace{\langle k_t\rvert x\rangle}_{0} =0.
So the rank-one write changes the response in the selected key direction while leaving every orthogonal direction alone.
2.2 Derivation two: take one step on reconstruction loss
The same update falls out of an online learning objective. Treat the current key-value pair as one training example for the linear map SS:
Lt(S)=12∥S∣kt⟩−∣vt⟩∥22.\mathcal L_t(S) = \frac12 \left\| S\lvert k_t\rangle-\lvert v_t\rangle \right\|_2^2.
Its gradient with respect to the state is
∇SLt(S)=(S∣kt⟩−∣vt⟩)⟨kt∣.\nabla_S\mathcal L_t(S) = \left( S\lvert k_t\rangle-\lvert v_t\rangle \right) \langle k_t\rvert.
This is visibly an outer product: a value-space prediction error times the key bra at which that error was observed. Take one gradient-descent step of size βt\beta_t from St−1S_{t-1}:
St=St−1−βt∇SLt(St−1)=St−1−βt(St−1∣kt⟩−∣vt⟩)⟨kt∣=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣.\begin{aligned} S_t &= S_{t-1} - \beta_t\nabla_S\mathcal L_t(S_{t-1})\\ &= S_{t-1} - \beta_t \left( S_{t-1}\lvert k_t\rangle-\lvert v_t\rangle \right) \langle k_t\rvert\\ &= S_{t-1} + \beta_t \left( \lvert v_t\rangle-S_{t-1}\lvert k_t\rangle \right) \langle k_t\rvert. \end{aligned}
This is exactly the update we got by requiring immediate reconstruction. The two interpretations are the same:
as a memory operation, βt\beta_t controls how strongly to replace the old association;
as online learning, βt\beta_t is the step size;
as linear algebra, the change is a rank-one outer product.
2.3 The DeltaNet state transition
Expanding the error exposes DeltaNet as a structured state transition plus a new input:
St=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣=St−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1} + \beta_t \left( \lvert v_t\rangle-S_{t-1}\lvert k_t\rangle \right) \langle k_t\rvert\\ &= S_{t-1} \left( I-\beta_t\lvert k_t\rangle\langle k_t\rvert \right) + \beta_t\lvert v_t\rangle\langle k_t\rvert. \end{aligned}
For a unit key, I−βt∣kt⟩⟨kt∣I-\beta_t\lvert k_t\rangle\langle k_t\rvert has eigenvalue 1−βt1-\beta_t in the current key direction and eigenvalue 11 in every orthogonal direction. It removes the old association along the current key before adding the new one.
DeltaNet fixes the write. It does not yet fix the lifetime of the state.
3. Gated DeltaNet: sometimes old information should disappear
The linear state compresses the whole history into one matrix. A read
St∣q⟩=∑i≤t⟨ki∣q⟩∣vi⟩S_t\lvert q\rangle = \sum_{i\leq t} \langle k_i\rvert q\rangle\lvert v_i\rangle
cannot choose to skip an individual old token after that token has been folded into StS_t. Every stored direction that overlaps the query contributes. The delta rule can correct the state around the current key, but stale information in other directions remains available and can distort future reads.
We therefore need a way to forget the old state before using it. Let αt∈[0,1]\alpha_t\in[0,1] be a learned scalar retention gate:
S~t=αtSt−1.\widetilde S_t = \alpha_t S_{t-1}.
Run the same delta rule against this gated state:
S~t=αtSt−1,forget,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write.\boxed{ \begin{aligned} \widetilde S_t &= \alpha_tS_{t-1}, &&\text{forget},\\ \lvert\widehat v_t\rangle &= \widetilde S_t\lvert k_t\rangle, &&\text{predict},\\ \lvert e_t\rangle &= \beta_t \left( \lvert v_t\rangle-\lvert\widehat v_t\rangle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\rangle\langle k_t\rvert, &&\text{write}. \end{aligned} }
This is Gated DeltaNet. The order matters: forget first, predict from the retained state, then correct that prediction. If we predicted before forgetting, the error would describe a different memory from the one we update.
Expanding the recurrence gives
St=αtSt−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.S_t = \alpha_tS_{t-1} \left( I-\beta_t\lvert k_t\rangle\langle k_t\rvert \right) + \beta_t\lvert v_t\rangle\langle k_t\rvert.
The delta rule gives targeted replacement; the scalar gate gives global erasure. They solve different problems and are complementary.
But αt\alpha_t still makes one decision for the entire matrix. The model must retain or forget every key channel at the same rate.
4. Kimi Delta Attention: forget each channel independently
Kimi Delta Attention replaces Gated DeltaNet’s scalar retention with a vector αt∈[0,1]dk\alpha_t\in[0,1]^{d_k}. Put the vector on the diagonal:
Dt=Diag(αt)∈Rdk×dk.D_t = \operatorname{Diag}(\alpha_t) \in\mathbb R^{d_k\times d_k}.
Our state maps keys to values, so the key channels are the columns of SS. Right-multiplication applies a different retention factor to every one:
S~t=St−1Dt.\widetilde S_t = S_{t-1}D_t.
Everything else is the delta rule we have already derived:
S~t=St−1Dt,forget each key channel,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write,∣ot⟩=St(s∣qt⟩),s=dk−1/2,read.\boxed{ \begin{aligned} \widetilde S_t &= S_{t-1}D_t, &&\text{forget each key channel},\\ \lvert\widehat v_t\rangle &= \widetilde S_t\lvert k_t\rangle, &&\text{predict},\\ \lvert e_t\rangle &= \beta_t \left( \lvert v_t\rangle-\lvert\widehat v_t\rangle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\rangle\langle k_t\rvert, &&\text{write},\\ \lvert o_t\rangle &= S_t(s\lvert q_t\rangle), \qquad s=d_k^{-1/2}, &&\text{read}. \end{aligned} }
That is KDA. Compared with Gated DeltaNet, the conceptual change is only the promotion
αt⟶Dt=Diag(αt).\alpha_t \quad\longrightarrow\quad D_t=\operatorname{Diag}(\alpha_t).
The effect is substantial: one channel can be cleared while another is retained.
4.1 Why the transition is diagonal-plus-low-rank
Expand the KDA correction:
St=St−1Dt+βt(∣vt⟩−St−1Dt∣kt⟩)⟨kt∣=St−1Dt(I−βt∣kt⟩⟨kt∣)⏟At+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1}D_t + \beta_t \left( \lvert v_t\rangle - S_{t-1}D_t\lvert k_t\rangle \right) \langle k_t\rvert\\ &= S_{t-1} \underbrace{ D_t \left( I-\beta_t\lvert k_t\rangle\langle k_t\rvert \right) }_{A_t} + \beta_t\lvert v_t\rangle\langle k_t\rvert. \end{aligned}
The key-space transition is
At=Dt−βtDt∣kt⟩⟨kt∣=Dt−∣bt⟩⟨at∣,\begin{aligned} A_t &= D_t-\beta_tD_t\lvert k_t\rangle\langle k_t\rvert\\ &= D_t-\lvert b_t\rangle\langle a_t\rvert, \end{aligned}
where
∣bt⟩=Dt∣kt⟩,⟨at∣=βt⟨kt∣.\lvert b_t\rangle=D_t\lvert k_t\rangle, \qquad \langle a_t\rvert=\beta_t\langle k_t\rvert.
So AtA_t is a diagonal matrix minus a rank-one matrix: a diagonal-plus-low-rank, or DPLR, transition. “DPLR” describes the dk×dkd_k\times d_k transition acting on key space. The memory state itself is still the dv×dkd_v\times d_k matrix StS_t.
The full journey can now be summarized compactly:
The implementation usually stores gt=logαtg_t=\log\alpha_t with gt≤0g_t\leq0, then obtains the retention factors as exp(gt)\exp(g_t). In the transposed dk×dvd_k\times d_v layout used by the reference code, the recurrence is only five lines:
state = state * g_t.exp().unsqueeze(-1) prediction = einsum(“bhkv,bhk->bhv”, state, k_t) residual = beta_t.unsqueeze(-1) * (v_t - prediction) state = state + einsum(“bhk,bhv->bhkv”, k_t, residual) output = einsum(“bhk,bhkv->bhv”, q_t * scale, state)
See the official naive_recurrent_kda reference.
Casa Grande, Arizona
They Are Not Just Reading Plates
Flock cameras are sold as license plate readers. Their own patent describes something bigger: video analysis that can identify people, assign traits, and make real-world activity searchable.
Safety without mass surveillance
Who We Are
Deflock Casa Grande is a nonpartisan civic campaign.
We are asking Casa Grande leaders for public safety that supports law enforcement with clear guidelines, public oversight, audits, and accountability.
Not political
Not left-wing, not right-wing, and not affiliated with any political party, candidate, or partisan organization.
Independent
Not funded by anyone, not backed by any outside group, and not associated with other Deflock websites.
Pro law enforcement
We support public safety and the people responsible for protecting Casa Grande.
Pro accountability
Powerful surveillance systems need rules, oversight, audits, and consequences for misuse.
Latest updates
Updates on Flock in Casa Grande
Recent developments show why residents need clear rules, public notice, and enforceable limits before the city expands Flock surveillance.
New local record
Casa Grande Email Says Cameras Monitored a Student Protest
The January 30, 2026 email lists “Monitoring of a student protest at City Hall” among recent camera-system wins.
The same page says the system assisted in locating missing or endangered individuals and allowed police to “identify involved parties” and “monitor activity” before patrol arrived.
That matters because residents are told these cameras track vehicles, not individuals. The city’s own email describes uses involving people, suspects, and protesters.
The source document says “student protest.” We describe it as peaceful because the email does not allege violence, a crime, or a public-safety threat tied to that protest.
Source: January 30, 2026 Casa Grande police email
New Flock audio concern
Flock Is Now Recording Conversations
Flock has started early access for what it calls distress monitoring, expanding gunshot-detection microphones toward listening for human distress sounds.
Flock and the city have been adding what they call gunshot detection systems. Now those microphone systems can be upgraded to include “distress.”
For a microphone to capture distress, it has to be listening for events. That turns these devices into always-listening stations around the city.
EFF warned about this risk in 2025, when it reported that Flock gunshot-detection microphones were being expanded to listen for human voices and distress sounds.
Source: EFF: Flock’s Gunshot Detection Microphones Will Start Listening for Human Voices
Flock video camera concern
Flock ALPRs Are Not Just License Plate Readers
Flock’s own product launch says existing LPR cameras can become video-enabled through a cloud software update.
Most people think Flock is just a safety system for license plates looking for bad actors. But Flock says its LPR cameras can provide video context too.
That means the system can capture more than a plate. It can collect video of vehicles moving through neighborhoods, roads, and cities, regardless of whether the driver is suspected of a crime.
When video, location, time, and AI search are combined, residents face a bigger risk: movement records stored in a powerful database, exposed to broad authorized-user searches, misuse, warrantless access, and security breaches.
Source: Flock Safety Q2 2025 Law Enforcement Product Launch Summary
School camera concern
Flock Is Now in Your Elementary Schools
Casa Grande approved connecting existing elementary school cameras to the Flock ecosystem through Flock gateways.
On May 4, 2026, City Council voted to approve connecting existing Casa Grande elementary school cameras to the Flock ecosystem through Flock gateways.
Parents deserve clear notice before school cameras connect to a massive AI surveillance network.
The agreements raise serious questions about remote access, undefined authorized users, vague public safety purposes, AI analytics, retention, audits, and parent notification.
Plain-English explanation
What Are We Really Installing?
Flock cameras are often described as license plate readers. That sounds narrow, routine, and harmless.
But Flock’s own patent describes something more powerful: a system that analyzes video, identifies objects, including people, assigns attributes, and turns that information into searchable records.
In plain English, this is not just a camera looking for stolen cars. It is the blueprint for a searchable surveillance network. Casa Grande residents deserve to know exactly what it can collect, who can search it, and what rules stop abuse.
Casa Grande coverage
The City Is Building a Surveillance Grid
Full deployment of all technology will take approximately twelve months. After full deployment, the city says there will be:
100
License Plate Reader Cameras
100
Pan Tilt Zoom Cameras
31
Gun Shot Detection Systems
This Is Beyond Plate Reading
License Plate Reader (LPR) Cameras
The city says they identify vehicles associated with investigations, not individuals. But LPRs capture image records, plate reads, timestamps, and locations that can reconstruct where people travel over time.
Pan-Tilt-Zoom (PTZ) Cameras
The city says trained personnel operate them during incidents. In practice, zoomable cameras can follow any person in view, including women and children, not just people suspected of crimes.
Drone as First Responder (DFR)
Marketed as aerial support for faster response, but drones can also let police look into backyards and other private outdoor spaces.
Gunshot Detection System
The city says it detects gunfire and does not record conversations. EFF warns Flock gunshot-detection microphones are being expanded to listen for human voices and distress sounds.
Read the EFF warning
Flock OS Platform
Connects the surveillance network into one platform. Flock has described integrations that can combine camera data with OSINT-style information, including online profiles, posts, addresses, and identity clues.
Taxpayer tradeoff
Cameras do not patrol, de-escalate, interview witnesses, or build trust.
Casa Grande should prove this mass-surveillance spend reduces serious crime before spending roughly $1,000,000 a year on cameras, sensors, drones, and software.
Better use of tax money
We could have funded 13 more dedicated officers keeping us safe instead of mass surveillance.
A million dollars a year buys people, not just cameras.
Casa Grande approved a roughly 10-year, $10 million Safe City/Flock contract, which averages about $1 million per year before considering reported taxes and later-year annual costs.
Officer staffing math
13 officers
Casa Grande lists Police Officer base pay at $63,714-$87,824, with a $74,762 midpoint. At that midpoint, $1,000,000 / $74,762 = about 13 officer base salaries.
Benefits, equipment, training, and vehicles add real costs, so this is not a full staffing budget. But it shows the tradeoff: the same annual money could fund multiple additional sworn officers instead of locking Casa Grande into a permanent surveillance network.
Documented abuse pattern
Abuse Is Already Happening Across the Country
Flock and other ALPR systems can reveal where a person goes over time. That kind of location history is sensitive for the same reason cell phone location records are sensitive, but ALPR searches are often not regulated with a warrant standard.
The harm is not hypothetical. The Institute for Justice and continued reporting have documented at least 20 reported cases of police using ALPR access for romantic stalking, including searches of spouses, ex-partners, romantic interests, and their friends.
Source: Institute for Justice: Police Have Reportedly Used License Plate Readers to Stalk Romantic Interests at Least 14 Times in Recent Years
Note: deflockcg updated this list from 14 to 20 documented cases on 6/29/26.
Timeline of documented abuse - latest first
2026 Prairie Grove, Illinois Officer William C. Copp, who also served as the police chief of nearby Holiday Hills, was arrested after searching Flock for several former romantic partners and at least one of their new partners. Copp has been fired from his Prairie Grove position and his employment with Holiday Hills is under review.
2026
Prairie Grove, Illinois
Officer William C. Copp, who also served as the police chief of nearby Holiday Hills, was arrested after searching Flock for several former romantic partners and at least one of their new partners. Copp has been fired from his Prairie Grove position and his employment with Holiday Hills is under review.
2026 Winnebago County, Illinois Former Sheriff’s Deputy Tyler Bryan was charged with stalking and official misconduct after allegedly using the department’s ALPR system to monitor the locations of an ex-girlfriend and her new partner. The misconduct came to light after the victims filed for an order of protection against Bryan.
2026
Winnebago County, Illinois
Former Sheriff’s Deputy Tyler Bryan was charged with stalking and official misconduct after allegedly using the department’s ALPR system to monitor the locations of an ex-girlfriend and her new partner. The misconduct came to light after the victims filed for an order of protection against Bryan.
2026 Niceville, Florida Former Niceville Officer Coty Hall pleaded no contest to several charges after using the department’s Flock system to track another officer and that officer’s spouse. Hall’s misconduct was discovered via an internal audit; Hall was fired following his arrest in October 2025.
2026
Niceville, Florida
To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".
10HN is also available as an iOS App
If you visit 10HN only rarely, check out the the best articles from the past week.
Visit pancik.com for more.