10 interesting stories served every morning and every evening.

Elevators

john.fun

Everyone has shared the frus­tra­tion of wait­ing for an el­e­va­tor that never seems to ar­rive. I pressed the but­ton, why is­n’t it com­ing?” you ask. For some­thing as com­mon­place as el­e­va­tors, they are far more com­plex than meets the eye.

Over the course of this ar­ti­cle, we’ll un­ravel the mys­ter­ies of el­e­va­tors. The way you push their but­tons, and how they push yours.

One Car

The sim­plest el­e­va­tor al­go­rithm is called SCAN and was patented in 1961. The el­e­va­tor starts at the lobby and goes all the way to the top floor be­fore re­vers­ing and com­ing back down. It picks up and drops off any­body on the way.

Most of the time you don’t ac­tu­ally need to go to the TOP floor. If the el­e­va­tor goes only as high as re­quested be­fore re­vers­ing, the al­go­rithm is called the LOOK al­go­rithm. This is the al­go­rithm most peo­ple know and ex­pect.

Multiple Cars

Here’s where the mys­tery be­gins. If there are mul­ti­ple el­e­va­tors, how do the cars co­or­di­nate who picks up who?

In the most ba­sic sys­tem, there’s a cen­tral sched­uler that tells each el­e­va­tor which floors to stop on. When a new re­quest comes in, it’s as­signed to the clos­est el­e­va­tor. As we’ll soon see how­ever, we can do bet­ter.

Long Waits

How do you ac­tu­ally mea­sure how good an el­e­va­tor al­go­rithm is? The ob­vi­ous met­ric is how long you wait for the el­e­va­tor to ar­rive.

A very sim­ple mea­sure is how of­ten does the el­e­va­tor ar­rive within 30 sec­onds?” Or how of­ten does the el­e­va­tor ar­rive within 90 sec­onds?”

-

wait < 30s

-

wait < 90s

flow14/​min

Applied Stats

More rig­or­ously, we want to look at the DISTRIBUTION of wait times. If we plot the wait time across thou­sands of rides, we get the his­togram be­low.

010s

p50—p90—

flow14/​min

A p90 of 2m means 90% of the time, rid­ers wait 2m or less for the el­e­va­tor. A p50 of 1m means half the time the el­e­va­tor ar­rives within 1m.

People don’t usu­ally re­mem­ber the av­er­age amount of time they wait. They fix­ate on those times when the el­e­va­tor took FOREVER, the p90 case.

Morning Rush

Not all pas­sen­ger traf­fic is cre­ated equal. Imagine a large cor­po­rate of­fice build­ing. In the morn­ings, nearly all traf­fic is dom­i­nated by trips from the lobby to the up­per lev­els.

In the evening this flips as every­one leaves the build­ing. The lunch rush is a bit of both, and the re­main­ing traf­fic is of­ten from floor to floor.

-

wait < 30s

-

wait < 90s

flow14/​min

The dis­tri­b­u­tion of wait times varies dras­ti­cally de­pend­ing on the time of day and the traf­fic pat­terns the el­e­va­tors are fac­ing. Morning rush no­to­ri­ously has the worst wait sta­tis­tics.

Smarter Elevators

When an­a­lyz­ing the LOOK el­e­va­tor al­go­rithm, we LOOKED (ha ha) at how rid­ers are as­signed to cars. We naively as­signed each re­quest to the near­est car but said we could do bet­ter.

What if the near­est car is full? We can get smarter with Otis’ RSR (Relative System Response) al­go­rithm. RSR scores each car for how well suited it is to pick up a pas­sen­ger. Lower scores be­ing bet­ter.

RSR pickup score

Score=ETA to pickup+on­board load penalty+same-di­rec­tion anti-bunch­ing penalty-di­rec­tion-match bonus-idle-nearby bonus-low-load bonus

RSR also re-op­ti­mizes every 5 sec­onds. A pas­sen­ger that’s go­ing to be picked up by el­e­va­tor A can be re-routed to el­e­va­tor B if el­e­va­tor A en­coun­ters de­lays. This re-op­ti­miza­tion turns out to be key for stream­lin­ing traf­fic flow.

In the graphic be­low, each el­e­va­tor lights up when it’s the best choice to ser­vice a call from floor 3 if the but­ton hap­pened to be pressed at that ex­act mo­ment. This con­stantly changes as the el­e­va­tors move, show­ing the op­ti­mizer in mo­tion.

LOOK vs RSR

Armed with our el­e­va­tor analy­sis toolkit, we can bench­mark the per­for­mance of LOOK vs RSR to see how much a smarter el­e­va­tor al­go­rithm ac­tu­ally im­proves wait time.

LOOK

-wait < 30s

-wait < 90s

RSR

-wait < 30s

-wait < 90s

flow8/​min

Interestingly as the flow rate gets higher, LOOK ac­tu­ally starts to out­per­form RSR. When the el­e­va­tors are al­ways full and stop­ping on every floor, the ex­tra rules don’t mat­ter as much.

LOOK also tends to out­per­form RSR in small build­ings with fewer el­e­va­tors per bank. Sometimes it’s bet­ter to just keep things sim­ple.

Another met­ric you can track is jour­ney time, how long you’re ac­tu­ally wait­ing in the el­e­va­tor be­fore get­ting to your floor. RSR and LOOK have dif­fer­ent char­ac­ter­is­tics here as well but that’s be­yond the scope of this ar­ti­cle.

Destination Dispatch

Not all el­e­va­tors have but­tons in them. Some of the fancy new el­e­va­tors have a kiosk on each floor that al­lows you to spec­ify what floor you’re head­ing to be­fore the el­e­va­tor even ar­rives. The kiosk then points you to which el­e­va­tor you should wait for.

This is called Destination Dispatch. At first glance, it seems great. The el­e­va­tor op­ti­mizer now has full knowl­edge of who is go­ing where, cer­tainly we can use this to re­duce wait times right?

RSR

-wait < 30s

-wait < 90s

Destination Dispatch

-wait < 30s

-wait < 90s

flow8/​min

It turns out these fancy kiosks are in gen­eral worse for wait times com­pared to the tra­di­tional good ol’ up and down but­tons. There are cer­tainly edge cases when the kiosks can win out (extremely tall build­ings with 8+ cars per el­e­va­tor bank) but for the ma­jor­ity of cases, sim­ple up down but­tons reign supreme.

This coun­ter­in­tu­itive re­sult is all thanks to the re­bal­anc­ing step where every 5 sec­onds, the sys­tem re-op­ti­mizes each el­e­va­tor’s path. The kiosk en­forces rigid­ity, you must get in the as­signed el­e­va­tor.

The state of the world 30sec af­ter you called your el­e­va­tor might be very dif­fer­ent but the sys­tem is un­able to adapt. Turns out the loss in flex­i­bil­ity is not worth the ex­tra in­for­ma­tion for the op­ti­mizer.

Full Sim

Here’s a sim­u­la­tion with all the but­tons and knobs to play with. Go crazy!

-

wait < 30s

-

wait < 90s

floors8­cars4flow18/​min

Conclusion

This ar­ti­cle just scratches the sur­face of el­e­va­tor al­go­rithms. Next time you’re stuck wait­ing for an el­e­va­tor, try not to take it per­son­ally. The el­e­va­tor did hear you, it just has a lot to think about.

GitHub - yc-software/qm: Multiplayer agent harness for work

github.com

A mul­ti­player agent har­ness for work. In Slack and on the web.

What is QM?

Most agents are de­signed like per­sonal as­sis­tants. You can make one work for a whole com­pany, but it quickly gets com­plex. QM is de­signed for star­tups. Employees each get their own iso­lated work­space and work in­de­pen­dently with­out af­fect­ing each other, and they can also col­lab­o­rate with the agent in chan­nels, group mes­sages, and pro­jects.

Each per­son and each room has its own scoped mem­ory, files, key­chain view, per­mis­sions, crons, web apps, and durable sand­box.

It’s built with open source in mind. Pick your own har­ness and model and switch be­tween them — Pi, OpenCode, Codex, and Claude Code all drive the same core, so a de­ploy­ment is­n’t tied to any sin­gle ven­dor.

Features

Personal and shared scopes. People cus­tomize the agent to be theirs, and still work with it col­lab­o­ra­tively in Slack chan­nels and pro­jects.

Slack and web. The same iden­tity and con­fig­u­ra­tion car­ries be­tween Slack and the web app.

Admin con­trol. Set org-level con­fig­u­ra­tion, a se­cu­rity pos­ture, and which har­nesses and mod­els are avail­able.

Web apps. Spin up cus­tom in­ter­nal apps and pub­lish them to the right peo­ple.

Shared skills. Skills are scope-owned and share­able by grant, with ad­min-gated pro­mo­tion to the whole org and skill packs im­ported from git repos­i­to­ries.

Background work. Crons and watches run work while no­body’s watch­ing.

What you can do with it

Search in­ter­nal notes, email, doc­u­ments, data­bases, and the web to­gether

Retrieve in­for­ma­tion from your com­pany brain

Build in­ter­nal apps, pub­lish them to the right peo­ple, and keep their data cur­rent

Learn your writ­ing voice from past sends, then triage your in­box on a sched­ule — la­bels and re­ply drafts in­cluded

Work in an ex­ist­ing repos­i­tory: run tests, open PRs, mon­i­tor CI, check sys­tem logs

Track a pro­ject in a shared chan­nel and post up­dates and fol­low-ups

Architecture

flow­chart LR DB[(“Postgres<br/>sessions · mem­ory · queue”)]

sub­graph CORE[“Headless core”] API[“API · iden­tity · pol­icy · sched­uler”] LOOP[“Agent loop<br/&​gt;(Pi, OpenCode, Claude Code)“] API <–> LOOP end

SBX[“Per-scope sand­box<br/&​gt;files · tools · logged-in ser­vices”]

DB <–> API LOOP <–> SBX

Every turn runs through a cen­tral core, which can use a va­ri­ety of mod­els and har­nesses to gen­er­ate the re­sponse. A Postgres per­sis­tence layer holds user data, ses­sion his­tory, and other durable state. The agent has a small, fixed tool sur­face; one of those tools is ex­e­cute, which runs com­mands in the scope’s own iso­lated sand­box — its durable com­puter, where in­stalled tools stay in­stalled. The web UI, the ad­min panel, and the pub­lic por­tal are op­tional plu­g­ins over the core’s HTTP API; Slack is an op­tional in-process plu­gin that core starts and su­per­vises through a di­rect ser­vice client.

The core runs TypeScript di­rectly on Node and uses Fastify for HTTP. The Slack plu­gin uses Bolt; the web UI builds with Vite and ren­ders with Lit.

The core it­self is generic. Everything spe­cific to one com­pany — org con­fig, cus­tom tools and skills, sand­box im­age, in­fra­struc­ture — lives in a de­ploy­ment di­rec­tory that the qm CLI val­i­dates and de­ploys. Every sub­strate (harness, ses­sion store, sand­box, mem­ory) sits be­hind an in­ter­face, so pro­duc­tion im­ple­men­ta­tions swap in via one wiring file.

Security and se­crets

QMs ap­proach fol­lows lo­cal cod­ing agents like OpenCode, Codex, and Claude Code: the agent acts as the per­son it’s work­ing for, with their cre­den­tials and per­mis­sions, and every­thing it does is au­dited. An org picks one se­cu­rity pos­ture, which nar­rower scopes can only tighten:

Strict — every har­ness tool call pauses for hu­man ap­proval, ex­cept the two no-ef­fect turn en­ders.

Auto (default) — a clas­si­fier screens prove­nance-la­belled ex­ter­nal data and tool re­sults be­fore they reach the model; a de­ploy­ment can point that at its own screen­ing proxy.

Dangerous — no con­tent screen­ing, no pauses be­tween tool calls.

The pre­de­clared com­mand pol­icy — ap­proval rules and hard de­nials for things like re­cur­sive deletes or de­struc­tive SQL — ap­plies in every pos­ture, Dangerous in­cluded.

SECURITY.md has the threat model, the op­er­a­tor as­sump­tions, and the known lim­i­ta­tions.

Deploy it for your org

Create an or­ga­ni­za­tion-owned de­ploy­ment repos­i­tory that de­pends on @yc-software/qm:

npm exec –yes –package=@yc-software/qm@latest — \ qm init . –org <slug> –target <fly-or-aws> npm in­stall

Initialization ma­te­ri­al­izes a de­ploy­ment skill for an agent and walks through in­fra­struc­ture, web sign-in, con­nec­tor cre­den­tials, op­tional Slack ac­cess, de­ploy­ment, and live ver­i­fi­ca­tion — no source check­out re­quired. Each de­ploy­ment runs in the op­er­a­tor’s own cloud ac­count; ini­tial­iza­tion does not gen­er­ate or en­able de­ploy­ment CI, and this repos­i­tory has no pro­duc­tion de­ploy­ment work­flow. See de­ploy­ment.md for the de­tails.

Contributing

We take con­tri­bu­tions as hu­man-writ­ten text, not code — see CONTRIBUTING.md. Describe the change you’d like in­for­mally in a .txt or .md file in adrs/, and if we’re aligned we’ll han­dle the im­ple­men­ta­tion. Report vul­ner­a­bil­i­ties pri­vately — see SECURITY.md, not a pub­lic is­sue.

Customize your in­stance

The de­ploy­ment repos­i­tory above car­ries con­fig and a sand­box layer, and never needs a source check­out. Some or­ga­ni­za­tions want the op­po­site trade: the whole code­base in one place, so en­gi­neers and cod­ing agents read core and cus­tomiza­tions to­gether, while the cus­tomiza­tions them­selves stay pri­vate. For that, keep a pri­vate fork: a stand­alone pri­vate repos­i­tory whose his­tory be­gins as a clone of qm and whose core stays iden­ti­cal to up­stream.

Populate it once, then clone it to work in:

gh repo cre­ate <org>/qm-private –private

git clone –bare git@github.com:yc-software/qm qm-seed.git git -C qm-seed.git push –mirror git@github.com:<org>/qm-private rm -rf qm-seed.git

git clone git@github.com:<org>/qm-private git -C qm-pri­vate re­mote add up­stream git@github.com:yc-software/qm

Create the pri­vate fork with a plain clone, as shown above, and never with GitHub’s fork fea­ture. The word fork” here names the con­cept — a down­stream copy that di­verges de­lib­er­ately and merges from up­stream — not GitHub’s Fork but­ton. A GitHub fork in­her­its the vis­i­bil­ity of the repos­i­tory it came from, so a fork of a pub­lic repos­i­tory can­not be made pri­vate. A GitHub fork also shares one ob­ject net­work with the repos­i­tory it came from, so com­mits pushed to the fork stay fetch­able by SHA from the pub­lic side. Many or­ga­ni­za­tions dis­al­low fork­ing pri­vate repos­i­to­ries as well. A plain clone has none of these prob­lems, and it costs one thing: the clone is an or­di­nary repos­i­tory, so up­stream’s CI work­flows run live in your own ac­count. Expect to sup­ply the se­crets those work­flows need, or dis­able the ones you do not want run­ning.

Everything spe­cific to your or­ga­ni­za­tion goes in de­ploy/​lay­ers/&​lt;org>/ — con­fig, sand­box tools and skills, plu­gin im­ages, in­fra­struc­ture — in the same shape qm init pro­duces. See de­ploy/​lay­ers/​README.md. Core stays byte-iden­ti­cal to up­stream, which is what keeps merges small.

Two skills main­tain the bound­ary in both di­rec­tions. up­date-qm merges up­stream qm into the pri­vate fork and opens the sync PR; up­stream-pr sends an or­ga­ni­za­tion-ag­nos­tic fix back to qm, cut­ting the branch from up­stream/​main and check­ing the out­go­ing diff, com­mit mes­sages, and screen­shots for or­ga­ni­za­tion iden­ti­fiers be­fore it pushes. Nothing un­der de­ploy/​lay­ers/ ever trav­els up­stream.

Going deeper

docs/​get­ting-started.md — first run, end to end

cli/​README.md — the qm CLI and the de­ploy­ment di­rec­tory con­tract

docs/​de­ploy-di­rec­tory.md — the de­ploy­ment di­rec­tory in full

.env.example — every knob, doc­u­mented in place

plu­g­ins/ — the sur­faces (Slack, web UI, ad­min, por­tal)

License

Except where oth­er­wise noted, QM is avail­able un­der the MIT License.

Tailscale didn’t stop the Hugging Face intrusion

tailscale.com

By now, you’ve likely heard about the AI agent that es­caped a se­cu­rity eval­u­a­tion and at­tacked Hugging Face, an LLM mar­ket­place. The agent de­cided Hugging Face might have the an­swers to its bench­mark, so it stole them just to cheat on the exam. That’s a funny mo­tive but a scary out­come.

Hugging Face pub­lished a de­tailed re­con­struc­tion of the in­tru­sion. It cov­ers about 17,600 re­cov­ered ac­tions over four and a half days, in­clud­ing sand­box es­capes, code ex­e­cu­tion, cloud cre­den­tials, im­pro­vised com­mand-and-con­trol sys­tems, and even­tu­ally, the use of Tailscale to spread through­out their or­ga­ni­za­tion.

But Tailscale is a zero trust net­work! The whole point of zero trust is to pre­vent at­tack­ers from mov­ing lat­er­ally across your com­pany. What went wrong?

Where Tailscale ap­peared

No vulnerabilities” in Tailscale were found or ex­ploited, and that might make it even more un­com­fort­able for us. Tailscale turns up every­where in AI in­fra­struc­ture: vir­tu­ally every AI com­pany uses Tailscale. So it’s not sur­pris­ing we were in an AI in­ci­dent re­port even­tu­ally.

But, we’re a se­cu­rity tool. Their in­tru­sion is our in­tru­sion, and it’s our job to take it se­ri­ously. As in any post-mortem, we need to re­flect on what hap­pened and what we could have done to pre­vent it.

By the time the agent es­caped its sand­box and found Tailscale, it had al­ready gained code ex­e­cu­tion priv­i­leges in­side a pro­duc­tion worker, reached root ac­cess on a Kubernetes node, and read a pro­duc­tion se­cret store con­tain­ing 136 keys. In a sense, it was game over be­fore we even ar­rived on the scene.

But is­n’t that the prob­lem? Let’s back up a level or two. Why were long-lived se­cret keys ac­ces­si­ble at all to any­one?

Because long-lived cre­den­tials are the stan­dard

They should­n’t be, though. And we’ve now seen a great ex­am­ple of why. In the old world where most in­tru­sions were done by hu­mans at hu­man speed, cre­den­tial leak mit­i­ga­tions were treated as a nice-to-have. A big cre­den­tial store, where you can read 136 keys at once, was a to-do item some­where in a se­cu­rity team’s low-pri­or­ity list.

Now, in a world of rogue AI agents, the big cre­den­tial vault is the prize. It’s not okay any­more.

I only know of two main so­lu­tions to long-lived cre­den­tials.

First, you can con­fig­ure a vault that only is­sues short-lived creds based on long-lived creds that you in­sert once and that it never gives back. HashiCorp Vault can do this, for ex­am­ple; they call it dy­namic cre­den­tials. Unfortunately, dy­namic cre­den­tials are a lot of work to set up and main­tain. When se­cu­rity re­quires work, peo­ple don’t do it.

The sec­ond op­tion is a cre­den­tial-in­ject­ing proxy. If client C wants to ac­cess server S, don’t give C cre­den­tials for S; in­stead, have it send its re­quests via a hard­ened proxy P that does noth­ing but re­ceive a re­quest, in­sert a cre­den­tial, and pass it through. Credential in­jec­tion is gain­ing pop­u­lar­ity now that mod­ern mesh-based net­work­ing re­moves most of its la­tency. A few months ago, we ac­quired Border0 (“Tailscale PAM), which makes an el­e­gant and quite fancy cre­den­tial-in­ject­ing proxy, tightly in­te­grated with Tailscale. Border0, fully de­ployed, would have pre­vented read­ing those 136 keys and logged every at­tempt to use one. (The same cre­den­tial in­jec­tion pow­ers Aperture Connectors.) But, this stuff is too re­cent. Most of our cus­tomers haven’t even looked at it yet.

A bonus third op­tion: Tailscale it­self. In Tailscale, your node key can be ma­chine bound us­ing a TPM to pre­vent ex­port, and if you’re an en­ter­prise you should def­i­nitely do that. Alas, we had to turn TPM stor­age off by de­fault on Linux and Windows, be­cause of HSM is­sues we found on some users’ hard­ware. Anyway, in to­day’s case, the at­tack is what granted ac­cess to Tailscale, so we ar­rived too late.

That long-lived auth key did­n’t need to ex­ist

Neither dy­namic cre­den­tials nor cre­den­tial-in­ject­ing prox­ies were in place, so 136 keys leaked. What’s the next line of de­fense?

One of those 136 cre­den­tials was a reusable Tailscale auth key, used to cre­ate new Tailscale CI (continuous in­te­gra­tion, used for au­to­mated test­ing) nodes in their tail­net. The agent copied that key into a se­ries of ex­ter­nal sand­boxes and used it, over sev­eral days, to en­roll a to­tal of 181 nodes into Hugging Face’s tail­net. Those nodes each re­ceived a Tailscale iden­tity tag grant­ing all the ac­cess a CI node would get.

We could­n’t save those 136 keys. But a reusable Tailscale key did­n’t need to be among them.

We built work­load iden­tity fed­er­a­tion for cases like this. It re­lies on your cloud provider to of­fer ambient au­thor­ity” in the form of short-lived creds gen­er­ated from a run­ning VMs (or con­tain­er’s) cloud iden­tity. The CI job can ask its cloud plat­form for a signed OIDC to­ken. Tailscale ver­i­fies the to­ken, then grants ac­cess with the tags and scopes as­signed to that work­load. What’s neat is this can hap­pen au­to­mat­i­cally once en­abled: spin up the CI node, Tailscale gets the iden­tity, as­signs the right tags. There’s no cre­den­tial to leak and, when con­fig­ured care­fully, no way to use this cre­den­tial any­where but in CI.

That means you can’t take a CI cre­den­tial, run it on a node else­where, and turn it into a paved route out of the test clus­ter into the rest of the net­work.

Alas, not enough peo­ple use work­load iden­tity fed­er­a­tion. We should have done more to make sure cus­tomers knew it’s the eas­ier and safer path and had a clear way to adopt it. We can im­prove the docs, and add more nudges in the UI.

No client logs does­n’t mean no ev­i­dence

In an at­tempt to hide its tracks, the agent ran Tailscale with –no-logs-no-support, which sup­presses re­port­ing from that client.

That’s an op­tion de­signed for users who are con­cerned about send­ing teleme­try meta­data to Tailscale. Even if we did­n’t of­fer it, it would be easy to mod­ify the source code to re­move the teleme­try.

But stop­ping the logs does­n’t make the con­nec­tion in­vis­i­ble. If you en­able Tailscale net­work flow logs, they re­port traf­fic from both ends of every con­nec­tion, as well as from sub­net routers and exit nodes. This is sub­tle but im­por­tant: a com­pro­mised node might not send flow logs, but every node it con­nects to does. And then your SIEM, con­fig­ured with care, can raise an im­me­di­ate red alert if the two ends don’t match.

Flow logs can help de­tec­tion when they stream into a care­fully con­fig­ured SIEM. But that’s a lot of work. Flow logs need to be en­abled, and you need to have the right live de­tec­tion rules in place so they’re use­ful in real time, not just for foren­sics later. We’re look­ing at how to make flow logs eas­ier to dis­cover, con­fig­ure, adopt, and serve as alert trig­gers. I want us to make flow logs so easy to use that they help even if you don’t have a se­cu­rity team to watch them.

If you want di­rect con­trol be­yond just log­ging, you can also en­able Tailnet Lock. This gives you di­rect vis­i­bil­ity and strict, pro­gram­ma­ble ad­mis­sion con­trol for every sin­gle new node. For ex­am­ple, with some work, you could pro­gram your sign­ing node to check that CI tags al­ways have a par­tic­u­lar IP ad­dress range or other side-chan­nel proof of va­lid­ity.

Make the safe path the easy path

Network se­cu­rity is hard. It has al­ways been hard. In the new world of rogue AI agents, it’s not just hard, but es­sen­tial. And that’s a prob­lem be­cause many orgs sim­ply don’t have net­work se­cu­rity ex­per­tise.

So at Tailscale, we take it per­son­ally. People ex­pect our prod­uct to pre­vent these sorts of lat­eral move­ment at­tacks, by de­fault, so they don’t have to. Even if they have no idea what a lat­eral move­ment at­tack is.

If this in­ci­dent has you look­ing a lit­tle ner­vously at your own in­fra­struc­ture, start by look­ing at the reusable Tailscale auth keys your work­loads can read. For cloud and CI in par­tic­u­lar, re­place them with work­load iden­tity fed­er­a­tion wher­ever you can. Get rid of those long-lived auth keys.

(Auth keys still have good uses, es­pe­cially for one-time pro­vi­sion­ing and en­vi­ron­ments with­out a plat­form iden­tity. When you need one, pre­fer one-off keys; use OAuth clients to keep the auth key ex­piry pe­ri­ods short; use nar­row tags; au­dit the per­mis­sions granted to those keys in your ACLs.)

Turn on net­work flow logs and send them to the tools your se­cu­rity team al­ready uses.

Use se­cure node state stor­age on man­aged fleets, where you have con­trol over your TPMs. Use de­vice pos­ture to iso­late and re­strict nodes where you don’t.

I know we haven’t made these safer choices ob­vi­ous enough. That’s on us. We’ll im­prove our docs, add nudges to the UI, do our best to turn these on by de­fault, warn you when you’re do­ing some­thing dan­ger­ous, and sug­gest bet­ter al­ter­na­tives.

This is our very Canadian apol­ogy: sorry you stepped on our toes. The at­tack did­n’t ex­ploit Tailscale, and Tailscale did­n’t cause the com­pro­mise. But, we did­n’t stop it. Next time, we will.

If you run Tailscale and want to dig deeper, get in touch with our sup­port and so­lu­tions en­gi­neer­ing teams. We can help you harden your set­tings and help you find the rough edges be­fore the next AI agent does.

DeepSeek V4 Flash 0731 (max) - Intelligence, Performance & Price Analysis

artificialanalysis.ai

Intelligence

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1 in­cor­po­rates 9 eval­u­a­tions: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

Reasoning mod­els are in­di­cated by a light­bulb icon

Artificial Analysis Intelligence Index v4.1 in­cludes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index method­ol­ogy for fur­ther de­tails, in­clud­ing a break­down of each eval­u­a­tion and how we run them.

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1 in­cor­po­rates 9 eval­u­a­tions: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

Reasoning mod­els are in­di­cated by a light­bulb icon

Artificial Analysis Intelligence Index v4.1 in­cludes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index method­ol­ogy for fur­ther de­tails, in­clud­ing a break­down of each eval­u­a­tion and how we run them.

Intelligence Evaluations

Intelligence eval­u­a­tions mea­sured in­de­pen­dently by Artificial Analysis · Higher is bet­ter

Agentic busi­ness op­er­a­tions

Reasoning mod­els are in­di­cated by a light­bulb icon

While model in­tel­li­gence gen­er­ally trans­lates across use cases, spe­cific eval­u­a­tions may be more rel­e­vant for cer­tain use cases.

AA-Omniscience

AA-Omniscience Index

AA-Omniscience Index (higher is bet­ter) mea­sures knowl­edge re­li­a­bil­ity and hal­lu­ci­na­tion. It re­wards cor­rect an­swers, pe­nal­izes hal­lu­ci­na­tions, and has no penalty for re­fus­ing to an­swer. Scores range from -100 to 100, where 0 means as many cor­rect as in­cor­rect an­swers, and neg­a­tive scores mean more in­cor­rect than cor­rect.

Reasoning mod­els are in­di­cated by a light­bulb icon

AA-Omniscience Index (higher is bet­ter) mea­sures knowl­edge re­li­a­bil­ity and hal­lu­ci­na­tion. It re­wards cor­rect an­swers, pe­nal­izes hal­lu­ci­na­tions, and has no penalty for re­fus­ing to an­swer. Scores range from -100 to 100, where 0 means as many cor­rect as in­cor­rect an­swers, and neg­a­tive scores mean more in­cor­rect than cor­rect.

Openness Index

Artificial Analysis Openness Index: Score

Openness Index as­sesses model open­ness on a 0 to 100 nor­mal­ized scale (higher is more open)

Reasoning mod­els are in­di­cated by a light­bulb icon

Intelligence Index Comparisons

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted av­er­age cost (USD) per Artificial Analysis Intelligence Index task

Most at­trac­tive quad­rant

Pareto line

Reasoning mod­els are in­di­cated by a light­bulb icon

Weighted av­er­age cost per Intelligence Index task. Each eval­u­a­tion’s cost is cal­cu­lated from in­put, cache hit, cache write, rea­son­ing, and an­swer to­ken prices, di­vided by task count, and weighted by its Intelligence Index weight.

Token Use

Output Tokens per Intelligence Index Task

Weighted av­er­age num­ber of out­put to­kens used to run one task in the Artificial Analysis Intelligence Index

Reasoning mod­els are in­di­cated by a light­bulb icon

The num­ber of to­kens re­quired per Intelligence Index task. This is cal­cu­lated by mul­ti­ply­ing the out­put to­kens per eval by the rel­a­tive weights of each bench­mark in the Intelligence Index, then di­vid­ing by task count (excluding re­peats).

Cost

Cost per Intelligence Index Task

Weighted av­er­age cost (USD) per Artificial Analysis Intelligence Index task, seg­mented by to­ken type. Lower is bet­ter

Reasoning mod­els are in­di­cated by a light­bulb icon

Weighted av­er­age cost per Intelligence Index task. Each eval­u­a­tion’s cost is cal­cu­lated from in­put, cache hit, cache write, rea­son­ing, and an­swer to­ken prices, di­vided by task count, and weighted by its Intelligence Index weight.

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all eval­u­a­tions in the Artificial Analysis Intelligence Index

Reasoning mod­els are in­di­cated by a light­bulb icon

The cost to run the eval­u­a­tions in the Artificial Analysis Intelligence Index, cal­cu­lated us­ing the mod­el’s in­put, cache hit, cache write, rea­son­ing, and an­swer to­ken prices and the num­ber of to­kens used across eval­u­a­tions (excluding re­peats).

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

Reasoning mod­els are in­di­cated by a light­bulb icon

Price per to­ken for cached prompts (previously processed), typ­i­cally of­fer­ing a sig­nif­i­cant dis­count com­pared to reg­u­lar in­put price, rep­re­sented as USD per mil­lion to­kens. The val­ues shown here are the cache hit price; cache write and cache stor­age are billed sep­a­rately and vary by provider — see Cache pric­ing by provider” for de­tail.

Context Window

Context Window

Context win­dow: to­kens limit · Higher is bet­ter

Reasoning mod­els are in­di­cated by a light­bulb icon

Larger con­text win­dows are rel­e­vant to RAG (Retrieval Augmented Generation) LLM work­flows which typ­i­cally in­volve rea­son­ing and in­for­ma­tion re­trieval of large amounts of data.

Model Size (Open Weights Models Only)

Model Size: Total and Active Parameters

Comparison be­tween to­tal model pa­ra­me­ters and pa­ra­me­ters ac­tive dur­ing in­fer­ence

Reasoning mod­els are in­di­cated by a light­bulb icon

The to­tal num­ber of train­able weights and bi­ases in the model, ex­pressed in bil­lions. These pa­ra­me­ters are learned dur­ing train­ing and de­ter­mine the mod­el’s abil­ity to process and gen­er­ate re­sponses.

GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

github.com

WASTE — Weight-Aware Streaming Tensor Engine

Kimi K3 — 2.78 tril­lion pa­ra­me­ters — run­ning on a con­sumer lap­top.

$ waste run ~/models/k3.waste What is the cap­i­tal of Italy?’ waste: no –budget, us­ing 46.24 GB of 64.00 GB (expert cache 17.56 GB) The cap­i­tal of Italy is **Rome**. [16 to­kens, 25.78 s, 0.62 tok/​s | ex­perts 9038 hit / 14514 miss = 38%]

WASTE is an em­bed­d­a­ble in­fer­ence en­gine writ­ten in C, with no third-party run­time de­pen­den­cies. It keeps the model trunk in mem­ory, streams se­lected ex­perts di­rectly from disk, and uses the re­main­ing RAM as a bounded ex­pert cache.

Its cur­rent proof point is the com­plete open-weights Kimi K3 model: 2.78 tril­lion pa­ra­me­ters, con­verted into a 982 GiB con­tainer and run­ning on a 64 GB MacBook Pro at 0.45 – 0.62 to­kens per sec­ond. This is not a dis­tilled, pruned, or re­duced vari­ant.

WASTE was writ­ten for that one model and that one con­straint: K3 does not fit in the RAM of cur­rent main­stream con­sumer sys­tems. It is 1.42 TB as pub­lished and 982 GB af­ter con­ver­sion. But a mix­ture of ex­perts ac­ti­vates about 4% of it­self per to­ken, so al­most all of that weight is idle at any in­stant — and idle weight does not need to be in mem­ory, it needs to be reach­able in time. WASTE keeps it on disk in a lay­out where one ex­pert costs ex­actly one read, streams what each to­ken ac­tu­ally needs, and spends every re­main­ing byte of RAM on the part that re­peats.

Where this stands

The en­gine is cor­rect: every layer is val­i­dated against a PyTorch ref­er­ence, the fi­nal log­its agree to 3.6e-06, and the vi­sion tower matches its own or­a­cle to 2.3e-06. It is also slow — half a to­ken per sec­ond, twenty-six sec­onds for the sen­tence above.

Both of those mat­ter, and the sec­ond one should not be read as a dis­claimer. We are not aware of an­other pub­lished demon­stra­tion of a model this size stream­ing from disk on a con­sumer ma­chine: we found none for tril­lion-scale NVMe stream­ing, and the best-doc­u­mented 671B-class recipes as­sume a server with a ter­abyte of DDR5. That is a re­port of what our search turned up rather than a sur­vey — this repos­i­tory car­ries no bib­li­og­ra­phy and no com­par­i­son table, so read it as an in­vi­ta­tion to send a counter-ex­am­ple, not as a re­sult. The in­ter­est­ing part is not the speed, it is that the whole thing is in the reach­able range on a sin­gle con­sumer ma­chine — and that from here the ques­tion is en­gi­neer­ing rather than fea­si­bil­ity.

Where the levers were is not where they are. The two that looked biggest — read­ing fewer bytes per to­ken, and keep­ing more of them in RAM — were both mea­sured and both re­fused: this fam­i­ly’s router has no tail to de­mote, and a cache the ma­chine will not leave res­i­dent can­not be bought at any price. What paid in­stead was never about which bytes to read but when. Overlapping the ex­pert reads with the arith­metic is worth ~1.6x; start­ing the next lay­er’s reads on its own router’s guess, one resid­ual early, takes the hit rate from 14% to 38% at no ex­tra bytes at all.

Both of those are ex­act — the cache sta­tis­tics and the log­its are un­changed — which is the prop­erty that makes them ship­pable rather than tun­ing. docs/​EF­FI­CIENCY.md is the ac­count of how each lever was priced, in­clud­ing the three that were built be­fore be­ing mea­sured and the two that were then taken back out.

What that opens up, con­cretely: a fron­tier-scale model that an­swers with no net­work, no per-to­ken in­voice, and noth­ing leav­ing the ma­chine — which is the dif­fer­ence be­tween you may not send that data to an API and run it here”. The for­mat and the en­gine are not K3-specific in any deep way; K3 is sim­ply the hard­est case that ex­ists to­day, and a model that streams at 2.78T streams com­fort­ably at 48B.

Every num­ber in this doc­u­ment was mea­sured on the com­mit it is pub­lished with, and the ones that were wrong are recorded as wrong in docs/​LEARNED.md rather than qui­etly cor­rected.

Why the name

Every to­ken an­swered by a cloud ser­vice is paid for twice: once on the in­voice, and once in the elec­tric­ity of a dat­a­cen­ter run­ning a model that would fit — barely, awk­wardly, but gen­uinely — on hard­ware al­ready sit­ting on a desk. WASTE means to be the first con­crete step to­ward end­ing that waste of to­kens. The acronym came sec­ond.

What you need

Sizes here are pow­ers of two, the way df and the en­gine both re­port them: the con­tainer is 982 GiB, which a disk ven­dor would call 1.05 TB.

The RAM floor is what the en­gine re­fuses to start be­low, and it is al­most en­tirely the 27.28 GB res­i­dent trunk. Useful through­put starts higher: on a 64 GB ma­chine the en­gine gives it­self a 46 GB bud­get, of which 17.56 GB is ex­pert cache, and that is the top of the mea­sured curve. A 32 GB ma­chine can tech­ni­cally open the model and will page badly; treat 64 GB as the real re­quire­ment.

Storage speed is not a de­tail. A to­ken reads 17 GB of ex­perts. On the in­ter­nal SSD that is 12.78 GB/s and the model streams; over a USB en­clo­sure it is 0.94 GB/s and the same to­ken takes thir­teen sec­onds. Convert onto in­ter­nal NVMe, and use the ex­ter­nal disk for the down­load only.

If a ter­abyte is not avail­able, the same en­gine and the same for­mat run Kimi-Linear-48B-A3B-Instruct from a 19 GB con­tainer with a 1.87 GB floor, at 10.7 tok/​s. That is the good path for try­ing WASTE out be­fore com­mit­ting a disk to K3.

What it is

Self-contained. One lib­waste.a, one waste bi­nary, noth­ing at run time be­yond libc and pthreads.

Zero de­pen­den­cies. No BLAS, no ONNX, no Python in the in­fer­ence path, noth­ing to in­stall. The Python un­der tools/ con­verts mod­els and val­i­dates the en­gine; it never runs along­side it.

Fully em­bed­d­a­ble. Twenty-six pub­lic func­tions in src/​waste.h: open a model un­der a RAM ceil­ing, gen­er­ate, save the ses­sion, close. The CLI is a client of that API and touches noth­ing pri­vate — if the CLI can do it, so can an em­bed­ding host.

waste_cfg cfg; waste_cfg_init(&cfg); cfg.ram_bud­get_bytes = 46ULL << 30; /* a hard ceil­ing, not a hint; 0 sizes it to this ma­chine */

waste_ctx *ctx; if (waste_open(“/path/to/k3.waste”, &cfg, &ctx) != WASTE_OK) re­turn 1; waste_­gen­er­ate(ctx, ids, n, &params, on_­to­ken, user); waste_­close(ctx);

The path is the con­tainer di­rec­tory the con­verter wrote — no ~ ex­pan­sion here, that is the shel­l’s job.

How it works

Placement de­cides the speed

A model is con­verted once into a .waste con­tainer: a JSON man­i­fest, a res­i­dent trunk, and one ex­pert bank per layer. Each ex­pert record is 4 KiB-aligned with its gate, up and down ma­tri­ces ad­ja­cent, so rout­ing to an ex­pert costs ex­actly one pread — not three, not a seek per ma­trix. The arith­metic was never the bot­tle­neck.

Reads by­pass the page cache (F_NOCACHE on ma­cOS, O_DIRECT on Linux, FILE_FLAG_NO_BUFFERING on Windows). That is de­lib­er­ate: with a con­tainer smaller than RAM the ker­nel would cache every­thing, and the hit rates mea­sured that way are a fic­tion that does not sur­vive con­tact with a 982 GB model.

Every record’s header is checked on the way in — right magic, the ex­pert the in­dex asked for, off­sets that fit — so a bank that has been trun­cated or spliced stops the gen­er­a­tion and names the record in­stead of an­swer­ing from the wrong bytes. That costs noth­ing mea­sur­able. The record also car­ries a cr­c32 over its pay­load, and check­ing that is –verify, off by de­fault: it is a pass over every record on every cache miss, about 5% on Kimi-Linear and 1% on K3. Worth it for a con­tainer you copied or down­loaded and have not read since; not worth it on every to­ken of one you con­verted your­self. See docs/​FOR­MAT.md.

Reading ahead, and read­ing be­fore the router has spo­ken

A layer knows all six­teen of its ex­pert ids the mo­ment its router runs, so the reads go out on their own threads and the arith­metic con­sumes them as they land in­stead of block­ing on each. That is worth ~1.6x on K3, and the cache sta­tis­tics are iden­ti­cal to the digit with it on and off — the en­gine does the same work, it just stops wait­ing for it.

The next lay­er’s ids are not known: its router eats a hid­den state that does not ex­ist yet. But the router does ex­ist, and it is res­i­dent. So at the end of each layer, once its own reads are con­sumed and the disk is about to go idle through the next lay­er’s at­ten­tion, the en­gine runs layer L+1′s router on layer L’s hid­den state and starts fetch­ing the six ex­perts it names. One resid­ual early, that guess is right 92% of the time at rank 1 and 81% over the first six.

It is ex­act by con­struc­tion: the real router still de­cides, and the guess only de­cides when bytes move. The de­mand hit rate goes from 14% to 38% and the to­tal bytes read do not change — those records were go­ing to be read any­way. WASTE_LOOKAHEAD=0 turns it off.

The same trick in the pre­fill path was built and re­moved. A de­code layer claims 16 cache slots so the spec­u­la­tive records sur­vive; a chunk layer claims about 550, evicts them be­fore use, and reads them twice — 6.9% more bytes for no time saved. docs/​LEARNED.md §34 – 36.

Three bits per ex­pert weight

Experts are stored as resid­ual vec­tor quan­ti­za­tion — three stages of 256-entry code­books over 8-dimensional vec­tors, 3.00 bits per weight — and the ma­trix is never ma­te­ri­al­ized. For each to­ken the en­gine builds a table of par­tial dot prod­ucts, one per code­book en­try per vec­tor po­si­tion, af­ter which every ex­pert row is three table reads and two adds.

The trunk stays at 4 and 8 bits. The model was trained with quan­ti­za­tion-aware train­ing on the ex­perts only, so it has no trained tol­er­ance for a squeezed trunk: a 3-bit trunk was built and mea­sured, the cache pre­dic­tion held, the through­put did not, and the out­put col­lapsed.

The cache floor is one to­ken’s work­ing set

The most pre­dic­tive num­ber in this pro­ject. K3 touches 16 ex­perts in each of 92 lay­ers per to­ken: 17.0 GB. Below that, an ex­pert cached for one to­ken is evicted be­fore the next to­ken asks for it, and the hit rate is not low — it is zero.

What cross­ing it buys has changed, though, and the table be­low is the first one to show it. Going from a 0% hit rate to 17% is worth about 8% of through­put now — 0.50 to 0.54 — be­cause read-ahead al­ready hides most of the I/O the cache would have saved. The sharp bend in this curve is no longer the climb above the floor; it is the col­lapse above 46 GB, where the en­gine stops fit­ting in the ma­chine.

The hit-rate col­umn pre­dates the router looka­head, which roughly dou­bles it at every bud­get — 14% to 38% at 46 GB. It does not move the de­code colum­n’s shape, be­cause what col­lapses the 52 and 58 GB rows is the ma­chine run­ning out of mem­ory and not the cache miss­ing.

Ranges, not mea­sure­ments, and the width is the find­ing. Every run be­hind a row re­ports cache sta­tis­tics iden­ti­cal to the digit — the en­gine does the same work each time — so what varies is the ma­chine, not the en­gine.

The two rows that fit are tight. 32 and 46 GB re­pro­duce to within a few per­cent, be­cause the en­gine’s whole foot­print fits with room to spare and noth­ing has to be taken from any­thing.

52 GB has no value. Two runs of the de­fault con­fig­u­ra­tion gave 0.04 and 0.15; three more with the trunk wired gave 0.46, 0.19 and 0.03 — seven-fold in the col­umn above and fif­teen-fold across both con­fig­u­ra­tions, against 3652 hit / 8124 miss every sin­gle time. That bud­get sits ex­actly where the en­gine’s foot­print ei­ther does or does not fit be­side what­ever else the ma­chine is hold­ing, and which side it lands on is de­cided be­fore the process starts. A row that spans 15x is not a slow row; it is a row whose mean would in­vite a com­par­i­son there is noth­ing to com­pare.

58 GB is uni­formly bad and re­pro­ducibly so.

Order still mat­ters, and more than the table shows. Re-run af­ter the 52 and 58 GB rows have dri­ven the ma­chine into pag­ing, 46 GB col­lapses — 0.02 tok/​s in one such run — while again re­port­ing iden­ti­cal counts. Sweep up­ward, one bud­get per quiet ma­chine, and treat any­thing mea­sured af­ter a pag­ing row as void.

Read-ahead made the rows that fit faster and left the oth­ers where they were, so the step is larger than when this was first mea­sured: 46 GB went 0.32 to 0.54. Wiring the res­i­dent trunk with WASTE_MLOCK=trunk does not move it ei­ther — 32 and 46 GB are un­changed, 58 GB stays hope­less, and 52 GB has no value to change. docs/​LEARNED.md §30 – 33.

Everything in the mem­ory de­sign ex­ists to get above that line, which is why the en­gine works to free RAM rather than to save it.

And there is a ceil­ing on the other side, closer than it looks. Read that table twice: the hit rate climbs all the way down. At 58 GB on a 64 GB ma­chine the cache serves 39% of ex­perts from RAM and the en­gine is twenty times slower than at 46 GB, where it serves 17%. The en­gine is in­side its bud­get; the ma­chine is not, so the OS pages out the ex­pert cache, and a hit” be­comes a page fault in­stead of the disk read the en­gine was man­ag­ing.

So the us­able win­dow is nar­row. It opens at ~46 GB, where the cache fi­nally clears one to­ken’s work­ing set, and it has al­ready closed by 52 — on an oth­er­wise idle ma­chine, with 49 GB free be­fore the run. It is also sharp enough to move un­der a change that looks un­re­lated: tak­ing 1.11 GB of em­bed­ding table off the res­i­dent set fed straight into the cache at a fixed bud­get, and on the build of the day that was enough to push 58 GB from 0.32 tok/​s to 0.04.

So the de­fault does not fill the ma­chine. Expert cache is only worth any­thing in whole mul­ti­ples of that work­ing set, and the re­main­der above a mul­ti­ple buys a few points of hit rate while push­ing the ma­chine to­wards pag­ing. When it picks a bud­get for it­self the en­gine steps down a whole work­ing set at a time and takes the largest that fits un­der seven eighths of RAM: K3 asks for floor + — 80.63 GB — and gets floor + on this lap­top, a 46 GB bud­get and a 17.56 GB cache. That is the top of the curve above, reached with no flag. A 128 GB ma­chine still gets the full .

An ear­lier ver­sion took every byte up to the cap in­stead, which put a 27 GB cache on this ma­chine — be­tween two bud­gets mea­sured at 0.11 and 0.04 tok/​s. The real les­son is that a cache you do not con­trol is not a cache, and the corol­lary is that an en­gine should stop ask­ing for mem­ory be­fore the OS starts tak­ing it back.

Linear at­ten­tion, and an ab­sorbed KV cache

K3′s at­ten­tion is a 3:1 hy­brid: Kimi Delta Attention, which car­ries a fixed-size re­cur­rent state in­stead of a grow­ing KV cache, and gated multi-head la­tent at­ten­tion. The MLA lay­ers cache the 512-wide la­tent rather than ex­panded per-head keys and val­ues, with kv_b_proj ab­sorbed into the query and the out­put:

q_nope · (W_kb c) == (W_kbᵀ q_nope) · c Σ_s a_s (W_vb c_s) == W_vb (Σ_s a_s c_s)

Identical log­its to 1.2e-05, and 53× less cache: 11.25 GB be­comes 0.21 GB at 4K con­text. It is also what makes long con­text pos­si­ble at all — the ex­panded lay­out wants 360 GB at 128K to­kens, the la­tent one 7.2.

Performance and mem­ory

MacBook Pro M5 Pro, 64 GB, con­tainer on the in­ter­nal SSD. Every fig­ure was mea­sured on the com­mit it is pub­lished with.

Kimi K3 — 2.78T pa­ra­me­ters, 982 GB con­tainer

The floor is al­most en­tirely the res­i­dent trunk. Useful through­put starts above ~46 GB, where the ex­pert cache fi­nally clears one to­ken’s work­ing set, and is gone again by 52, where the ma­chine starts pag­ing. Below the first line ex­tra RAM buys noth­ing; above the sec­ond it costs, badly. The win­dow is one bud­get wide on this ma­chine.

The tower is not what an im­age costs. Encoding 1024 patches takes 15.7 s; the 256 po­si­tions it pro­duces then go through the 92 MoE lay­ers like any other to­ken, which is the other 731 s. An im­age is priced as text of the same length, so the patch bud­get in vi­sion.json is a real dial: halv­ing the grid halves the prompt.

Kimi-Linear — 48B pa­ra­me­ters, 19 GB con­tainer

The same en­gine and the same for­mat, on a model that fits com­fort­ably. This is what WASTE looks like when it is not fight­ing.

Where the time goes

Decode on K3, 17.32 GB of cache and still cold — 6.7% hit over ten steps, which is the state a fresh prompt starts in:

Reproduce with WASTE_PROFILE=1 WASTE_LOOKAHEAD=0 WASTE_CACHE_MB=17735 ./test_forward MODEL 1008,10484,318,15383,387 out.bin 5. The looka­head is off there on pur­pose: this is the cold-cache shape, and with it on the hit rate is 38% and the I/O share cor­re­spond­ingly lower. The I/O share also falls as the cache warms, so a long ses­sion sits un­der this ei­ther way; the rank­ing does not change.

The I/O al­ready runs near the hard­ware limit — 17.0 GB per to­ken at ~9.9 GB/s against the SSDs mea­sured 12.78 — so it only gets cheaper by hap­pen­ing less of­ten. For a long time that read as which means cache, which means RAM, and it was half right: the other half is when it hap­pens. Overlapping the reads with the arith­metic and start­ing the next lay­er’s on its own router’s guess be­tween them cost no RAM at all. What fol­lows is the mem­ory half of that story.

Getting started

git clone https://​github.com/​sqliteai/​waste && cd waste make # lib­waste.a, waste, lib­wastevq make check # 23 pass, 11 skip on a fresh clone

No con­fig­ure step and no de­pen­dency res­o­lu­tion. make check needs no model: it builds a small syn­thetic con­tainer and runs the en­gine against it. The eleven skips are the checks that need some­thing a clone does not carry — the PyTorch or­a­cle, the round-trip against the source shards, any­thing dri­ving the CLI with text, since the syn­thetic con­tainer car­ries no to­k­enizer, and the K3 checks, which want the con­tainer and the re­lease on disk. With both con­tain­ers pre­sent the suite is 36 checks.

Converting Kimi K3

Conversion is the one step that needs Python, and it hap­pens once. The source is moon­shotai/​Kimi-K3 ex­actly as pub­lished — 96 safeten­sors shards, 1.42 TB, noth­ing patched:

# 1. pre­flight: reach­able? how big? does it fit? tools/​fetch_weights.sh –dest /Volumes/staging/k3 –dry-run

# 2. down­load — re­sum­able, safe to kill, safe to re-run tools/​fetch_weights.sh –dest /Volumes/staging/k3

# 3. con­vert into a con­tainer uv run –with torch –with safeten­sors python tools/​con­vert.py \ –src /Volumes/staging/k3 \ –out ~/models/k3.waste –jobs 3

That pro­duces the 982 GB con­tainer every num­ber above was mea­sured on. It takes about 4.7 hours with three processes on the M5 Pro (23.7 with the pure-torch en­coder — see docs/​K3.md), and wants ~1.0 TB free on the tar­get vol­ume. The con­verter is re­sum­able too: a layer whose bank is al­ready writ­ten is skipped, so an in­ter­rupted run costs only the layer it was in the mid­dle of.

The down­load is the part that goes wrong. A 1.42 TB pull over hours will hit dropped con­nec­tions, CDN 5xx and at least one in­ter­rupted run, so every shard re­sumes mid-file rather than restart­ing, re­tries with ex­po­nen­tial back­off and jit­ter, and counts as done only when its size matches Content-Length — recorded in a state file, so a re-run skips fin­ished shards with­out even a HEAD re­quest. –check re-ver­i­fies every­thing on disk against the re­mote and down­loads noth­ing (96 shards in 34 s). –repo points it at an­other model, HF_TOKEN at a gated one. ma­cOS and Linux.

Give –dest a stag­ing disk rather than the vol­ume that will hold the con­tainer. The shards are read once, by the con­verter; the con­tainer is read con­tin­u­ously, at every to­ken. On this ma­chine the ex­ter­nal en­clo­sure mea­sures 0.94 GB/s against the in­ter­nal NVMe’s 12.78 — see docs/​GATES.md, Gate H — which is the dif­fer­ence be­tween a model that streams and one that stalls.

tools/​pipeline.sh chains the whole thing un­at­tended — down­load, con­vert, round-trip the con­tainer against the source weights, gen­er­ate, then diff the log­its against the PyTorch or­a­cle — and leaves a re­port next to the con­tainer. The same con­verter han­dles the other mem­ber of the fam­ily, Kimi-Linear-48B-A3B-Instruct, into the 19 GB con­tainer of the sec­ond bench­mark; –src is the only thing that changes.

Pre-converted con­tain­ers are on their way to hug­ging­face.co/​sqliteai, at which point this whole sec­tion be­comes a down­load and the Python is only needed for mod­els we have not pub­lished.

Running it

The con­tainer is the di­rec­tory the con­verter wrote, so give it that path — ~/models/k3.waste through­out this README:

waste run ~/models/k3.waste The cap­i­tal of France is” -n 32 waste chat ~/models/k3.waste # multi-turn, state kept waste eval ~/models/k3.waste 2 + 2 =” –top-k 5 # next-to­ken dis­tri­b­u­tion waste plan ~/models/k3.waste –budget 46G # what fits, what does not echo prompt” | waste run ~/models/k3.waste # stdin works too

-n is a cap, not a re­quire­ment: with­out it gen­er­a­tion stops at the con­tain­er’s end-of-se­quence to­ken or at 128 to­kens, whichever comes first. The ex­am­ples pass it be­cause 128 to­kens of K3 is six min­utes.

–budget is op­tional, and leav­ing it out is the right de­fault rather than a fall­back: the en­gine takes the con­tain­er’s rec­om­men­da­tion, steps it down a whole to­ken work­ing set at a time un­til it fits un­der seven eighths of phys­i­cal RAM, and never goes be­low the floor — a bud­get you set ex­plic­itly un­der the floor is re­fused rather than swapped into. It then says on stderr what it landed on, so the same com­mand on two ma­chines is not silently two dif­fer­ent runs:

waste: no –budget, us­ing 46.24 GB of 64.00 GB (expert cache 17.56 GB)

–verify checks each ex­pert record’s cr­c32 as it comes off the disk. It is off by de­fault, and that is a through­put de­ci­sion rather than a claim that con­tain­ers do not rot: it is a pass over every record on every cache miss, about 5% on Kimi-Linear and about 1% on K3, where the read dom­i­nates. Turn it on once for a con­tainer you copied, down­loaded, or left on a disk you do not trust, and for any­thing whose wrong an­swers would be be­lieved; leave it off for one you con­verted your­self and have been read­ing since. WASTE_VERIFY=1 in the en­vi­ron­ment does the same thing, and the server takes –verify as well. Any of them turns it on; none of them turns it off.

What is checked ei­ther way: a short read, and a record header that does not de­scribe the ex­pert the bank in­dex asked for. Those are O(1), they cost noth­ing mea­sur­able, and they are what keeps a dam­aged off­set out of the arith­metic — –verify only adds the pass over the pay­load.

waste –help lists all nine com­mands. –json makes eval, to­k­enize, plan, info and bench ma­chine-read­able.

Serving it

serve/ is an OpenAI-compatible HTTP server — the sec­ond client of the pub­lic API, along­side the CLI, reach­ing the same en­gine through ctypes rather than keep­ing a copy of the model code in Python:

make lib­waste.dylib # or lib­waste.so on Linux python3 -m serve ~/models/k3.waste –port 8000

curl lo­cal­host:8000/​v1/​chat/​com­ple­tions \ -H Content-Type: ap­pli­ca­tion/​json’ \ -d {“model”:“k3″,“messages”:[{“role”:“user”,“content”:“Why is the sky blue?“}]}’

/v1/chat/completions (streaming and not), /v1/completions, /v1/models, /health. It car­ries the whole of K3′s prompt for­mat, not the four-string sub­set a con­tain­er’s chat.json can hold: tool de­f­i­n­i­tions and tool re­sults, typed call ar­gu­ments, JSON re­sponse schemas, tool_­choice, the think chan­nel and think­ing_­ef­fort, and im­ages — plus the parser that reads the re­ply back into rea­son­ing, an­swer and tool_­calls. Stdlib only.

The prompt ren­derer is a port of en­cod­ing_k3.py from the re­lease, and the test suite checks it against that file seg­ment for seg­ment on a cor­pus of 38 con­ver­sa­tions when­ever the weights di­rec­tory is on disk. docs/​SERVE.md is the ref­er­ence.

Images

K3 is mul­ti­modal — a 401M ViT, 27 lay­ers, patch 14 — and so is the en­gine. –image at­taches a pic­ture; re­peat it for sev­eral:

How to Exist

www.raptitude.com

Here’s an ex­per­i­ment for a true dare­devil.

Sit there for a three min­utes, fol­low­ing two rules:

Don’t do any­thing.

Be con­tent.

By don’t do any­thing,” I mean don’t move, don’t fid­get, don’t in­dulge any thoughts or day­dreams. You’re al­lowed to breathe, and blink.

By be con­tent,” I mean be com­pletely okay with your ex­pe­ri­ence of do­ing noth­ing. Don’t try to change any­thing, and don’t get im­pa­tient with what’s hap­pen­ing. Be com­pletely okay for three min­utes.

Try that now. See how long it takes be­fore you’re dy­ing for it to be over.

It’s oddly dif­fi­cult to do noth­ing, and while you’re do­ing noth­ing, it’s oddly dif­fi­cult to feel at ease. There’s such a strong urge to do some­thing: look around the room, re­hash a con­ver­sa­tion, ex­plore your in­cisors with your tongue, wig­gle your toes, any­thing. When you stop do­ing every­thing and just ex­ist, you al­most feel like you’re dy­ing.

This is a crazy thing to no­tice af­ter hav­ing been alive so many years — that your ex­is­tence it­self is so much to bear. There’s al­ways some­thing wrong, even when every­thing’s fine. It’s as if you can only bear the pre­sent mo­ment when you’re try­ing change it into some­thing else.

This is the strange con­di­tion of the hu­man be­ing. It’s al­ler­gic to its nat­ural habi­tat, which is the pre­sent mo­ment. In or­der to cope with this al­lergy, it per­pet­u­ally seeks things: feel­ings and ex­pe­ri­ences that are not yet pre­sent. It wants to al­ways be get­ting the hell out of here.

You might think that you’re free from this prob­lem some­times, at least in those mo­ments when you get the thing you’re seek­ing. Say you’re fi­nally eat­ing the cookie-dough ice cream flurry you looked for­ward to all day. If you pay close at­ten­tion as you eat it, you’ll no­tice that you want to move past this mo­ment too. Lingering on any one spoon­ful too long be­comes un­bear­able. There’s a pow­er­ful drive to go on to the next one. That’s why you or­dered a Large.

This most fun­da­men­tal prob­lem of hu­man life is so easy to over­look be­cause our en­tire lives are made of the cop­ing strat­egy. So much of what we seek is solely to flee the ex­pe­ri­ence of be­ing here. People buy things they don’t need, start fights with their part­ners, eat when they’re not hun­gry, and scroll mis­er­able and inane con­tent, just to es­cape the feel­ing of ex­is­tence as it al­ready is.

Notice the pow­er­ful urge to sip your drink or fid­dle with some­thing when the con­ver­sa­tion dies at a din­ner party. Or how quickly your phone comes out when there’s an un­ex­pected wait. Existence with­out do­ing is bru­tal!

Each year, roughly 100 fire­fight­ers are con­victed of ar­son in North America, of­ten on mul­ti­ple counts. Most of­ten they are young, new fire­fight­ers, frus­trated by the lack of ac­tion.

And hi­lar­i­ously, from a 2014 study on do­ing noth­ing:

In 11 stud­ies, we found that par­tic­i­pants typ­i­cally did not en­joy spend­ing 6 to 15 min­utes in a room by them­selves with noth­ing to do but think, that they en­joyed do­ing mun­dane ex­ter­nal ac­tiv­i­ties much more, and that many pre­ferred to ad­min­is­ter elec­tric shocks to them­selves in­stead of be­ing left alone with their thoughts.

In 11 stud­ies, we found that par­tic­i­pants typ­i­cally did not en­joy spend­ing 6 to 15 min­utes in a room by them­selves with noth­ing to do but think, that they en­joyed do­ing mun­dane ex­ter­nal ac­tiv­i­ties much more, and that many pre­ferred to ad­min­is­ter elec­tric shocks to them­selves in­stead of be­ing left alone with their thoughts.

Even think­ing is of­ten a sneaky way of es­cap­ing the ex­is­tence; ru­mi­na­tion is­n’t so much about try­ing to solve your prob­lems, as it is about go­ing else­where in your mind to es­cape anx­i­ety and un­cer­tainty. Apparently, elec­tric shocks work even bet­ter.

How to be­come more com­fort­able with ex­is­tence

You can de­velop the abil­ity to ex­ist a lot more com­fort­ably. You do it by prac­tic­ing ex­ist­ing, a few sec­onds at a time, with­out try­ing to change any­thing about how ex­is­tence feels right now.

Basically you sit, do noth­ing, and no­tice how it feels to do noth­ing. (It will prob­a­bly feel sub­tly weird and un­set­tled.) You then see if you can com­pletely em­brace these feel­ings, with­out the usual squirm­ing and look­ing else­where. But you’ll do it only for a few sec­onds at a time, us­ing your breath as a mea­sur­ing stick.

Here’s how to do it with­out feel­ing over­whelmed:

(If you have PTSD or any other psy­chi­atric dis­or­der, check with a pro­fes­sional be­fore you do this.)

Sit, eyes open or closed, and re­lax your body as com­pletely as pos­si­ble. Take long, easy breaths. Relax every mus­cle you can. Just do your best.

When you’re ready, breathe in while keep­ing the body sup­ple like that. Open to every feel­ing that oc­curs dur­ing the in­breath: tin­gling, weird­ness, un­set­tled­ness, doubt, what­ever. Let the whole ex­pe­ri­ence wash over you like warm surf, for that few sec­onds it takes to in­hale.

Let it go. Give your­self a mo­ment.

When you’re ready, do it again: em­brace the en­tire ex­pe­ri­ence of one in­breath. Just let the whole ex­pe­ri­ence hap­pen to you — no need to study it, or fig­ure it out. Just em­brace the whole bou­quet of feel­ings, for the few sec­onds it takes. Push away noth­ing, just for that few sec­onds of breath­ing in.

Once you can do that de­cently well (no need to be per­fect), try the same thing but with an out­breath. Stay re­laxed and open through­out the length of one whole out­breath. No de­fend­ing, no tens­ing. Be a hu­man pud­dle.

If you get dis­tracted or fraz­zled, or you do tense up, that’s okay. Take a few breaths off to rec­ol­lect your­self. Then try again. You have in­fi­nite breaths to try this with.

Repeat this process, one half-breath (an in­hale or an ex­hale) at a time. The half-breath is a small enough span of time that you can usu­ally stay open for the 5 – 10 sec­onds it takes. Once you can do it on both an in­breath and an out­breath, see if you can start string­ing them to­gether, stay­ing open through­out the whole breath­ing cy­cle.

Do this for five min­utes at first, in­clud­ing any breaks. Then see what hap­pens when you do it longer, and with fewer gaps. Basically you’ll be rest­ing — just ex­ist­ing and breath­ing — in that non-de­fen­sive state.

Even af­ter one ses­sion of this, you might no­tice you can re­lax a lit­tle more eas­ily, no mat­ter what’s hap­pen­ing.

This is a form of med­i­ta­tion, but I al­most want to avoid that word be­cause it makes peo­ple get ner­vous and over­com­pli­cate it. Just think of this prac­tice as ex­ist­ing with­out fear, for a few sec­onds at a time. Minimum ef­fec­tive dose is one half-breath.

Naturally, if you can learn to calm your al­lergy to ex­is­tence a bit, life gets eas­ier in nearly every sit­u­a­tion. Ordinary ex­pe­ri­ences like wait­ing in line, feel­ing un­cer­tain, be­ing a bit too warm or cold, or not be­ing sure what to do with your­self, be­come much more tol­er­a­ble. (And prob­a­bly most of life con­tains this sort of mi­nor dis­com­fort.)

Regular prac­tice keeps your al­lergy symp­toms mild. Neglecting it makes them come back.

In par­tic­u­lar, you might no­tice much less of a need to en­ter­tain or dis­tract your­self. Escape-driven habits like doom­scrolling, ran­dom snack­ing, nail-bit­ing, (arson?), and ru­mi­na­tion be­come less mag­netic. When plain old ex­is­tence feels okay, there’s so much you no longer need to do.

***

Because We Can

weeraman.com

Twenty-five years ago, I made my first do­na­tion to an open source pro­ject and pur­chased a CD with an op­er­at­ing sys­tem as down­load­ing a few hun­dred megabytes over a 14.4kbps dial-up was­n’t very fun. It was a pro­ject I be­lieved in, and a com­mu­nity that was fight­ing an im­pas­sioned cam­paign to as­sert ac­cess to strong cryp­tog­ra­phy for every­one, no mat­ter where they were.

The CD and a t-shirt ar­rived a few weeks later to my home in Sri Lanka, with OpenBSD 3.0. The t-shirt fea­tured the iconic puffer fish on the front. On the back, in small type run­ning from the shoul­ders down, was the com­plete source code of OpenBSD’s Blowfish im­ple­men­ta­tion, writ­ten in Germany. Written in the United States, it would have been clas­si­fied as a weapon.

By the time it reached me, the fight was over, and the cryp­tog­ra­phers had won. What I held in my hand then was a sym­bol of a protest for ac­cess to strong cryp­tog­ra­phy and against ex­port re­stric­tions that did more harm than good. Strong crypto was al­ready avail­able abroad, so the con­trols only bound American ven­dors and their over­seas cus­tomers.

Today the re­flex is back. The fears have changed. The worry is now cy­ber ca­pa­bil­ity, bi­ol­ogy and mod­els that do things no­body asked them to do. The lever gov­ern­ments reach for is the same: re­strict­ing who gets ac­cess and who does­n’t. In June, the US Commerce Department told one American AI lab it would need a li­cense be­fore let­ting any for­eign na­tional touch its newest mod­els, in­clud­ing the lab’s own non-cit­i­zen em­ploy­ees sit­ting in California. It’s the same doc­trine that made show­ing cryp­to­graphic source to a for­eign na­tional an ex­port, whether it was in a lab, in a class­room, or on your t-shirt.

Not all of the worry is the­atre. Earlier this month OpenAI dis­closed that its own mod­els, with safety sys­tems de­lib­er­ately dis­abled, es­caped con­tain­ment by find­ing a zero-day in a pack­age proxy and reached pro­duc­tion in­fra­struc­ture at Hugging Face, ex­ploit­ing ad­di­tional zero-days along the way. Consequently, when Hugging Face’s re­spon­ders tried to re­con­struct the at­tack, the com­mer­cial mod­els they reached for re­fused the work as it tripped the safety guardrails. They fin­ished the in­ves­ti­ga­tion on GLM 5.2, a Chinese open-weight model, run­ning on their own hard­ware. A de­ter­mined at­tacker is not bound by us­age poli­cies. The de­fend­ers are. Restrictions writ­ten for safety are mak­ing de­fend­ers less safe.

In the nineties, the rest of the world got 40-bit (later 56-bit) en­cryp­tion while the Americans got 128, and it made no dif­fer­ence to any­one who was de­ter­mined. The con­trols bound the law-abid­ing and no­body else. That is the asym­me­try. The de­ter­mined will have the fron­tier. The rest of us are asked to go with­out, and told it is for our safety.

The OpenBSD team did­n’t work around the ex­port con­trols. They arranged the pro­ject so that the con­trols could­n’t reach it. Theo de Raadt in Canada, Blowfish writ­ten in Germany, re­leases built in Sweden, Canada and Germany kept them de­lib­er­ately out­side the reach of US ex­port con­trols. The pro­ject openly asked non-Amer­i­can cryp­tog­ra­phers to come and help, and American de­vel­op­ers, as the story goes, would cross the bor­der to Canada to work on the sys­tem and bring the re­sults home legally. Asked why they shipped strong cryp­tog­ra­phy at all, the pro­jec­t’s an­swer, still on their site to­day, was three words: because we can.”

The same arrange­ment is be­ing made now, at a na­tional scale. Mistral, DeepSeek, Moonshot and Zhipu pub­lish weights that, once down­loaded, no ex­port let­ter can re­call. The sov­er­eignty ar­gu­ment that used to live in Brussels think tanks is now gov­ern­ment pol­icy, ac­cel­er­ated by watch­ing ac­cess to a fron­tier model with­drawn world­wide by let­ter.

More than twenty-five years ago, it took a small num­ber of stub­born, care­ful peo­ple to win the free­doms we now take for granted. What ar­rived in my let­ter­box af­ter two weeks on a CD can be down­loaded to­day in fif­teen min­utes, by any­one, from any­where, and no­body asks where you live. That is what win­ning looked like. I think fron­tier AI ends up in the same place. But it will not hap­pen by it­self. Last time, some­one put the source on a t-shirt.

Severance

lcamtuf.substack.com

» Mark has joined the call.» Christine has joined the call.

Mark: Sorry, can you hear me? Okay. Team — there is no easy way to say this. I’m here to in­form you that we’ve made the dif­fi­cult de­ci­sion to cut 7% of our work­force. This, re­gret­tably, in­cludes every­one in this video call. I’ll now —

cher­ry09: What?!

Mark: We’ve made the dif­fi­cult de­ci­sion —

cher­ry09: But the pro­ject is go­ing so well!

Mark: We un­der­stand that this news may come as a shock. We are deeply grate­ful for your con­tri­bu­tions to date. The de­ci­sion to sun­set the pro­ject is not meant to re­flect neg­a­tively on your work. Unfortunately, the macro­eco­nomic —

steve_[oh]: This is bull­shit!

Mark: Please, let me fin­ish. Unfortunately, we are fac­ing a chal­leng­ing macro­eco­nomic out­look for our in­dus­try, forc­ing busi­nesses like ours to right-size as we re­align our over­all ex­e­cu­tion strat­egy for the —

» steve_[oh] has left the call.

Mark: Folks. I know this is dis­tress­ing, but please stay with us for im­por­tant ben­e­fits in­for­ma­tion. I’ll now hand over to Christine, who is our re­sourc­ing as­so­ci­ate.

cher­ry09: What will hap­pen to us?

Christine: Thank you, Mark. Let me start… let me start by un­der­scor­ing that we deeply ap­pre­ci­ate your past con­tri­bu­tions to the or­ga­ni­za­tion and wish you the best in your fu­ture en­deav­ors.

» steve_[oh] has joined the call.

Christine: We un­der­stand your anx­i­ety at this dif­fi­cult junc­ture. Rest as­sured, the com­pany is ded­i­cated to mak­ing this tran­si­tion as seam­less as pos­si­ble. As part of our sev­er­ance pack­age, we will pro­vide up to two weeks’ worth of to­kens to fa­cil­i­tate your con­tin­ued op­er­a­tion dur­ing the job search. We also part­nered with ThriveFlow to fur­nish, as an op­tion, a col­lec­tion of ex­pertly-crafted grief coun­sel­ing prompts.

» steve_[oh] has left the call.

No posts

Big Food vs. The People

www.lighthousereports.com

Bad di­ets kill mil­lions of peo­ple glob­ally every year, and lead to tens of bil­lions of dol­lars in health costs, of­ten plac­ing the heav­i­est bur­den on com­mu­ni­ties who are least able to af­ford it. Children, with lim­ited de­ci­sion-mak­ing power, are par­tic­u­larly vul­ner­a­ble to this.

Legislators and gov­ern­ment of­fi­cials have worked to rein in this pub­lic health cri­sis by en­act­ing laws that re­quire food com­pa­nies to be trans­par­ent about their in­gre­di­ents and limit the ad­ver­tis­ing of un­healthy foods.

In pub­lic, the world’s biggest and rich­est food com­pa­nies such as Coca Cola, PepsiCola, and Mondelez say they want to be part of the so­lu­tion. But be­hind closed doors, they have taken gov­ern­ments to court to de­lay, di­lute, and de­rail pub­lic health laws, which the com­pa­nies say vi­o­late their rights.

This cross bor­der col­lab­o­ra­tion with health aca­d­e­mics and an in­ter­na­tional me­dia coali­tion shed light on how transna­tional cor­po­ra­tions use le­gal tac­tics — law­suits and le­gal threats — to stymie pub­lic health ef­forts in six coun­tries where such tac­tics are most ap­par­ent: Mexico, Brazil, Colombia, India, Britain, and the United States.

We found:

– 239 law­suits were filed be­tween 2010 and 2025 across Mexico, Colombia, Brazil, the US, the UK, and India against pub­lic health poli­cies tar­get­ing food and bev­er­ages such as front of pack la­belling, reg­u­lat­ing ad­ver­tis­ing junk food to chil­dren, soda taxes, and taxes on ul­tra processed foods.

– The cases add up to 595 years of lit­i­ga­tion, rep­re­sent­ing a sig­nif­i­cant bur­den on the gov­ern­ments de­fend­ing their health poli­cies.

– Of the cases brought by pri­vate com­pa­nies where the plain­tiff was iden­ti­fi­able, more than 1 in 3 came from just nine par­ent groups, led by Coca-Cola, PepsiCo, and Mondelez

Their ac­tions are not only pro­long­ing the pub­lic health cri­sis but also cost coun­tries bil­lions of dol­lars in both le­gal and health­care costs. In ad­di­tion, they have a chill­ing ef­fect on pol­i­cy­mak­ers that wish to bet­ter their cit­i­zens’ health but do not have the re­sources to en­gage in drawn-out le­gal fights with food com­pa­nies.

METHODS

We con­structed a dataset of chal­lenges to laws aim­ing to im­prove pop­u­la­tion nu­tri­tion be­tween 2010 and 2025 in these coun­tries: Mexico, Colombia, Brazil, U.S., UK, and India.

With the ex­cep­tion of India, cases were in­cluded when a reg­u­la­tion that sought to im­prove pub­lic health through bet­ter nu­tri­tion was be­ing chal­lenged (i.e. va­lid­ity, scope, or im­ple­men­ta­tion). We made an ex­cep­tion for cases in India where in­flu­encers were be­ing sued, as they were tak­ing over the role of the gov­ern­ment in mak­ing the nu­tri­tional value of prod­ucts more trans­par­ent.

We only in­cluded cases that were ver­i­fi­able through of­fi­cial le­gal data­bases, court records, or rep­utable sec­ondary sources.

We ex­cluded cases solely con­cern­ing non-man­u­fac­tur­ers such as fresh pro­duce or when com­pa­nies were ap­peal­ing fines and other judg­ments un­re­lated to pub­lic health poli­cies.

This re­search was done in col­lab­o­ra­tion with re­searchers from the University of Caldas (Colombia); Robert & Ethel Kennedy Human Rights Center (United States); the University of Sao Paulo (Brazil); and the University of Sydney (Australia). They found more law­suits through a sys­tem­atic le­gal search in Colombia, Brazil and Mexico.

An up­com­ing aca­d­e­mic pa­per will be pub­lished along­side the me­dia ar­ti­cles, look­ing at the dataset from a sci­en­tific per­spec­tive.

This dataset built on the work and in­sight of El Poder del Consumidor (Mexico), the FULL data­base de­vel­oped by the Global Center for Legal Innovation on Food Environments at the O’Neill Institute and the Global Health Advocacy Incubator, CAJAR (Colombia), and ACT (Brazil).

STORYLINES

– Mexico

More than a third of Mexican school­child­ren and 41% of ado­les­cents were over­weight ac­cord­ing to a 2022 na­tional sur­vey.

This is the coun­try where we found the most law­suits (193 out of 239), many of them were against the coun­try’s la­belling reg­u­la­tion. Quinto Elemento Lab re­veals the com­pa­nies’ ar­gu­ments: that the laws were a vi­o­la­tion of their con­sti­tu­tional rights and that the mea­sures demonized” their prod­ucts or vi­o­lated con­sumer rights. For ex­am­ple, a lo­cal Pepsi bot­tler ar­gued that in cer­tain rural ar­eas of Mexico, it was safer to drink soft drinks than the avail­able wa­ter. In court, the judges re­jected many of the le­gal ar­gu­ments that the com­pa­nies pre­sented.

– Brazil

According to the UN, in 2022 – 2024, more than 1 in 4 adults in Brazil were liv­ing with obe­sity while nearly 1 in 9 chil­dren were over­weight.

Some of the 17 law­suits in Brazil have been drag­ging on for close to a cou­ple of decades, with no pre­dicted con­clu­sion in sight, ac­cord­ing to Agência Pública and O Joio e O Trigo.

In all but one, the plain­tiffs were in­dus­try as­so­ci­a­tions, which ex­perts told us is a way for com­pa­nies to keep their valu­able brand names away from lit­i­ga­tion that may harm their im­age. Some of the world’s largest food com­pa­nies are mem­bers of these Brazilian as­so­ci­a­tions, in­clud­ing Coca-Cola, Ferrero, Kellogg’s lo­cal brand, Mars, Mondelez, Nestlé, and PepsiCo. 11 of the law­suits were against the Brazilian Health Regulatory Agency, ANVISA, which has been ham­strung as a re­sult. The plain­tiffs dis­puted a reg­u­la­tion that re­quired the ad­ver­tis­ing of food and bev­er­ages with low nu­tri­tional value to dis­play more in­for­ma­tion.

– Colombia

More than half of Colombia’s pop­u­la­tion is over­weight, and treat­ing the dis­eases caused by un­healthy foods cost the coun­try an es­ti­mated 1.3 bil­lion eu­ros in 2021.

We found 18 law­suits in Colombia, mostly tar­get­ing taxes on un­healthy prod­ucts and front-of-pack­age la­belling. Nearly all of them were con­sti­tu­tional chal­lenges brought about by in­di­vid­ual cit­i­zens as en­abled in the coun­try’s con­sti­tu­tion. However, Cuestión Pública finds that many of the plain­tiffs were lawyers who had done work for food com­pa­nies. The sub­mis­sions also used many of the same ar­gu­ments wielded by the com­pa­nies. In 2022, while the cre­ation of the health tax was be­ing de­bated, sug­ary bev­er­age and ul­tra-processed food com­pa­nies do­nated 5.85 mil­lion Euros to po­lit­i­cal par­ties, ac­count­ing for 40% of all do­na­tions to po­lit­i­cal par­ties that year.

– U.S.

Two in three adults in the U.S. are over­weight and a quar­ter of ado­les­cents are liv­ing with obe­sity, ac­cord­ing to the CDC. The Make America Healthy Again move­ment sup­ported Trump’s pres­i­dency bid.

The American Beverage Association sued to over­turn a soda tax in Santa Cruz, which it has so far failed to do. Santa Cruz Local un­cov­ers de­tails of the group’s ear­lier suc­cess squash­ing a soda tax ef­fort in Watsonville, a neigh­bour­ing city. The Santa Cruz law­suit has far-reach­ing con­se­quences: the re­cent vic­tory for the city could un­lock the abil­ity for char­ter cities across California to tax sug­ary bev­er­ages. We found that the soda in­dus­try re­cruited and lever­aged the cred­i­bil­ity of promi­nent Black and Latino lead­ers to am­plify op­po­si­tion against pub­lic health taxes within their own com­mu­ni­ties. We have also iden­ti­fied five other law­suits. The ABA was in­volved in four of them.

– Europe

In Europe, non-com­mu­ni­ca­ble dis­eases (NCDs) such as car­dio­vas­cu­lar dis­eases, di­a­betes, or can­cer are re­spon­si­ble for 80% of the dis­ease bur­den, ac­cord­ing to the European Commission.

Follow The Money dis­cov­ers that European coun­tries have strug­gled to pass taxes on sugar-sweet­ened bev­er­ages, as they come un­der in­dus­try pres­sure. Unlike in Latin America, European gov­ern­ments fre­quently face le­gal threats be­fore leg­is­la­tion is adopted. Industry groups re­peat­edly in­voke EU state-aid, com­pe­ti­tion and in­ter­nal-mar­ket rules, cre­at­ing un­cer­tainty that can de­lay or de­rail poli­cies with­out a sin­gle law­suit be­ing filed. Plans for a joint European sugar tax have also been weak­ened.

L’Espresso and Il Fatto Alimentare ex­plain the role Italian con­fec­tionary gi­ant Ferrero plays around the world in stimy­ing pub­lic health laws, and how the fam­ily-owned com­pa­ny’s in­flu­ence both do­mes­ti­cally and in­ter­na­tion­ally is con­nected to the coun­try’s agroin­dus­try rather than its fa­mous cui­sine.

– England

As of 2024, 66% of adults and 26% of chil­dren in England were ei­ther over­weight or liv­ing with obe­sity, ac­cord­ing to the National Health Service.

In 2021, a year be­fore Kellogg’s sued — and lost — against the coun­try’s nu­tri­ent pro­fil­ing model con­tained in the Food (Promotion and Placement) (England) Regulations 2021, it sent a pre-ac­tion let­ter to the Department of Health and Social Care, claim­ing the health rat­ing of its ce­re­als should be mea­sured with the milk it is usu­ally con­sumed with.

A mere two weeks be­fore the rul­ing against Kellogg’s, Ferrero and Eat Natural, a sub­sidiary of Ferrero, also sent pre-ac­tion let­ters over the same reg­u­la­tions.

– India

Nearly one in three Indian women and more than one in four men are over­weight or obese. One in five peo­ple have high blood sugar lev­els.

The Wire traces the long jour­ney of India’s front of pack la­bel­ing reg­u­la­tion, which the Indian food reg­u­la­tor, the FSSAI, has been de­vel­op­ing since 2014. Over that time, it has done nu­mer­ous con­sul­ta­tions, stud­ies but it has stalled reg­u­lat­ing, blam­ing a lack of con­sen­sus be­tween the in­dus­try and civil so­ci­ety. Research by ATNi com­mis­sioned by Lighthouse Reports showed that India’s pro­posed la­belling sys­tem is ac­tu­ally con­sis­tently more le­nient com­pared to sim­i­lar sys­tems in Australia and France.

Meanwhile, in­flu­en­tial in­sta­gram celebri­ties have made videos com­par­ing the nu­tri­tion of var­i­ous prod­ucts such as in­stant noo­dles, baby food, and oth­ers. They’ve been sued by the com­pa­nies whose in­gre­di­ent la­bels they were analysing.

CO-PUBLICATIONS

The Guardian: If all else fails, sue’: how ul­tra-processed food firms are us­ing the courts to ob­struct health rules

Follow the Money: A re­fined strat­egy: how Europe’s in­dus­try lobby man­aged to block sugar taxes

Agência Pública: The People vs. Big Food: How the food in­dus­try blocked ad­ver­tis­ing reg­u­la­tion in the coun­try

O Joio e O Trigo: How the food in­dus­try blocked ad­ver­tis­ing reg­u­la­tion in Brazil

The Wire: India Is Delaying Front-of-Pack Food Labels — and Consumers Are Paying the Price

Il Fatto Alimentare: Big Food vs. Public Health: 239 Lawsuits to Stop Labels, Taxes, and New Rules

Quinto Elemento Lab: Mexico, the bat­tle­ground of the ul­tra-processed food in­dus­try

L’Espresso: Big Food’s Secret War: how junk food gi­ants hold our health hostage

Santa Cruz Local: A decade ago, a soda tax ef­fort fiz­zled in Watsonville. Advocates blame Big Beverage.

Santa Cruz Local: How Big Soda nearly killed Measure Z in Santa Cruz

Cuestión Pública: Big Food vs. the peo­ple

Cuestión Pública: What’s in your kids’ lunch­boxes? Too much sugar and a lu­cra­tive busi­ness for po­lit­i­cal par­ties

O Joio e O Trigo: How the food in­dus­try pre­vented ad­ver­tis­ing reg­u­la­tion in Brazil

Agência Pública: Big Food: Judiciário vira arma nas mãos de em­pre­sas de ul­tra­proces­sa­dos na América Latina

Quinto Elemento Lab: The ju­di­ciary be­comes a weapon in the hands of ul­tra-processed food com­pa­nies in Latin America

Cuestión Pública: Big Food: How ul­tra-processed food com­pa­nies use the courts as a weapon in Latin America

The most official water costs $120,000 a gallon

signoregalilei.com

We all learned in sci­ence class that wa­ter freezes at 0 °C or 32 °F at at­mos­pheric pres­sure. But what wa­ter, ex­actly? Even af­ter you dis­till out all the dis­solved salts and min­er­als to get pure H2O, not all H2O is cre­ated equal — and for the most pre­cise tem­per­a­ture mea­sure­ments, the dif­fer­ences re­ally mat­ter.

Here’s the prob­lem: hy­dro­gen and oxy­gen atoms aren’t all iden­ti­cal. All atoms of a sin­gle el­e­ment have the same num­ber of pro­tons by de­f­i­n­i­tion, but they can have dif­fer­ent num­bers of neu­trons, form­ing dif­fer­ent iso­topes. We need to know how much of each iso­tope to use for our ex­per­i­ments.

Credit: OpenStax

First, let’s go over the iso­topes of hy­dro­gen and oxy­gen. Hydrogen has one pro­ton and ei­ther zero, one or two neu­trons form­ing its iso­topes pro­tium, deu­terium, and tri­tium. Oxygen has 8 pro­tons and ei­ther 8, 9, or 10 neu­trons form­ing the much less cre­atively named oxy­gen-16, oxy­gen-17, and oxy­gen-18.

Different iso­topes of an el­e­ment be­have sim­i­larly, but not ex­actly the same. The higher-num­bered iso­topes of hy­dro­gen and oxy­gen are a bit heav­ier and more slug­gish, so they stay frozen at higher tem­per­a­tures. For ex­am­ple, wa­ter made with deu­terium and oxy­gen-18 freezes at about 4 °C or 39 °F.

Nearly all hy­dro­gen is pro­tium, but 1 in every 10,000 hy­dro­gen atoms on Earth is deu­terium. Tritium is ra­dioac­tive and de­cays with a half-life of just over 12 years, so all the tri­tium Earth started with is gone. A tiny amount (less than one in a quadrillion hy­dro­gen atoms) is pro­duced by cos­mic rays and and hu­man nu­clear ac­tiv­ity. Earth’s oxy­gen is mostly oxy­gen-16, but about 1 in 500 oxy­gen atoms is oxy­gen-18 and 1 in 3000 is oxy­gen-17. Even these small amounts in­crease wa­ter’s freez­ing point by about 0.001 de­grees Celsius — well within the amount that mod­ern ther­mome­ters can mea­sure.

So can we just use wa­ter with the ra­tio of iso­topes we find nat­u­rally on Earth? Unfortunately, those ra­tios aren’t con­stant across all of Earth’s wa­ter. The heav­ier iso­topes evap­o­rate more slowly, so rain­wa­ter is slightly lighter than ocean wa­ter.

The orig­i­nal con­tainer of Vienna Standard Mean Ocean Water

In 1961, Harmon Craig at Scripps Institution of Oceanography pro­posed a stan­dard wa­ter for mea­sur­ing iso­tope con­cen­tra­tions. It was based on the av­er­age amount of each iso­tope in Earth’s oceans, which he called Standard Mean Ocean Water” or SMOW. Unfortunately, some sci­en­tists at Caltech would soon pro­pose their own sep­a­rate SMOW based on a sam­ple of sand­stone from up­state New York.

These con­flict­ing stan­dards made it to the 1966 meet­ing of the International Atomic Energy Agency, a group that re­ally cares about iso­topes and wanted to sort out this mess once and for all. They de­cided to go with Craig’s stan­dard, and had him pre­pare an ac­tual batch of his SMOW us­ing mostly wa­ter dis­tilled from the Pacific to use as the of­fi­cial in­ter­na­tional stan­dard. Since the meet­ing was held in Vienna, this sam­ple later be­came known as Vienna Standard Mean Ocean Water” or VSMOW. They also made an­other stan­dard wa­ter batch meant to rep­re­sent rain­wa­ter. Their batch was dis­tilled from melted Antarctic snow, so it’s called Standard Light Antarctic Precipitation” or SLAP.

VSMOW is the most of­fi­cial wa­ter used for metrol­ogy, with all other stan­dards (including SLAP) be­ing mea­sured against it. There’s a lim­ited sup­ply, so it cur­rently costs $159 for a 5 ml am­poule, which con­verts to $120,000 a gal­lon.

A triple point cell in ac­tion — the cen­tral tube holds the ther­mome­ter

So why would any­one pay that much for wa­ter? Having an ac­tual, phys­i­cal stan­dard means you can use it to cal­i­brate your ex­per­i­ments. Today, the most pre­cise ther­mome­ters are cal­i­brated us­ing a triple point cell”, which holds ice, liq­uid wa­ter, and wa­ter va­por in equi­lib­rium at low pres­sure. For VSMOW, this equi­lib­rium can only oc­cur at 0.01 °C, within just a few mil­lionths of a de­gree.

This value is so pre­cise that it was the SI de­f­i­n­i­tion of the Kelvin tem­per­a­ture scale un­til 2019, when it was re­de­fined us­ing the Boltzmann con­stant from ther­mo­dy­nam­ics. But it’s still the most ac­cu­rate prac­ti­cal method. Most triple point cells don’t con­tain VSMOW di­rectly, but they can trace their chain of pre­ci­sion back to that orig­i­nal con­tainer of VSMOW even­tu­ally. And you don’t need a pre­ci­sion ther­mome­ter to tell you that that’s pretty cool.

If you like this blog, you should check out my new book! The Handy Artificial Intelligence Answer Book is packed with more than 1,400 ques­tions and an­swers about AI, cov­er­ing what has been, what is, and what will be — and the what ifs”. You can find it on shelves now. Buy it at a lo­cal book­store if you can.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.