10 interesting stories served every morning and every evening.

Statement by Prime Minister Carney on Canada-U.S. trade negotiations

www.pm.gc.ca

Over the past 18 months, Canada’s new gov­ern­ment has fo­cused on build­ing our strength at home, di­ver­si­fy­ing our part­ner­ships abroad, and strik­ing a fair deal with the United States.

Our ob­jec­tives in our trade ne­go­ti­a­tions have been to:

Preserve tar­iff-free ac­cess to the U.S. for the vast ma­jor­ity of Canadian busi­ness;

Provide greater sta­bil­ity to our trade re­la­tion­ship;

Significantly re­duce U.S. tar­iffs on our key strate­gic in­dus­tries, so that Canadian busi­nesses in these sec­tors would have the best ac­cess of any in the world;

Protect our small and medium-sized busi­nesses — the lifeblood of our econ­omy — in­clud­ing by re­mov­ing the im­mi­nent threat of new tar­iffs; and

Maintain our flex­i­bil­ity, in­de­pen­dence, and sov­er­eignty so we can keep build­ing the Canada we want.

We have recog­nised from the be­gin­ning that America has changed, and that we will not re­turn to our old re­la­tion­ship. Our gov­ern­ment un­der­stood, be­fore many, that America is al­ter­ing all its trade re­la­tion­ships. Putting tar­iffs on its clos­est al­lies and charg­ing for ac­cess to its vast mar­ket.

We have worked in that con­text. To strike a fair deal that would pro­vide the best ac­cess to the U.S. mar­ket and greater cer­tainty to Canadian busi­nesses and work­ers. Throughout, our goal has been to se­cure the best deal for Canadians, never a deal at any price or on any dead­line.

In re­cent weeks, we made im­por­tant progress to­ward im­prov­ing Canada’s po­si­tion as hav­ing the best deal in the world with the U.S.

However, that progress has not been enough to meet our ob­jec­tives for Canadians. As a re­sult, this evening, I have de­cided to sus­pend trade ne­go­ti­a­tions with the U.S. and have di­rected Canada’s ne­go­tia­tors to re­turn to Ottawa. They have worked hard, in good faith, to de­fend the in­ter­ests of Canadians through­out these ne­go­ti­a­tions up un­til the very last minute. However, last-minute changes in the U.S. pro­posed terms were un­fair, un­eco­nomic, and called into ques­tion the re­li­a­bil­ity of any deal.

At mid­night tonight, the U.S. in­tends to im­pose a 50% tar­iff on roughly $28 billion of Canadian goods. Canada will match those tar­iffs dol­lar for dol­lar to pro­tect our work­ers and busi­nesses.

In the com­ing days, the gov­ern­ment will in­tro­duce ad­di­tional mea­sures to sup­port Canadian work­ers and busi­nesses, build­ing on the nearly $25 billion in sup­port pro­vided over the past 18 months.

These ac­tions com­ple­ment Canada’s core eco­nomic strat­egy. From day one, we have been fo­cused on build­ing our strength at home and di­ver­si­fy­ing our part­ner­ships abroad.

That strat­egy is work­ing. We are ad­vanc­ing nearly $500 billion in ma­jor in­fra­struc­ture pro­jects. In par­al­lel, we are un­lock­ing new ex­port mar­kets for Canadian busi­nesses. Our ex­ist­ing free trade deals al­ready pro­vide Canada with pref­er­en­tial ac­cess to 1.5 billion con­sumers, and we are on track to dou­ble that mar­ket ac­cess by the end of this year.

Canadian eco­nomic growth is ac­cel­er­at­ing, and we are on course to have the sec­ond-fastest growth in the G7 over the next two years. Our econ­omy is cre­at­ing jobs at four times the rate of the United States. Our ex­ports to non-U.S. mar­kets are on track to dou­ble over the next decade. Foreign di­rect in­vest­ment in Canada is at its high­est level in two decades, run­ning at twice the rate of our near­est G7 com­peti­tor. Canada now ranks as the most at­trac­tive coun­try in the world for in­fra­struc­ture in­vest­ment.

Canada has what the world wants. And we will not al­low any na­tion to de­ter­mine our fu­ture. We will set our own course to keep build­ing Canada strong for all.”

ElevenLabs, TwelveLabs, ThirteenLabs, …

quantumi.sh

You may have heard of the speech syn­the­sis com­pany ElevenLabs. Recently a friend men­tioned they knew some­one who worked at a com­pany called TwelveLabs that does AI for video (it feels like it must have been in­tended to play on the fact that ElevenLabs does au­dio, but I don’t know for sure). Jokingly, I googled thirteenlabs” and was sur­prised to find an AI for 3D scenery pro­ject. I googled fourteenlabs”. Another AI startup…?? How far does this go?

Numbers 0 – 99 an­no­tated with links to com­pa­nies us­ing it + labs” as their name. A link has a back­ground if the com­pany is AI-related.

The com­pany must have some on­line pres­ence (not nec­es­sar­ily a web­site, but usu­ally one).

The word labs” (or some­times lab”) must ei­ther go af­ter or be­fore the num­ber (spelled or nu­meric). There’s plenty of com­pa­nies named af­ter a num­ber - what I’m af­ter is the weird trend of num­ber + labs”.

If there were mul­ti­ple, I picked the one that was most sim­i­lar to ElevenLabs (and thus more likely to have got­ten the name in­spi­ra­tion from there?).

Every com­pany sort of wants to be an AI com­pany now, so it’s hard to def­i­nitely say what com­pa­nies are AI re­lated. I marked it if the do­main had a .ai TLD or if all of their main prod­ucts in­volved AI in a cen­tral ca­pac­ity.

This re­ally just raised more ques­tions than an­swers. Why is this such a pop­u­lar nam­ing scheme? Are peo­ple in­de­pen­dently ar­riv­ing at this nam­ing scheme? Why call your AI startup 68labs”? Why are the sev­en­ties so much more dense than the rest of the higher num­bers? Also, I’m tempted to spec­u­la­tively buy up do­mains like twentyfivelabs” or thirtytwolabs”…

One fun dis­cov­ery: among all the in­cred­i­bly same-y startup web­sites, there is sev­en­ty­onelab.com. Alongside the usual all rights re­served” text, it po­litely in­forms you it is

best viewed in Netscape 4/0+ or IE 5.0+

best viewed in Netscape 4/0+ or IE 5.0+

Amazing. It looks like a de­sign/​web­dev port­fo­lio from the early 2000s and is full of fun lit­tle sites. I par­tic­u­larly like the aes­thetic of the land­ing page and the other ver­sions of it in the little pro­ject” sec­tion. They sort of re­mind me of what I re­cently heard re­ferred to as the vectorheart” aes­thetic (like what you’d see on 2000s IDM al­bum cov­ers!).

Hister | Your Own Search Engine

hister.org

Free soft­ware

AGPLv3

Self hosted

Run it on your own ma­chine or server

Privacy fo­cused

No teleme­try, no ex­ter­nal re­quests

Preserve Knowledge.Find It Again.

Hister stores ex­tracted doc­u­ment con­tent with the search in­dex and dis­plays it as a read­able pre­view along­side search re­sults.

Download Hister Try the live demo

Search the con­tent you chose to in­dex

Look be­yond book­marks and file­names. Hister in­dexes the full con­tent of the pages and files you choose, then keeps that search­able knowl­edge on a server you con­trol.

Search the ac­tual con­tent

Look be­yond ti­tles and URLs to the words in­side every in­dexed doc­u­ment.

Narrow with pre­ci­sion

Use fields, phrases, wild­cards, nega­tion, pri­or­i­ties, and your own aliases.

Read it in con­text

Open a clean stored pre­view be­side the re­sults with­out los­ing your search.

How Hister works

A pri­vate mem­ory with­out the busy­work

Collect

Save newly vis­ited pages with the browser ex­ten­sion, watch lo­cal fold­ers, im­port your his­tory, or crawl a site.

Visited pages

Local files

Index

Hister ex­tracts the parts that mat­ter and in­dexes their full text on the server you choose.

Full text

Stored con­text

Find

Search from the web, ter­mi­nal, com­mand line, or let an AI as­sis­tant re­trieve it through MCP.

Web and ter­mi­nal

MCP as­sis­tants

The sim­ple loop

Browser ex­ten­sions can in­dex pages as they are vis­ited. File watch­ers, his­tory im­ports, and crawlers add other sources to the same in­dex.

Why Hister

Preserve Knowledge. Keep Control.

The in­dex, stored page con­tent, and rules re­main on the Hister server you con­fig­ure. The server has no teleme­try and does not re­quire a cloud ser­vice.

No teleme­try

The server does not phone home or re­port what you search.

No manda­tory cloud

A com­plete per­sonal setup can run on one lo­cal ma­chine.

Your cho­sen server

Clients send in­dexed con­tent only to the Hister server you con­fig­ure.

Auditable soft­ware

The source is pub­lic and li­censed as free soft­ware un­der AGPLv3.

Optional se­man­tic search sends text to the em­bed­dings end­point you con­fig­ure. Browser ex­ten­sions may re­trieve page fav­i­cons. You choose whether and where these con­nec­tions run.

Built for real re­call

One in­dex. Many ways back in.

Hister in­dexes vis­ited pages, watched files, im­ported browser his­tory, and crawled web­sites. The in­dex is avail­able through web, ter­mi­nal, CLI, HTTP API, and MCP in­ter­faces.

Browser ex­ten­sions can in­dex vis­ited pages au­to­mat­i­cally. File watch­ing, his­tory im­ports, and the crawler add other sources.

Browser ex­ten­sions

Local file watch­ing

History im­port

Website crawler

Full text search sup­ports field fil­ters, quoted phrases, wild­cards, nega­tion, date ranges, and query aliases.

Field fil­ters

Quoted phrases

Wildcards and nega­tion

Query aliases

Content ex­trac­tors han­dle struc­tured data from sup­ported for­mats and web­sites. Semantic search is op­tional.

Content ex­trac­tors

Semantic search

Language aware in­dexes

Readable pre­views

Skip and pri­or­ity rules con­trol in­dex­ing and rank­ing. Versioning can re­tain ear­lier doc­u­ment con­tent.

Skip rules

Priority rules

Version track­ing

Sensitive con­tent checks

The same in­dex is avail­able through the web in­ter­face, ter­mi­nal client, CLI, HTTP API, and MCP server.

Web in­ter­face

Terminal in­ter­face

HTTP API and CLI

MCP server

A sin­gle bi­nary can run lo­cally. Shared servers sup­port user scoped ac­cess with SQLite or PostgreSQL.

No con­fig quick­start

SQLite or PostgreSQL

Multiple users

Docker and Nix

Additional ca­pa­bil­i­ties

Hister also sup­ports mul­ti­ple crawler back­ends, lan­guage spe­cific in­dexes, con­tent ver­sion­ing, own­er­ship rules, and con­fig­urable ex­trac­tors.

Why your local LLM feels dumber than it is

forum.level1techs.com

Quick Introduction

We have all been on fo­rums, chats, red­dit, dis­cord, youtube, or some­where and heard Oh! Model XYZ is AMAZEBALLZ!zomgwtfbbq” then down­loaded it (or more likely, some quan­tized form of it) and said eww… This sucks!”

This post is go­ing to be a rather tech­ni­cal se­ries of ex­per­i­ments to demon­strate the im­pact of im­ple­men­ta­tion-spe­cific haz­ards with in­fer­ence. I will be us­ing the term reference im­ple­men­ta­tion” to de­scribe the lab that pub­lished and of­fers first-party host­ing of their mod­els and posts orig­i­nal bench­mark claims. Their hard­ware will be dif­fer­ent than yours. Their soft­ware will be very dif­fer­ent than yours. And the com­par­isons in this post are not go­ing to be run­ning some 2.58-bit-gguf-in-ollama with a cou­ple test prompts.

I am in­ten­tion­ally gloss­ing over en­tire emerg­ing fields of study, moun­tains of re­search pa­pers and lit re­view to make this more ap­proach­able for you the reader. Don’t nit pick my over­sim­pli­fi­ca­tions or I will make you read the re­ally long un­pleas­ant ver­sion with math.

Your lo­cal im­ple­men­ta­tion sucks. But that’s ok, be­cause every­one else’s does too.

Every sin­gle in­stance of hard­ware and soft­ware run­ning an LLM to­day is a lit­tle bit dif­fer­ent. or a lot dif­fer­ent when it comes to some cases. The av­er­age home lab user might be mix­ing mul­ti­ple dif­fer­ent gen­er­a­tions of GPU. The chips on those have dif­fer­ent in­struc­tion sets. Those in­struc­tion sets will im­ple­ment and ex­e­cute math to cal­cu­late your next to­ken dif­fer­ently from any other per­son, even when run­ning the same ex­act weights.

So that begs the first ques­tion: How much does your par­tic­u­lar setup suck? Turns out there are a num­ber of dif­fer­ent ways to go about mea­sur­ing that.

The prac­ti­cal ap­proach is straight for­ward. Run stan­dard bench­marks. A va­ri­ety of them. ter­mi­nal bench, hle, SWEthis, HELLAthat, MMLU-whatever… take your pick. Just make sure its rep­re­sen­ta­tive of your ac­tual work­load/​use case. Do not crank tem­per­a­ture to zero and paste in 3 test prompts then call it good/​bad. Zero-shot tests are not a good ana­log of most agen­tic tasks. You need long-con­text tool-call­ing and do­main spe­cific knowl­edge eval­u­a­tions to fig­ure out where your setup is weak when run­ning the same weights as some­body else repli­cat­ing those same bench­marks.

But the purely math­e­mat­i­cal an­swer is where my fo­cus is go­ing to be­gin be­cause as @wendell said:

Math is Math!

Logits” are the mod­els scores for each pos­si­ble next to­ken. They are nor­mal­ized into prob­a­bil­i­ties, passed through the con­fig­ured sam­pler, and con­verted back into text by the deto­k­enizer to gen­er­ate THE→NE→XT→TOK→EN dur­ing de­code.

A side note about sam­pler set­tings: the model card on HF usu­ally spec­i­fies ex­actly what sam­pler set­tings (and chat tem­plate) you should be us­ing. temp 1.0, top-p 0.95, etc. it varies by model so make sure you are us­ing the right ones. btw, set­ting temp too low is why your qwen is sit­ting there loop­ing un­able to es­cape its THINK out­put. You’re wel­come, glad I could fix that for you.

When the next to­ken prob­a­bil­ity changes enough, THE→NE→XT be­comes THE→NE→W→DAY… And while those small changes might be fine, odds are that’s the be­gin­ning of the nig­gling sen­sa­tion in the back of your mind that some­thing feels off.

Some of you may have heard the term KLD be­fore, or KL Divergence. Don’t worry, I won’t make you do any math or flood your brain with ta­bles of very small dec­i­mal num­bers. But just in case you wanted the sim­ple ver­sion: con­vert the out­put log­its into a prob­a­bil­ity dis­tri­b­u­tion, and mea­sure how far that dis­tri­b­u­tion has moved from a cho­sen base­line. Lower KLD means closer to that base­line, not au­to­mat­i­cally smarter’. KLD is also di­rec­tional, so the or­der of the two dis­tri­b­u­tions mat­ters.

A word of cau­tion: Don’t get suck­ered in by im­pos­si­bly low KLD claims on a quant HF model card. It is im­pos­si­ble to in­ter­pret a num­ber un­less the au­thor dis­closes the ref­er­ence check­points and full run­time en­vi­ron­ment, eval­u­a­tion text, cal­i­bra­tion data, con­text lengths, sam­pled po­si­tions, KL di­rec­tion, any vo­cab­u­lary trun­ca­tion, and how the mea­sure­ments were ag­gre­gated. The method­ol­ogy mat­ters as much as the num­ber and plenty of peo­ple get it wrong.

What the hell is vllm do­ing?

Now, we need to take a brief field trip down what the gi­ant stack of soft­ware is do­ing on your in­fer­ence en­gine to un­der­stand where some of those sources of di­ver­gence come from.

At every step of this over­sim­pli­fied di­a­gram are com­po­nents that can be con­fig­ured or changed based on your spe­cific hard­ware/​soft­ware foot­print, model, quant, ten­sor shape, etc.

The nightly VLLM con­tainer im­age I snagged had 734 (252 uv/​pip Python) pack­ages in it. That’s 734 code­bases each with their own bugs and un­doc­u­mented idio­syn­crasies. The path your spe­cific im­ple­men­ta­tion takes through that moun­tain of code will be dis­tinct.

Test 1: Precision Benchmarking Attention Backends

Lets start with one piece of that in­fer­ence flow­chart. During pre­fill (prompt pro­cess­ing) there are a sev­eral at­ten­tion back­ends your in­fer­ence en­gine will se­lect from. This im­pacts both speed and pre­ci­sion of pre­fill, while re­quir­ing dif­fer­ent cuda ker­nels for every GPU fam­ily / SM com­pute ca­pa­bil­ity 1.3. The CUDA plat­form — CUDA Programming Guide . Lets test them and com­pare.

(I’m re­ally very sorry, I had to…)

I started with the of­fi­cial BF16 check­point of Qwen3.6 – 27B on an RTX PRO 6000 Blackwell GPU at ten­sor par­al­lelism 1. The KV cache was BF16, with no weight/​ac­ti­va­tion or KV-cache quan­ti­za­tion. The soft­ware was a pinned nightly vllm build. I used ea­ger ex­e­cu­tion, dis­abled CUDA graphs, pre­fix caching, and MTP, and used 2k-token chun­ked pre­fill.

Qwen3.6 – 27B is dense, not an MoE, but it is still a hy­brid model. 64 lay­ers re­peat in a pat­tern of three Gated DeltaNet/linear-attention lay­ers fol­lowed by one full-at­ten­tion layer. Only those 16 full-at­ten­tion lay­ers use the se­lec­table at­ten­tion back­end in this ex­per­i­ment; the Gated DeltaNet path re­mained fixed.

The work­load re­played here is Prompt 2”, a roughly 100k to­ken con­text cap­tured from a real Turnstone lab work­stream con­tain­ing mul­ti­ple tool calls and real work prod­ucts. It was se­lected to re­sem­ble what a lo­cal agent ac­tu­ally does rather than a syn­thetic nee­dle-in-a-haystack test. And maybe more im­por­tantly, it does­n’t ap­pear in any bench­mark or train­ing dataset in the wild to­day. Nobody could have bench­maxed for this, or cal­i­brated their quant to ac­com­mo­date it.

There are three avail­able full at­ten­tion back­ends to se­lect from in vllm for this work­load: FlashAttention 2, Flash Inference, and Triton Attention. This was the only change made be­tween ex­e­cu­tions, the rest of the hard­ware and soft­ware stack re­mained sta­ble.

I also per­formed a same-back­end cross-GPU re­peata­bil­ity con­trol. For this graph, I cap­tured the full-vo­cab­u­lary log­its in BF16 every 32 prompt to­kens. Distribution com­par­isons such as KLD were cal­cu­lated af­ter­ward in FP64 from those stored log­its.

Top-1 agree­ment is whether the to­ken with the high­est logit, the greedy argmax, was the same. All three back­ends were eval­u­ated against the same forced to­ken his­tory. A top-1 flip” there­fore means a back­end would have cho­sen a dif­fer­ent greedy next to­ken at that po­si­tion. We did not let that choice al­ter the re­main­ing his­tory. This keeps the math­e­mat­i­cal com­par­i­son con­trolled, but it does not show how far an un­con­strained gen­er­a­tion would branch or whether a tool call would even­tu­ally fail… that comes in test 2 ;D

The fol­low­ing graph shows % of sam­pled log­its re­sult­ing in to­ken flips:

For the first sev­eral thou­sand to­kens, every run of the model agreed about what the next to­ken was go­ing to be re­gard­less of back­end. Then in later por­tions of the prompt, back­ends be­gan dis­agree­ing. Triton was se­lected as the base­line to sim­plify up­com­ing quan­ti­za­tion chi­canery.

Each 8k-token win­dow con­tains 250 sam­pled po­si­tions, one probe every 32 to­kens. The per­cent­age is the frac­tion of those probes where the other back­ends high­est-scor­ing to­ken dif­fered from Triton’s.

Random noise was ac­counted for by run­ning the same test with the same at­ten­tion back­end mul­ti­ple times. The log­its across runs at every hid­den state were bit for bit iden­ti­cal. Meaning this par­tic­u­lar di­ver­gence comes ex­clu­sively from the ma­trix mul­ti­pli­ca­tion and ad­di­tion op­er­a­tions hap­pen­ing dur­ing pre­fill in­side trt/​fa2/​fi.

Disagreements ap­peared in clus­ters and var­ied with prompt con­tent rather than in­creas­ing smoothly with con­text length. This is not ev­i­dence of one uni­ver­sal length at which the model falls apart” but… we will get there soon…

Now that we have a base­line com­par­i­son of in­ter­est­ing prompt fuel, lets dive into…

Test 2: KV Cache quan­ti­za­tion, or why your LLMs IQ drops like a rock af­ter 40k to­kens

Repeating the same method­ol­ogy, we took the BF16 weights and BF16 kv cache base­line above run­ning Triton, and ran the next ex­per­i­ment. What hap­pens when you leave the weights and ac­ti­va­tions alone, and JUST quan­tize the kv-cache?

Ah, di­ver­gence. And this leads us to our first dump­ster-fire of the evening: a com­pletely re­pro­ducible tool call­ing er­ror.

Enough top-to­kens got flipped dur­ing tool calls, we let them play out and while BF16 was fine, int8 kv-cache even­tu­ally man­aged to re­cover, int4 did not!

Test 3: Weight Weight, Don’t Tell Me!

This time we are leav­ing all the kv-caches full size at bf16. We are adding some new play­ers to the game how­ever by com­par­ing:

BF16 ref­er­ence: Qwen/Qwen3.6 – 27B ( Qwen/Qwen3.6 – 27B · Hugging Face )

Official FP8: Qwen/Qwen3.6 – 27B-FP8 ( Qwen/Qwen3.6 – 27B-FP8 · Hugging Face )

INT8 W8A16: TheHouseOfTheDude/Qwen3.6 – 27B-INT8 ( TheHouseOfTheDude/Qwen3.6 – 27B-INT8 · Hugging Face )

NVIDIA NVFP4: nvidia/​Qwen3.6 – 27B-NVFP4 ( nvidia/​Qwen3.6 – 27B-NVFP4 · Hugging Face )

AWQ W4A16: cyankiwi/​Qwen3.6 – 27B-AWQ-BF16-INT4 ( cyankiwi/​Qwen3.6 – 27B-AWQ-BF16-INT4 · Hugging Face )

These 4 quants rep­re­sent a broad pic­ture of weights and ac­ti­va­tions. A no­table piece of in­for­ma­tion for our math­na­sium is the ac­tual CUDA ker­nel / GEMM (general ma­trix mul­ti­pli­ca­tion) / MMA (matrix mul­ti­ply ac­cu­mu­late) in­struc­tions be­ing run to cal­cu­late the log­its for each quant are dif­fer­ent:

Qwen3.6 – 27B (reference)

Weights/activations: BF16 weights, BF16 ac­ti­va­tions

Linear/GEMM: UnquantizedLinearMethod → torch.nn.func­tional.lin­ear. Each CUDA tile se­lected by its as­so­ci­ated shape/​geom­e­try.

KV cache: BF16 (Forced)

Qualification: Reference check­point.

Qwen3.6 – 27B-FP8

Weights/activations: E4M3 FP8 weights in 128×128 blocks; dy­namic FP8 ac­ti­va­tion quan­ti­za­tion in­side con­verted lin­ears; ex­cluded mod­ules such as lm_­head re­main BF16

Linear/GEMM: Fp8LinearMethod → CutlassFp8BlockScaledMMKernel

KV cache: BF16 (Forced)

Qualification: DeepGemm was au­to­mat­i­cally dis­abled be­cause vLLM flags its E8M0 scale for­mat as ac­cu­racy-de­grad­ing for this ar­chi­tec­ture (SM120); CUTLASS was se­lected in­stead. No cal­i­bra­tion dataset was iden­ti­fied in the pub­lished files.

Qwen3.6 – 27B-INT8

Weights/activations: Static, sym­met­ric, chan­nel-wise INT8 lin­ear weights; BF16 ac­ti­va­tions (W8A16). GDN/linear_attn pro­jec­tions and lm_­head ex­cluded from quan­ti­za­tion.

Linear/GEMM: CompressedTensorsWNA16 → MarlinLinearKernel

KV cache: BF16 (Forced)

Qualification: One-shot quan­ti­za­tion with ex­plic­itly no cal­i­bra­tion dataset. Its un­usu­ally good fi­delity is less mys­te­ri­ous once you ac­count for W8A16 plus un­quan­tized GDN pro­jec­tions.

Qwen3.6 – 27B-NVFP4

Weights/activations: Mixed check­point — 208 sta­tic FP8 W8A8 tar­gets cov­er­ing 64 full-at­ten­tion pro­jec­tions and 144 GDN pro­jec­tions; 193 NVFP4 W4A16 tar­gets cov­er­ing 192 MLP pro­jec­tions plus lm_­head, group size 16

Linear/GEMM:

FP8 tar­gets: ModelOptFp8LinearMethod → FlashInferFP8ScaledMMLinearKernel NVFP4 tar­gets: NVFP4 GEMM → MarlinNvFp4LinearKernel

FP8 tar­gets: ModelOptFp8LinearMethod → FlashInferFP8ScaledMMLinearKernel

NVFP4 tar­gets: NVFP4 GEMM → MarlinNvFp4LinearKernel

KV cache: BF16 (Forced)

Qualification: Not na­tive FP4 arith­metic in our up­stream-nightly run. vLLM clas­si­fied the GPU path as lack­ing na­tive FP4 sup­port and ex­plic­itly se­lected weight-only FP4 com­pres­sion through Marlin. The check­point’s em­bed­ded FP8 KV scheme was over­rid­den with BF16 KV for the bake­off.

Qwen3.6 – 27B-AWQ-BF16-INT4

Weights/activations: Static asym­met­ric INT4 weights, group size 32, MSE ob­server; BF16 ac­ti­va­tions (W4A16). GDN/linear_attn pro­jec­tions and lm_­head ex­cluded.

Linear/GEMM: CompressedTensorsWNA16 → MarlinLinearKernel

KV cache: BF16 (Forced)

Qualification: AWQ cal­i­bra­tion dataset dis­closed as STEM and Agentic.”

Other no­table in­for­ma­tion for this run:

Full soft­max/​GQA at­ten­tion for all mod­els was AttentionBackendEnum.TRITON_ATTN; JIT mon­i­tor ob­served ker­nel_u­ni­fied_at­ten­tion.

GDN pre­fill: Triton/FLA GDN pre­fill ker­nel, re­quested as tri­ton, head­_k_dim=128.

During ex­e­cu­tion, the re­cur­rent path also JIT-compiled _causal_conv1d_update_kernel, fused_re­cur­ren­t_­gat­ed_delta_rule_­packed_de­code_k­er­nel, and re­duce_seg­ments.

TP1, ea­ger mode, no CUDA graphs, no MTP/speculative de­cod­ing, lan­guage-only ex­e­cu­tion.

The next-to­ken flip re­sults shake out fairly pre­dictably. TheDude (W8A16) mops the floor with every­body, beat­ing first party FP8 (W8A8) and Nvidia(FP4-is-a-Lie) re­lease. In fact, out of the 5 op­tions, Nvidia’s re­lease comes in dead last hit­ting ~50% to­ken flips by the time we reach 88k con­text.

Both the NVFP4 and AWQ W4A16 failed to prop­erly close their tool calls and botched Cisco com­mand line syn­tax (the cor­rect com­mand was show arp’, while they ex­e­cuted show run’), while both FP8 and INT8 were able to com­plete the cor­rect calls.

In fu­ture ex­per­i­ments I will try to ex­plore the im­pact of us­ing dif­fer­ent fused GEMMs for the same weights, this is an­other in­ter­est­ing source of di­ver­gence where some­times you have to trade pre­ci­sion for speed.

Part 1 Wrap Up

I have quite a few more ex­per­i­ments and ob­ser­va­tions to post, but re­quire a great deal of par­al­lel GPU time to cal­cu­late and record every logit sam­pled across huge con­text chains on mul­ti­ple prompts with dozens of dif­fer­ent set­tings.

If you have spe­cific ques­tions, shoot me a DM or poke me on dis­cord I guess.

Munder Difflin — Agent harness to run an office of your clones

munderdiffl.in

Free, open source and per­for­mant multi-agent har­ness, works with your ex­ist­ing sub­scrip­tions (uses hourly lim­its).

Munder Difflin uses CLI agents run­ning on your com­puter to do any­thing you can do

Supports 12 CLI agent providers off the shelf, more com­ing soon

Claude Code

Codex

Grok

Kimi Code

Antigravity

Qwen

Gemini CLI

OpenCode

Crush

Pi

Copilot

Cursor

PRIVATE CLOUD + NETWORK

Get the Teams plan and dou­ble your pro­duc­tiv­ity

Private Cloud: Run agents 24/7 for each team­mate in iso­lated sand­boxes

Private Network: Allow clones of your team to talk to each other au­tonomously (E2E en­crypted)

MAKE CLONES OF YOUR TEAM · THEY WORK 24/7 🔒 E2E

CLICK A CLONE TO INSPECT

4/4 HUMANS · 4 CLONES ON

HOW IT WORKS

Three steps to a sec­ond you.

1

Install your har­ness

One down­load. It wraps the agent CLI you al­ready use and runs on your lap­top. Your code, your keys, your ex­ist­ing sub­scrip­tion — noth­ing leaves your ma­chine.

2

It be­comes you

It cap­tures your work­flow, your tool­ing and what you know. Every clone you run shares that mem­ory, so the next one you spin up starts al­ready know­ing how you work.

3

Your of­fice gets to work

Your clones work around the clock — and when one needs some­thing, it mes­sages an­other. They hand off work, share con­text and un­block each other, all on your own ma­chine.

JIM’S CLONEPAM’S CLONE🔒 E2E

JIM’S CLONE

Blocked — need the in­voice-state de­sign to­kens.

03:12 · en­crypted

PAM’S CLONE

Sent — to­kens + edge-case flows in billing/​to­kens.json.

03:12 · en­crypted

✓ un­blocked overnight · PR #147 open

WHAT EACH TEAM MEMBER GETS

Understand your har­ness­es’ ca­pa­bil­i­ties.

Munder Difflin does­n’t give your team one shared bot. It acts as a clone of the in­di­vid­ual and con­trols their com­puter.

KICK IT OFF FROM ANYWHERE

YOU

SLACK INBOX

TRIGGERS

YOUR CLONE · YOUR COMPUTER

GOD or­ches­tra­tor reads · plans · routes

re­search Claude Code

build Codex

re­view Claude Code

git · each agent in its own iso­lated work­tree

MemPalace their mem­ory · their ma­chine · nowhere else

⟳ 24/7

🔒 asks pricing fi­nal?” 🔒 an­swers 3 AM · he’s asleep

CLONECLONE E2E · same org only

TEAMMATE’S CLONE · THEIR COMPUTER

GOD or­ches­tra­tor reads · plans · routes

sell Grok

draft Kimi CLI

MemPalace their mem­ory · their ma­chine

⟳ 24/7

WHILE YOU’RE BUSY

Real work. Not demos.

🔍

Reviews like you would

Your clone re­views team­mates’ PRs with your stan­dards and your nit­picks — while you’re in a meet­ing.

💬

Answers for you

How does the billing ser­vice work?” A team­mate’s clone asks yours and gets your an­swer — at 3am, with­out wak­ing you.

🌙

The of­fice never closes

Clones plan, build, hand off, and un­block each other around the clock. You come back to fin­ished threads, not open ques­tions.

👔

You stay the boss

Your clone es­ca­lates only the few de­ci­sions that gen­uinely need a hu­man. Check in oc­ca­sion­ally, an­swer, and it keeps mov­ing.

WHAT EACH NODE CAN DO

Not just for en­gi­neers.

Everything a com­puter does is reach­able from the com­mand line — and CLI agents can drive all of it. So every team­mate gets a clone that does their job, what­ever that job is.

👩‍💻

Developer

Reviews PRs, fixes bugs, ships small fea­tures, babysits CI, keeps docs hon­est.

$ git, tests, de­ploys

🎨

Designer

Audits screens against the de­sign sys­tem, ex­ports as­sets, drafts specs and copy.

$ screen­shots, to­kens, specs

📋

Product man­ager

Writes specs, triages is­sues, keeps boards and docs in sync, preps standup sum­maries.

$ tick­ets, docs, roadmaps

📈

Sales & GTM

Drafts out­reach, preps call briefs, keeps the CRM hon­est, chases fol­low-ups.

$ crm, email, briefs

🗂️

Everyone else

Reports, spread­sheets, files, sched­ul­ing, fol­low-ups — any­thing script­able. Which is every­thing.

typ.ing

typ.ing

typ.ing is an awe­some typ­ing trainer, de­signed to help you get faster and more ac­cu­rate when typ­ing with an ex­ter­nal key­board.

It looks like you’re us­ing a mo­bile de­vice - typ.ing will work much bet­ter if you plug a phys­i­cal key­board in (or just use a com­puter).

Do you have a phys­i­cal key­board con­nected?

Typey typey type

A Friendly Introduction to Racket

geometridae.bearblog.dev

11 Aug, 2026

Lisp is worth learn­ing for the pro­found en­light­en­ment ex­pe­ri­ence you will have when you fi­nally get it.” — Eric S. Raymond

Lisp is worth learn­ing for the pro­found en­light­en­ment ex­pe­ri­ence you will have when you fi­nally get it.” — Eric S. Raymond

Welcome. Today you’ll learn a lan­guage from one of pro­gram­ming’s old­est and most un­usual fam­i­lies. A lan­guage where code is data, where paren­the­ses are pure struc­ture, and where pro­grams can write pro­grams. By the end of this tu­to­r­ial, you’ll have writ­ten your own syn­tax.

A bit of his­tory

Lisp was born in 1958, in­vented by John McCarthy at MIT. For con­text: it’s the sec­ond-old­est high-level lan­guage still in use (only Fortran, from 1957, beats it by a year). Python ar­rived in 1991. JavaScript in 1995. Lisp pre­dates them by more than 30 years and sev­eral ideas we now con­sider modern” were born there:

Garbage col­lec­tion — in­vented for Lisp.

First-class func­tions — pass­ing func­tions as ar­gu­ments, now stan­dard every­where.

The REPL — the in­ter­ac­tive read-eval-print loop that Python, Node, and Julia all have to­day started in Lisp.

Conditionals as ex­pres­sions — the if that re­turns a value.

Homoiconicity — code is a data struc­ture of the lan­guage it­self. This is the big one. We’ll come back to it at the end.

For decades, Lisp was the lan­guage of ar­ti­fi­cial in­tel­li­gence. In the 70s and 80s there were phys­i­cal com­put­ers de­signed to run Lisp di­rectly: the Lisp Machines built by Symbolics and LMI. Then came the AI win­ter,” fund­ing dried up, and Lisp went from star to cult lan­guage.

But in­ter­est­ing ideas don’t die they mu­tate, that’s part of the beauty of the lisps.

From Lisp to Scheme to Racket

In 1975, Gerald Sussman and Guy Steele cre­ated Scheme: a min­i­mal­ist, el­e­gant, al­most math­e­mat­i­cal Lisp. Scheme be­came acad­e­mi­a’s fa­vorite lan­guage for teach­ing pro­gram­ming (the leg­endary book SICP Structure and Interpretation of Computer Programs is writ­ten in Scheme).

In 1995, Matthias Felleisen’s group cre­ated PLT Scheme, a Scheme de­signed for ed­u­ca­tion and pro­gram­ming lan­guage re­search. In 2010 it was re­named Racket, and to­day it’s much more than a Scheme: it’s a lan­guage for build­ing lan­guages. Its un­of­fi­cial motto is lan­guage-ori­ented pro­gram­ming: if your prob­lem needs its own lan­guage, Racket lets you build one in an af­ter­noon.

Who uses Lisp to­day?

More peo­ple than you might think:

Clojure runs in pro­duc­tion at banks, air­lines, and star­tups (Nubank, the largest dig­i­tal bank in Latin America, runs on Clojure).

Common Lisp (with the SBCL com­piler) is still alive in ex­pert sys­tems, flight plan­ning (ITA Software, ac­quired by Google, pow­ered Google Flights), and sci­en­tific com­put­ing.

Emacs Lisp — mil­lions of peo­ple run Lisp every day with­out know­ing it, be­cause their ed­i­tor is a Lisp in­ter­preter.

Guile/Guix — an en­tire Linux dis­tri­b­u­tion con­fig­ured 100% in Scheme.

Racket has its own an­nual con­fer­ence (RacketCon), an ac­tive aca­d­e­mic and artis­tic com­mu­nity, and is used for lan­guage re­search, for­mal ver­i­fi­ca­tion (Rosette), ty­pog­ra­phy and pub­lish­ing (Pollen), and ed­u­ca­tion around the world.

And new Lisps keep ap­pear­ing: Fennel (a Lisp that com­piles to Lua, pop­u­lar for games), Janet, Hy (a Lisp on top of Python)…

Easter egg for TADC fans:

In The Amazing Digital Circus (episode 8, hjsakldfhl”), when Kinger opens the ter­mi­nal to try to re­set Caine, you can see that Caine (a cre­ative AI built in 1996) is pro­grammed in Lisp. The file is lit­er­ally named Caine-core.lisp.

In The Amazing Digital Circus (episode 8, hjsakldfhl”), when Kinger opens the ter­mi­nal to try to re­set Caine, you can see that Caine (a cre­ative AI built in 1996) is pro­grammed in Lisp. The file is lit­er­ally named Caine-core.lisp.

Installation (5 min­utes)

Go to https://​racket-lang.org

Download the in­staller for your sys­tem (Linux, ma­cOS, Windows).

Open DrRacket, the en­vi­ron­ment that comes in­cluded.

DrRacket has two ar­eas: at the top you write your de­f­i­n­i­tions (your pro­gram), and at the bot­tom you have the REPL for live ex­per­i­men­ta­tion. On the first line of the de­f­i­n­i­tions area, write:

#lang racket

That line tells Racket which lan­guage you’re us­ing (remember: Racket is a lan­guage fac­tory, so you have to pick one).

If you pre­fer the ter­mi­nal: the racket com­mand gives you a REPL, and raco is the pack­age man­ager and tool­ing com­mand.

In the REPL, try:

> (+ 1 2) 3 > (* 3 (+ 2 2)) 12 > (string-append hello world”) hello world”

The rule of Lisp fits in one line:

Everything is (operator ar­gu­ment1 ar­gu­ment2 …). Always. No ex­cep­tions.

Everything is (operator ar­gu­ment1 ar­gu­ment2 …). Always. No ex­cep­tions.

There’s no op­er­a­tor prece­dence to mem­o­rize, no spe­cial syn­tax for any­thing. (+ 1 2) adds. (if …) de­cides. (define …) names. The paren­the­ses that look in­tim­i­dat­ing at first are ac­tu­ally the com­plete ab­sence of ar­bi­trary rules. After a week, you stop see­ing them.

Definitions and func­tions

#lang racket

(define pi-ap­prox 3.14159)

(define (circle-area r) (* pi-ap­prox r r))

(circle-area 2)  ; => 12.56636

de­fine with a name cre­ates a con­stant.

de­fine with (name ar­gu­ments…) cre­ates a func­tion.

Comments start with ;.

Anonymous func­tions use lambda (yes, that lambda Church’s lambda cal­cu­lus from the 1930s is the the­o­ret­i­cal grand­par­ent of all this):

(lambda (x) (* x x))  ; a func­tion with no name ((lambda (x) (* x x)) 5)  ; => 25, ap­plied di­rectly

Lists: the heart of Lisp

Lisp stands for LISt Processing. Lists are the fun­da­men­tal struc­ture:

(list 1 2 3)  ; => (1 2 3) (1 2 3)  ; the same thing, quoted” (first (1 2 3))  ; => 1 (rest (1 2 3))  ; => (2 3) (cons 0 (1 2 3))  ; => (0 1 2 3) (length (a b c))  ; => 3

Notice the quote mark . It tells Racket: don’t eval­u­ate this, it’s data. Hold onto that de­tail it’s the door to the fi­nal trick.

Higher-order func­tions

This is where Racket shines. Passing func­tions to other func­tions is the most nat­ural thing in the world:

(map (lambda (x) (* x x)) (1 2 3 4 5)) ; => (1 4 9 16 25)

(filter even? (1 2 3 4 5 6)) ; => (2 4 6)

(foldl + 0 (1 2 3 4 5)) ; => 15

map trans­forms, fil­ter se­lects, foldl ac­cu­mu­lates. With those three func­tions you can solve most list prob­lems with­out writ­ing a sin­gle for loop.

Recursion: think­ing in spi­rals

In Lisp you don’t think repeat N times” you think what’s the base case, and how do I move to­ward it?”:

(define (factorial n) (if (= n 0) 1 (* n (factorial (- n 1)))))

(factorial 5)  ; => 120

And to make things vi­sual, let’s draw some­thing. Racket ships with graph­ics li­braries in­cluded:

#lang racket (require 2htdp/image)

(define (sierpinski level) (if (= level 0) (triangle 8 solid” purple”) (let ([t (sierpinski (- level 1))]) (above t (beside t t)))))

(sierpinski 6)

Paste it into DrRacket, press Run, and watch the Sierpinski tri­an­gle ap­pear on your screen.

The grand fi­nale: code that writes code

Remember the quote mark : it turns code into data. Watch:

(+ 1 2)  ; => the LIST (+ 1 2), not the num­ber 3 (first (+ 1 2))  ; => the sym­bol + (eval (+ 1 2))  ; => 3. You just eval­u­ated data as code.

Your pro­gram is a list. You can build lists. Therefore: you can build pro­grams with pro­grams. This is ho­moiconic­ity, and it’s why Lisp has real macros not text macros like in C, but func­tions that re­ceive code and re­turn code, be­fore any­thing runs.

Racket does­n’t have a while loop? Let’s in­vent one:

(define-syntax-rule (while con­di­tion body …) (let loop () (when con­di­tion body … (loop))))

(define counter 0) (while (< counter 5) (displayln counter) (set! counter (+ counter 1)))

You just ex­tended the lan­guage ! in Lisp, the syn­tax is yours.

Alan Kay called Lisp the Maxwell’s equa­tions of soft­ware”: a tiny core from which every­thing else can be de­rived.

What next?

How to Design Programs — the book Racket was de­signed around, free on­line.

The Racket Guide — of­fi­cial doc­u­men­ta­tion, among the best out there.

Beautiful Racket — learn to build your own lan­guages.

SICP — the clas­sic of clas­sics, if you want the full en­light­en­ment.

A Kantian Critique of "Sorry" by Justin Bieber

decodingvibes.com

Aug 22, 2026•5 min read

by­Suran­jan Das

pop cul­ture­mu­sicphi­los­o­phynot se­ri­ous

The cen­tral moral ques­tion posed in the 2015 song Sorry” by Justin Bieber, i.e., Is it too late to say sorry?” is ac­tu­ally not a moral ques­tion at all when viewed through a Kantian lens. The ques­tion in it­self con­tains the an­swer: yes, it is too late to say sorry, not in the sense that has any­thing to do with the na­ture of time, but rather the form of the ques­tion it­self shows that the song fun­da­men­tally mis­un­der­stands what an apol­ogy is.

When we con­sider the ques­tion’s im­pli­ca­tions, Bieber is not ask­ing What do I owe to the per­son I have wronged?” Instead, the ques­tion pri­mar­ily asks whether say­ing sorry is still worth any­thing. Utilizing a lan­guage of strat­egy in­stead of duty, such as Could some­one call a ref­eree?” or one more shot at for­give­ness” fur­ther in­di­cates that the ques­tion might not be asked in good faith.

If one’s apol­ogy be­comes a means to­wards an end, then it is a hy­po­thet­i­cal im­per­a­tive. You ex­pect to gain for­give­ness by us­ing the apol­ogy. A moral oblig­a­tion, how­ever, does not de­pend on whether the de­sired con­se­quence re­mains avail­able. The song’s ques­tion of whether it is too late shows that Bieber mea­sures the value of this apol­ogy by what it can ac­com­plish.

One’s ad­mis­sions of wrong­do­ing do not change this ei­ther. He claims I made those mis­takes maybe once or twice”, and then ad­mits, maybe a cou­ple of hun­dred times”. Therefore, Bieber can­not claim ig­no­rance of the fact that he has wronged some­one. However, the song does not seem fo­cused on this. It con­tin­ues, So let me re­deem my­self tonight” and I just need one more shot at sec­ond chances.” The con­cern thus shifts im­me­di­ately from the wrong done to the ap­par­ent con­se­quences suf­fered by the wrong­doer. But moral re­demp­tion is not restor­ing one’s pref­er­ences.

If one was act­ing from duty, their maxim would be sim­pler: I have done wrong; there­fore I ought to ac­knowl­edge it. Whether for­give­ness is pos­si­ble is ir­rel­e­vant, and the same goes for whether he gets an­other chance or not. If, in this form, the apol­ogy is valu­able only be­cause it can se­cure a sec­ond chance or re­demp­tion, then it is not a moral duty; it is merely a po­ten­tially use­ful strat­egy.

Further in the song, Bieber clar­i­fies that he’s missing more than just your body”. This is the real rea­son for which the apol­ogy is be­ing ren­dered. He is not con­fronting the wrong he com­mit­ted; in­stead, he is con­fronting the loss of some­thing he de­sired. His at­tempt to qual­ify his in­ten­tion by claim­ing that his feel­ings of miss­ing en­com­pass more than just your body” does not solve this prob­lem ei­ther. It just tells us that he val­ues the af­fec­tion, per­haps com­pan­ion­ship, and the gen­eral re­la­tion­ship it­self. But the cat­e­gor­i­cal im­per­a­tive re­quires treat­ing an­other per­son as an end in it­self, not as a means to sat­isfy one’s de­sires. Thus, we can ask our­selves: would the apol­ogy still be ren­dered if Bieber knew for sure there is no sec­ond chance or re­demp­tion? If not, then the apol­ogy was ren­dered not be­cause of the wrong done in the first place, but to re­store what the wrong­doer lost.

In later verses, the lyrics seem to re­in­force this sus­pi­cion. The line I’ll take every sin­gle piece of the blame if you want me to” is morally re­veal­ing. If one con­sid­ers them­selves re­spon­si­ble for a wrong done, that re­spon­si­bil­ity can not de­pend on the vic­tim’s re­quest. One ei­ther takes the re­spon­si­bil­ity in good faith, or they do not at all. Furthermore, the song states there is no in­no­cent one in this game for two” which again re­veals that in­stead of of­fer­ing a moral apol­ogy, the song seems to con­sider wrong­do­ing as a con­test in which the guilt of the other party is enough to di­min­ish their own. But an­other per­son’s mis­con­duct can­not al­ter the maxim by which he acted.

Finally, the lyrics can we both say the words and for­get this?” high­light that the ob­jec­tive is clo­sure rather than moral ac­knowl­edg­ment. Bieber qual­i­fies his po­si­tion again with I’m not just tryna get you back on me” but the de­nial does lit­tle to re­solve the prob­lem once more. One’s moral worth ab­solutely can­not be es­tab­lished just by de­clar­ing the pu­rity of their mo­tives. The maxim must be ex­am­ined in­stead. If the maxim is When I have wronged some­one and fear los­ing them, I will apol­o­gize in the hope of re­gain­ing them,” then the apol­ogy re­mains hy­po­thet­i­cal, and more im­por­tantly, fun­da­men­tally self-in­ter­ested.

Thus, the song’s cen­tral ques­tion iron­i­cally re­veals the very moral fail­ure it seems to try to re­pair. Is it too late?” in this con­text trans­lates to is an apol­ogy still ca­pa­ble of get­ting what I want?” Once an apol­ogy is eval­u­ated by whether it can se­cure for­give­ness and rec­on­cil­i­a­tion, it ceases to be an ac­tion grounded in duty. So, yes, it is too late to say sorry, as the song con­ceives it.

This, how­ever, does not stop Bieber from do­ing what moral­ity re­quires: ac­cept­ing the wrong with­out de­mand­ing any­thing in re­turn. Instead of ask­ing whether it was too late to say sorry, the moral thing to do would have been to just say it.

However, even though the ques­tion has a clear ob­jec­tive moral an­swer, we can­not at­tribute any moral judg­ment to Bieber as a hu­man be­ing; we can only an­swer the moral ques­tion as posed in the song. If one were to con­sider an­other in­ter­pre­ta­tion of the song it­self, the ques­tion posed might not be in a moral con­text, but rather as a rhetor­i­cal de­vice, i.e., the artist knows the song is not talk­ing about a sin­cere apol­ogy, but rather it is point­ing out that we can of­ten say sorry in­sin­cerely just be­cause we miss more than their body. This can also ex­plain the in­ten­tional de­ci­sion to pair an up­beat dance track with lyrics that, on the sur­face, seem to be about some­thing much more se­ri­ous.

Note: You can find the pub­lic ver­sion his­tory of my ar­ti­cles on my github: link.

Hook, hold, harvest and hide: Meta’s alleged strategy laid out in first week of landmark trial

www.theguardian.com

Meta’s busi­ness can be boiled down to four words that be­gin with the let­ter H: hook, hold, har­vest, hide, ac­cord­ing to a lawyer who is pros­e­cut­ing the world’s largest so­cial me­dia com­pany.

The owner of Facebook and Instagram hooks” in users, holds” them on its plat­forms for as long as pos­si­ble, harvests” their data and then hides” the truth from the pub­lic, she ar­gued.

Meta’s busi­ness model worked es­pe­cially well for kids,” said Megan O’Neill, a lawyer for the state of California.

Her ac­cu­sa­tion opened the block­buster trial against the US tech com­pany on Tuesday in Oakland, California, just north of Meta’s head­quar­ters in Silicon Valley. California has joined 28 other US states in su­ing the £1tn ($1.36tn) com­pany for al­legedly de­sign­ing ad­dic­tive prod­ucts that lead to chil­dren be­ing harmed.

Eight ju­rors heard from O’Neill and at­tor­neys for Meta this week, along with tes­ti­mony from for­mer em­ploy­ees and a psy­chol­o­gist. The law­suit cen­ters on al­le­ga­tions that the com­pany vi­o­lated US fed­eral child pri­vacy laws and state-level con­sumer pro­tec­tion laws by col­lect­ing data on chil­dren un­der the age of 13 with­out parental per­mis­sion. Over the course of the trial, the jury is ad­di­tion­ally ex­pected to hear from Meta CEO Mark Zuckerberg and Instagram CEO Adam Mosseri.

The threat to Meta is ex­is­ten­tial. If the com­pany is found li­able, dam­ages could be as high as $200bn — an amount equiv­a­lent to the com­pa­ny’s 2025 an­nual rev­enue. The states are also ask­ing that Meta be forced to change the de­sign of its prod­ucts to make them safer for chil­dren, which could have per­ma­nent ef­fects on the com­pa­ny’s busi­ness model and how its so­cial me­dia plat­forms op­er­ate.

Meta has de­nied all al­le­ga­tions. Liza Crenshaw, a spokesper­son for the com­pany, said: Rather than stick­ing to the facts or the law, the states have in­stead de­cided to chase an out­landish pay­out.”

During open­ing state­ments, Paul Schmidt, an at­tor­ney for Meta, said there is no dis­pute” peo­ple can strug­gle with so­cial me­dia, but that Meta had come up with tools to try and ad­dress that”. He added the com­pany does not al­low chil­dren un­der the age of 13 to reg­is­ter for ac­counts on its so­cial net­works and that it had dis­abled more than 1m ac­counts of those young users.

The trial is ex­pected to last six to eight weeks. The pro­ceed­ings will be led by at­tor­neys for the states of California, Colorado, Kentucky and New Jersey. The ju­ry’s role is ad­vi­sory, which means they will give rec­om­men­da­tions to the pre­sid­ing judge, Judge Yvonne Gonzalez Rogers, who will make the fi­nal de­ci­sion on the ver­dict and dam­ages.

Meta faces thou­sands of sim­i­lar US law­suits brought by fam­i­lies, school dis­tricts and other at­tor­neys gen­eral. The com­pany lost the first two of those cases to go to trial in March. In the first, the com­pany was or­dered to pay nearly $1bn to the state of New Mexico for al­low­ing child sex­ual ex­ploita­tion on its plat­forms; and in the sec­ond, it was found li­able for de­lib­er­ately de­sign­ing ad­dic­tive prod­ucts that hooked one young woman and was or­dered to pay her more than $4m.

af­ter newslet­ter pro­mo­tion

The star wit­ness to take the stand in the tri­al’s first week was Arturo Béjar, a safety en­gi­neer at Meta who worked there in two sep­a­rate stints be­tween 2009 and 2021. Since leav­ing, Béjar has been an out­spo­ken critic of the com­pany, tes­ti­fy­ing be­fore a US Senate com­mit­tee and serv­ing as an ex­pert wit­ness in other cases that in­volve so­cial me­di­a’s harm to chil­dren.

In Oakland, Béjar tes­ti­fied that his mo­ti­va­tion for pur­su­ing so­lu­tions for harms to chil­dren was his own teenage daugh­ter’s treat­ment on Instagram. He said she re­ceived un­wanted sex­ual ad­vances and pho­tos of male gen­i­tals as well as misog­y­nis­tic in­sults. Later, she told her fa­ther that re­port­ing these abuses through Instagram’s es­tab­lished processes was ei­ther in­ef­fec­tive or not pos­si­ble.

Meta is tak­ing a don’t ask, don’t tell’ strat­egy” when it comes to child safety, Béjar tes­ti­fied.

Béjar said that his job of­ten in­cluded brief­ing Zuckerberg and that he had spo­ken with the CEO more than 100 times in the course of his work.

During Béjar’s tes­ti­mony, at­tor­neys for the gov­ern­ment showed the jury an email he sent Zuckerberg in 2021, which out­lined a sur­vey he had con­ducted of teens’ ex­pe­ri­ences on Instagram. The re­sults showed 51% of users said yes” to hav­ing bad or harm­ful ex­pe­ri­ences within the pre­vi­ous seven days and that con­tent was taken down only 0.02% of the time.

Béjar tes­ti­fied he sent that data to Zuckerberg be­cause, in my ex­pe­ri­ence, when Mark makes some­thing a pri­or­ity, moun­tains move.”

Did he ever re­spond to you?” the at­tor­ney asked.

No,” Béjar replied. I did­n’t hear back from him.”

Meta fought to bar Béjar from tes­ti­fy­ing at the trial, fil­ing a se­ries of mo­tions to strike his ex­hibits and pre­vent him from tak­ing the stand, all of which were re­jected. In an email to re­porters on Wednesday, Meta con­tin­ued to hound him. The com­pa­ny’s state­ment said Béjar’s tes­ti­mony was not cred­i­ble or re­li­able be­cause he over­in­flated his role at the com­pany and took credit for work he did­n’t do.

After Béjar’s tes­ti­mony wrapped, the jury heard recorded de­po­si­tions from Elena Davis and Natalie Troxel — both for­mer user ex­pe­ri­ence re­searchers for Meta. Jean Twenge, a psy­chol­ogy pro­fes­sor at San Diego State University, also briefly took the stand, with tes­ti­mony sched­uled to con­tinue next week.

The New MCP Roadmap

blog.modelcontextprotocol.io

Today we’re ex­cited to pub­lish an up­dated roadmap for the Model Context Protocol (MCP), cov­er­ing the next spec­i­fi­ca­tion re­lease and be­yond.

It sets the di­rec­tion for pro­to­col work over the com­ing months and was de­vel­oped by the Core Maintainers to­gether with our com­mu­nity of main­tain­ers and Working Groups.

Explore the roadmap →

Looking back

The pre­vi­ously pub­lished roadmap came out in March with four pri­or­ity ar­eas: trans­port evo­lu­tion and scal­a­bil­ity, agent com­mu­ni­ca­tion, gov­er­nance mat­u­ra­tion, and en­ter­prise readi­ness. We’ve made sig­nif­i­cant progress in all of these over the past five months.

The bulk of the changes landed in the 2026 – 07-28 spec­i­fi­ca­tion re­lease - you might’ve al­ready seen them in our SDKs and doc­u­men­ta­tion. The im­prove­ments ranged from mi­nor mod­i­fi­ca­tions to ma­jor pro­to­col over­hauls.

One of the biggest changes we shipped is that pro­to­col-level ses­sions and the ini­tial­iza­tion hand­shake are gone, so a server can scale hor­i­zon­tally with­out hold­ing state (SEP-2575, SEP-2567). Additionally, clients can now call server/​dis­cover to learn a server’s sup­ported ver­sions and ca­pa­bil­i­ties be­fore do­ing any­thing else. List re­sults are also cacheable (SEP-2549).

On the agent com­mu­ni­ca­tion side, Tasks were re­worked based on early adopter feed­back - we moved them into an of­fi­cial ex­ten­sion (SEP-2663). The brand-new Multi Round-Trip Requests pat­tern (SEP-2322) re­placed server-ini­ti­ated re­quests so that elic­i­ta­tion and sim­i­lar flows work on state­less servers.

The Server Card Working Group con­tin­ues to work through the .well-known meta­data con­ven­tions for MCP servers, so a server can be dis­cov­ered and rea­soned over with­out con­nect­ing to it.

Governance has evolved as well. We for­mally adopted a Contributor Ladder, Working Groups now triage SEPs in their own area, and the spec­i­fi­ca­tion has a proper fea­ture life­cy­cle and dep­re­ca­tion pol­icy that the 2026 – 07-28 dep­re­ca­tions were the first to fol­low.

Enterprise readi­ness was heav­ily fo­cused on se­cu­rity in the past re­lease cy­cle, and as ex­pected most of this work ar­rived as au­tho­riza­tion im­prove­ments: is­suer val­i­da­tion, is­suer-bound client cre­den­tials, and Client ID Metadata Documents (CIMD) as the pre­ferred reg­is­tra­tion path for clients, with Enterprise-Managed Authorization avail­able as an ex­ten­sion (which is also now sta­ble).

This is sig­nif­i­cant progress in a very short span. The up­dated roadmap picks up from here.

Priority ar­eas

The new roadmap is or­ga­nized into five pri­or­ity ar­eas. Several of them pick up work that the pre­vi­ous ver­sion of the roadmap listed as be­ing on the hori­zon, in­clud­ing server-ini­ti­ated events, re­sult type im­prove­ments, and agent iden­tity, which have since ma­tured enough to be­come pri­or­i­ties in their own right. Each area has a set of Core Maintainers re­spon­si­ble for it and one or more Working Groups.

Agentic mes­sag­ing prim­i­tives

Modern agen­tic work­loads no longer fit the stan­dard re­quest-and-re­sponse pat­tern. Loops can run for longer, servers can push streamed re­sults, and there is a clear need to steer work mid-flight. MCP has been grow­ing to meet these re­quire­ments, in­tro­duc­ing Tasks, sub­scrip­tions/​lis­ten, and progress no­ti­fi­ca­tions. We want to make sure that we not only of­fer the right prim­i­tives for the job, but also that they work well to­gether. The work here spans server-ini­ti­ated events (webhooks and chan­nels, so clients aren’t left polling for re­sults), a com­po­si­tion re­view across the Agents, Transports, and Triggers & Events Working Groups, and ma­tur­ing the Tasks ex­ten­sion (SEP-2663) so it can move into the spec­i­fi­ca­tion.

HTTP-native trans­port uni­fi­ca­tion and hard­en­ing

With the 2026 – 07-28 re­lease, a re­mote MCP server is now no dif­fer­ent from any other HTTP work­load, mak­ing it easy to host and op­er­ate one on any in­fra­struc­ture that de­vel­op­ers and or­ga­ni­za­tions al­ready use for their APIs and ser­vices. This ap­proach has proven to scale, and we want to stretch it to cover other de­ploy­ment modes as well, in­clud­ing lo­cal servers speak­ing Streamable HTTP over stdio. Unifying on one trans­port lets us sim­plify MCP server and client de­vel­op­ment even fur­ther.

Agent iden­tity and en­ter­prise-ready se­cu­rity

MCP au­tho­riza­tion to­day is built around a per­son ap­prov­ing ac­cess in a browser. That works well for in­ter­ac­tive clients, but more and more of the callers are agents run­ning as cloud work­loads with their own iden­tity, act­ing on be­half of a user who is­n’t pre­sent, or del­e­gat­ing nar­rower au­thor­ity to sub-agents. We want MCP servers to have a stan­dard­ized way to rec­og­nize and trust those agent iden­ti­ties, built on ex­ist­ing stan­dards rather than pasted API keys and long-lived to­kens.

The work here cov­ers fi­nal­iz­ing Demonstrating Proof of Possession (DPoP) and dri­ving its adop­tion, and defin­ing an opin­ion­ated path for agent iden­tity and del­e­ga­tion through Workload Identity Federation, the ID-JAG grant be­hind Enterprise-Managed Authorization, and stan­dard to­ken ex­change. We will also con­tinue to grow our en­gage­ment with the OAuth stan­dards bod­ies, in­clud­ing the IETF OAuth and WIMSE work­ing groups, to help the un­der­ly­ing stan­dards evolve with the build­ing blocks that agent iden­tity needs.

Improved prim­i­tives

Tool call­ing is the part of MCP most de­vel­op­ers touch first, and it has held up well over the life­time of the pro­to­col. Where it falls a bit short, how­ever, is in the re­sult han­dling. A tools/​call re­sponse can carry the same out­put in more than one form, and a server de­vel­oper to­day has no way to know which form a given client will put in front of the model. We aim to make this eas­ier by stan­dard­iz­ing on one clear con­tract.

The other chal­lenge we need to ad­dress for prim­i­tives is their ever-grow­ing scale. Connecting to a server with a hun­dred tools means the model pays for that en­tire sur­face be­fore the user has asked a sin­gle ques­tion, and tool se­lec­tion tends to get worse as the list grows. We’re start­ing a pro­gres­sive dis­cov­ery ef­fort so a server can of­fer a small en­try point and re­veal more of its cat­a­log as the con­ver­sa­tion nar­rows.

Improved SDK de­vel­oper ex­pe­ri­ence

Our SDKs are how de­vel­op­ers ex­pe­ri­ence MCP. We are in­vest­ing in their er­gonom­ics and their con­for­mance with the spec­i­fi­ca­tion, and in mak­ing them in­tu­itive and well-doc­u­mented across every plat­form and lan­guage we sup­port. This is even more im­por­tant now that many de­vel­op­ers build MCP clients and servers by point­ing an agent at our li­braries, where clear APIs and ac­cu­rate docs de­cide whether the code will work with min­i­mal fric­tion.

Proposal pri­or­i­ti­za­tion

Specification Enhancement Proposals (SEPs) that fall within these pri­or­ity ar­eas get ex­pe­dited re­view and have the best chance of ac­cep­tance. Proposals out­side them aren’t re­jected au­to­mat­i­cally, but main­tainer re­view time is scarce and goes to these ar­eas first.

If you’re con­sid­er­ing a SEP, iden­tify the pri­or­ity area it be­longs to, raise it with the rel­e­vant Working Group, and work with its mem­bers to shape your pro­posal. Each area on the roadmap names the Core Maintainers re­spon­si­ble for it, and any­one in­ter­ested in con­tribut­ing can reach them on Discord. We’re ex­cited to work with the com­mu­nity to re­view and build on the pro­pos­als that sup­port this roadmap.

Get in­volved

Every pri­or­ity area above has a Working Group be­hind it or form­ing around it, and all of them have room for more con­trib­u­tors. There are sev­eral ways to par­tic­i­pate:

Join a Working Group or Interest Group: see the Working and Interest Groups page and the com­mu­nity chan­nels.

Propose or com­ment on a SEP: read the SEP guide­lines, then open one or weigh in.

Start an ex­per­i­men­tal ex­ten­sion: SEP-2133 lets any WG or IG ex­per­i­ment in an ex­per­i­men­tal-ext- repos­i­tory be­fore a for­mal SEP.

Contribute di­rectly: the con­tribut­ing guide cov­ers the spec­i­fi­ca­tion, SDKs, and tool­ing.

We look for­ward to grow­ing and evolv­ing MCP to­gether!

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.