10 interesting stories served every morning and every evening.

System Prompts

platform.claude.com

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Client Challenge

support.mozilla.org

A re­quired part of this site could­n’t load. This may be due to a browser ex­ten­sion, net­work is­sues, or browser set­tings. Please check your con­nec­tion, dis­able any ad block­ers, or try us­ing a dif­fer­ent browser.

A Third World Embedded Engineer Responds to "RISC-V: They Should Have Known Better"

rvembedded.com

Dmitry Grinberg pub­lished a long piece ex­plain­ing his dis­taste for RISC-V, you can read his ar­ti­cle here: RISC-V: They Should Have Known Better - Dmitry.GR. It went to the front page of Hacker News and it started a good ar­gu­ment on Lobsters. It is the most sub­stan­tial crit­i­cism the ar­chi­tec­ture has had in a while and though I switched my en­tire stack away from STM32 and ARM to RISC-V and did a video on it about a year ago Good­bye STM32 ARM — Meet the CH32 RISC-V Chips That Replaced It! , part of me is in­fu­ri­ated be­cause so much of what he said seems like a bi­ased per­spec­tive.

Look, I am not go­ing to de­fend the ISA com­mit­tee, RISC-V in­ter­na­tional de­nied me mem­ber­ship to their golden tower. On the ar­chi­tec­ture it­self, the com­pressed store off­sets re­ally are strange, Zicsr re­ally should not be a sep­a­rate thing you have to re­mem­ber to ask for, I have hit every one of these and I have writ­ten a book thats about 80% com­plete about hit­ting them on the CH32V003, which is one of the very RV32E type chip he men­tions.

Maybe I should say where I am writ­ing from, be­cause it changes which parts of this ar­gu­ment look im­por­tant from my per­spec­tive.

I work out of Trinidad and Tobago, a small is­land na­tion off the coast of Venezuela. When I want a de­vel­op­ment board I am not click­ing through to next day de­liv­ery, I am check­ing whether the seller ships here at all, what cus­toms will do to it (if I get it at all), and what the to­tal lands at in TT dol­lars. Free Shipping” from Digikey, Mouser or any US or European man­u­factuer dosen’t ap­ply to me. I pay any­where from US $60 to US $200 to ship one dol­lar chips that peo­ple every­where else get free ship­ping on. In fact a well known PCB com­pany who reached out to me con­sid­er­ing spon­sor­ship turned me down so­ley based on ship­ping to my lo­ca­tion. Have a look here:

The stu­dents I want to teach are in the same po­si­tion, and so are the ones in Nigeria and Bangladesh and every­where else the peo­ple in the in­dus­try does not think about when it writes its blog posts. From that po­si­tion, the dif­fer­ence be­tween a ten cent part and a one dol­lar part is not a round­ing er­ror and it is not a de­tail you get to wave past on the way to the in­ter­est­ing dis­cus­sion about en­cod­ings. It is the dif­fer­ence be­tween a class of thirty stu­dents each hav­ing their own chip and a class of thirty stu­dents watch­ing one demo board if any at all. Instruction set el­e­gance is a thing you can af­ford to care about once the hard­ware is al­ready on your desk. Whether the hard­ware can get to your desk at all comes first. That is why the para­graph most peo­ple scrolled past is, to me, the most im­por­tant one in the ar­ti­cle.

Grinberg missed that part that RISC-V cre­ates a space for the other 99% out­side of the world” (which in this space world” is mainly the US and Europe) and it has noth­ing to do with ar­chi­tec­ture.

He Derives the Requirements and Lands on RV32EC

Before the in­ter­rupt arith­metic, be­fore the en­cod­ing com­plaints, he does some­thing care­ful. He asks what a cheap mi­cro­con­troller core is ac­tu­ally for. His an­swer is that it sits in­side a larger chip prod­ding reg­is­ters and con­fig­ur­ing hard­ware blocks, in an MP3 player, an SD card, a USB stick”. The real work is done by cus­tom sil­i­con around it. From that he de­rives what such a core needs. Low in­ter­rupt la­tency a small die area and good code den­sity, be­cause the code lives in ROM or SRAM and both are ex­pen­sive per byte. No hard­ware di­vider, pos­si­bly not even a mul­ti­plier, since you are not do­ing much arith­metic. No priv­i­lege sep­a­ra­tion, be­cause noth­ing un­trusted ever runs there.

Then he writes the line him­self:

But,” you might say, you just de­scribed RV32IC (or RV32EC)!”

But,” you might say, you just de­scribed RV32IC (or RV32EC)!”

And ear­lier, plainly:

I am 100% sure that RISC-V will own the cheap-as-dirt sin­gle-use mi­cro­con­troller space even­tu­ally.

I am 100% sure that RISC-V will own the cheap-as-dirt sin­gle-use mi­cro­con­troller space even­tu­ally.

So the most cred­i­ble RISC-V critic of the month sat down, worked out from first prin­ci­ples what a cheap mi­cro­con­troller core should be, ar­rived at the in­struc­tion set a ten cent chip im­ple­ments, and stated that this seg­ment is go­ing to be RISC-V’s.

He de­rives the case for the chip and then spends the rest of the ar­ti­cle an­noyed that the chip ex­ists.

This is al­most satir­i­cal.

His quar­rel is with whether that out­come was earned. That is a real ques­tion and I un­der­stand why it both­ers him. It is not, how­ever, a ques­tion that af­fects any­body de­cid­ing what to learn on, be­cause the chip is on the shelf ei­ther way.

Where I Actually Disagree, Strongly.

His cen­tral claim is the first one in the ar­ti­cle, and it is big­ger than any of the en­cod­ing com­plaints:

Simply put, the things a high-end CPU needs are di­a­met­ri­cally op­posed to the things a small cost-sav­ing mi­cro­con­troller core needs.

Simply put, the things a high-end CPU needs are di­a­met­ri­cally op­posed to the things a small cost-sav­ing mi­cro­con­troller core needs.

The con­clu­sion he draws is that no sin­gle ISA can serve both ends, and that RISC-V fans are fool­ing them­selves, in the­ory the premise is true. The con­clu­sion does not fol­low, and I can show you why from three parts sit­ting on my desk as we speak.

CH32V003. This is the cheap RV32EC with six­teen reg­is­ters, no mul­ti­plier, no di­vider, ma­chine mode only, 2KB of SRAM, 16KB of flash, ten cents, it’s EXACTLY the core he spec­i­fied. I shipped two prod­ucts with these, one is a bin mon­i­tor that has a ToF sen­sor, an LED and an air tag. The other is an agri­cul­tural prod­uct for a client that opens and closes a door at a cer­tain time. It also makes a good throw away part, as I show case in my whis­tle switch Clap Switch Is Dead. Here’s the RISC-V Powered Whistle Switch! and which in my view is the BEST part to re­place the over­priced, out­dated Arduino Did Arduino Q Ruin Arduino? - Here’s how to Switch to RISC-V with the CH32V003.

CH32H417. A dual core MCU that is un­matched in per­for­mance to price point and is at the higher end of the MCU line of things. It has a QingKe V5F at 400 MHz along­side a V3F at 144 MHz, 896KB of SRAM, 960KB of flash. USB 3.2 Gen1 with an in­te­grated 5 Gbps trans­ceiver, 100M Ethernet MAC and PHY, a SerDes iso­lated trans­ceiver, a 500 MB/s high speed in­ter­face, SDMMC, a cam­era in­ter­face, a dis­play con­troller, a graph­ics ac­cel­er­a­tor etc etc. I got a web browser run­ning on this thing I Built a Web Browser on a RISC-V Microcontroller (No Linux) Quantum en­tropy based GAN cat gen­er­a­tion Schrödinger’s De/Motivational Quantum Cat: GAN Image Generation on CH32 RISC-V Microcontroller and real-time fa­cial recog­ni­tion Real Time Facial Recognition on The Edge With CH32H417 RISC-V MCU in un­der 150KB of ram. I got a host of other pro­jects run­ning but those are just some I got time to record and put up.

Baochip. A VexRISC-V with an MMU built around a stack thats open from sil­i­con to os Baochip-1x: A Mostly-Open, 22nm SoC for High Assurance Applications « bun­nie’s blog, that runs Xous be­trusted-io/​xous-core: The Xous mi­cro­ker­nel de­signed by leg­endary hard­ware hacker bunnie” Huang , a Rust mi­cro­ker­nel with real process iso­la­tion. Privilege sep­a­ra­tion, the ex­act thing he says the cheap end does not need and there­fore does not get. In ad­di­tion to Xous it also sup­ports op­er­at­ing sys­tems like SEL4 vk2seb/bao1x-seL4: seL4 port to baochip-1x and Linux pkoscik/​baochip-linux: An at­tempt to boot main­line Linux on a stock Dabao board. I wrote the bare metal C SDK for the chip Arm­strong­Subero/​dabao-sdk: Bare metal C SDK for the Baochip-1x RISC-V SoC and it was of course the chip in­side the badge of DEFCON 34 The New Defcon Badges Pack a Unique Open Source Chip That Doubles as a Security Key | WIRED this year.

I can also point to the NES em­u­la­tor I wrote for the $1 ESP32C3 RISC-V based chip NES Emulator on $1 ESP32-C3 RISC-V Microcontroller, or ex­per­i­ment­ing with Linux on the Orange Pi RV2 OrangePi RV2 5 Minute Unboxing and Setup | RISC-V Ubuntu Linux that takes 5 min­utes to setup and has been run­ning since the day I boot it up.

Point is I could go on and on about how di­verse and ac­ces­si­ble cur­rently ship­ping RISC-V parts are, but then we’ll be stray­ing too much from the topic at hand.

I linked all those to say this, that all these parts all have the same base in­struc­tion set and I gained ex­per­tise in all in un­der a year and un­der US $100 across the en­tire stack, from dis­posi­ble sil­i­con to PC level, of course mi­nus data cen­ter com­pute.

For un­der US $100 in­clud­ing ship­ping I was able to ex­plore an en­tire ver­ti­cal stack us­ing one ar­chi­tec­ture. Due to the AI race the OrangePi RV2 has now gone up in price but at re­lease it cost $30 and shipped free. For about 7 dol­lars I got 50 CH32V003s with a de­bug­ger, the CH32H417 board is $20 on ana­log lamb and uses the same cheap (and of­fi­cial) de­bug­ger for the CH32V003 and the Baochip Dabao board (which I wrote a book about by the way check it out here (The Dabao Book - Payhip) was $9.50 on crowd sup­ply when I bought it, two with ship­ping from crowd sup­ply cost me $35, un­der $100 in to­tal. A de­bug­ger for an ARM part alone a Segger J-Link costs about $600, though I guess for that $100, and add an­other $100 to ship,so about $200 I could get an EDU edi­tion J-link and no chips or boards. Yaay.

Back to RISC-V, across all these parts, the base set is the same. So that means the same reg­is­ter model, same call­ing con­ven­tion, same tool­chain. Yes the ex­ten­sions dif­fer, but the thing is what I learned writ­ing as­sem­bly on the ten cent CH32V003 part did not stop be­ing true on any of the oth­ers. A dual core MCU, an SBC run­ning Linux or an ad­vanced cus­tom se­cu­rity chip run­ning a novel op­er­at­ing sys­tem. My skills were trans­ferrable to the point that in each case within a few hours I had tool­chains setup, could fo­cus on my ap­pli­ca­tions and when de­bug­ging I felt at home. All I need to work with them is the ISA man­ual and a C com­piler.

Now price the same jour­ney on the other side, for­get x86 – 64 and that du­op­oly, patent mine­field, with multi-thou­sand dol­lar de­bug probes; we’ll take a look at ARM.

The equiv­a­lent to the CH32V003 is the Cortex-M0 is ARMv6-M so some­thing like an STM32F030, step it up we have a Cortex-M7 which is ARMv7-M, to get an MMU in a part for Linux or SEL4 and Xous, you’re look­ing at an ap­pli­ca­tion proces­sor like the ARMv8-A.  These are dif­fer­ent Arm pro­files with sig­nif­i­cantly dif­fer­ent priv­i­lege, ex­cep­tion, and sys­tem mod­els, so mov­ing up the stack in­volves sub­stan­tially more re­learn­ing than sim­ply en­abling an­other RISC-V ex­ten­sion. Trust me I’ve used them all.

And at the top of that range the gap is not even about learn­ing curves. There is no Cortex-M mi­cro­con­troller with an in­te­grated USB 3.0 SuperSpeed PHY. The near­est dual core Arm part is an STM32H747, which is a fine chip and does not have one. If you need USB 3.0 you leave the mi­cro­con­troller class en­tirely: an i.MX 8 or an RK3xxx, which means Cortex-A. You want an MMU, Linux, DDR, a PMIC, and a board you are not lay­ing out in two lay­ers. Or you keep the M7 and add an ex­ter­nal bridge chip.

The H417 eval­u­a­tion board is around twenty dol­lars. The H747 in TFBGA240 car­ries a twenty week man­u­fac­turer lead time, chip only, costs about the same, be­fore you have any­thing to plug in, and Mouser asks for ID be­fore you can or­der, Digikey has also been known to deny peo­ple parts de­pend­ing on where they are and their name as Hussein Ali, well known Youtuber from NorthridgeFix de­scribes Star­link Repair - Digi-key re­fused my or­der.. Oh and it’s about US $60 – 100+ to ship to my lo­ca­tion. I can pick up H417s on the of­fi­cial WCH store on Aliexpress with free ship­ping and no ver­i­fi­ca­tion hul­la­balu. We haven’t even started talk­ing about the Cortex-A parts that have MMUs or thier de­bug­ging tools and ecosys­tem frag­men­ta­tion.

The Boundary Is Not Technical

Here is the part that un­der­cuts his fram­ing most di­rectly, and it has noth­ing to do with en­cod­ings. He treats the gap be­tween a small core and a large one as an ar­chi­tec­tural fact, some­thing that falls out of op­posed re­quire­ments. On ARM chips it is not an ar­chi­tec­tural fact. It is a PRODUCT bound­ary, and it is en­forced by li­cens­ing. Has any­one tried adding an MMU to a Cortex-M? The phys­i­cal trade­offs are real, the dif­fer­ence is that with RISC-V, the ISA owner does not de­cide for you where that bound­ary must be drawn. If you want vir­tual mem­ory on ARM you li­cense a Cortex-A in­stead, which is a dif­fer­ent core fam­ily, a dif­fer­ent pro­file, a dif­fer­ent ne­go­ti­a­tion, and a dif­fer­ent roy­alty. There is no in­cre­men­tal path. there is a wall, with a sales team on the other side of it.

Compare what hap­pened with Baochip. The RISC-V priv­i­leged spec­i­fi­ca­tion de­fines su­per­vi­sor mode and Sv32 pag­ing as op­tional things an im­ple­men­ta­tion may pro­vide. VexRISC-V is an open core, some­body added an MMU to it. bun­nie built a chip around it and runs a mi­cro­ker­nel with real process iso­la­tion on it that me in Trinidad a coun­try who’s name does not even come up in ISA cir­cles can ex­per­i­ment with at low cost and teach to other peo­ple in the re­gion.

That’s what free­dom looks like.

Nobody asked per­mis­sion, no­body signed any­thing, no­body pays a roy­alty per unit shipped and any­body can learn down to the RTL the sil­i­con is built on.  So when Grinberg in his ar­ti­cle lists priv­i­lege sep­a­ra­tion among the things the cheap end does not need and there­fore does not get, it is de­scrib­ing a prop­erty of ARMs prod­uct seg­men­ta­tion and at­tribut­ing it to in­struc­tion set de­sign. On RISC-V it is a check­box in the priv­i­leged spec, you leave it off in a ten cent part be­cause it costs area you do not want to spend, and you turn it on when you do, and the in­struc­tion set un­der­neath is the same ei­ther way.

That is the real dif­fer­ence be­tween the two ecosys­tems, and it is why one ISA can­not serve both ends” reads dif­fer­ently de­pend­ing on which side you are stand­ing on. On one side the ends are sep­a­rated by physics and cost, on the other they are sep­a­rated by physics, cost, and a con­tract.

The Thing He Calls Fragmentation

Before I close I want to ad­dress his stance on frag­men­ta­tion. He is not wrong that the ex­ten­sion mech­a­nism frag­ments the stan­dard. Zcb split­ting off from C is an­noy­ing and Zicsr not be­ing im­plied by the base is an­noy­ing. Vendors adding pro­pri­etary in­ter­rupt hard­ware does frag­ment things fur­ther, I learned first hand port­ing NuttX to the CH32V307 Porting Apache NuttX RTOS to the WCH CH32V307: A Deep Dive into the PFIC and Everything That Went Wrong.

But that mech­a­nism is the an­swer to his own open­ing ques­tion. The rea­son one in­struc­tion set can sit in a ten cent part with six­teen reg­is­ters and also in a chip run­ning a pro­tected multi-process op­er­at­ing sys­tem is pre­cisely that the small part is not car­ry­ing the large part’s bag­gage. There is no com­pro­mise core in the mid­dle serv­ing both badly, which is what diametrically op­posed re­quire­ments” would nor­mally force. Fragmentation and scal­a­bil­ity are the same prop­erty, you do not get one with­out the other and whether the trade­off was worth it is a fair ar­gu­ment and I do not think it has an ob­vi­ous an­swer.

What I do think is that he is right about the im­por­tant part, and right in a way that favours the thing he is crit­i­cis­ing. RISC-V is not go­ing to take the cheap mi­cro­con­troller space be­cause its en­cod­ing is el­e­gant. It is go­ing to take it be­cause the part costs ten cents, and be­cause the lad­der above it is the same in­struc­tion set all the way up. It is go­ing there be­cause an em­bed­ded en­gi­neer in a 3rd world coun­try can shine a cheap LED and see the tran­sis­tors in the sil­i­con, In­fra-Red, In Situ (IRIS) Inspection of Silicon « bun­nie’s blog and get 50 chips with a de­bug­ger and free de­vel­op­ment tools for the price of a cup of cof­fee and shipped free. It also means that world class en­gi­neers can de­sign MMUs onto chips that the gate keep­ers will never give a li­cense for.

He writes that this will hap­pen not due to its ISA de­sign, but de­spite it,” and he means it as a mild in­dict­ment. Read it from here and it is not one. Winning on price and avail­abil­ity is not a lesser way to win. It de­cides who is in the room. An ar­chi­tec­ture that ar­rives in my coun­try at ten cents a part, with an open tool­chain and no li­cense to ne­go­ti­ate, puts em­bed­ded sys­tems within reach of peo­ple who were pre­vi­ously go­ing to watch some­body else’s demo board and con­sume thier prod­ucts with­out ever be­ing able to match what they have ac­cess to. That’s the power of free­dom, open­ness and is democ­racy in it’s truest sense.

That is a bet­ter rea­son than el­e­gance. and I want to tell Mr Grinberg, that the word priv­iledge he tosses around in his ar­ti­cle also ex­tends be­yond the ISA de­pend­ing on where you are in the world.

Nuff said.

Armstrong Subero is an em­bed­ded sys­tems en­gi­neer and pub­lished au­thor with Apress/Springer. He builds the Rovari RISC-V ed­u­ca­tion plat­form from Trinidad and Tobago.

Sorry...

scholar.google.com

We’re sorry…

… but your com­puter or net­work may be send­ing au­to­mated queries. To pro­tect our users, we can’t process your re­quest right now.

Models Are Getting Dumber on Purpose

w4g1.dev

Reasoning scores keep climb­ing while per-to­ken com­pute keeps drop­ping. GLM-5.2 scores 99.2% on AIME 2026 with about 40 bil­lion pa­ra­me­ters ac­tive per to­ken. Qwen3.5 scores 91.3% with 17 bil­lion ac­tive. DeepSeek V4-Flash runs 13 bil­lion ac­tive. For scale, GPT-4 was ru­mored to run around 280 bil­lion ac­tive pa­ra­me­ters in 2023, and it could barely solve an AIME prob­lem. At the small end, Qwen3.5 9B fits in 6GB of VRAM quan­tized and roughly dou­bles the score of the next best model un­der 10B pa­ra­me­ters on Artificial Analysis’s in­tel­li­gence in­dex. If you only looked at math and code bench­marks, you’d con­clude that mod­els are get­ting smarter per pa­ra­me­ter at an ab­surd rate.

They are, on those bench­marks. Ask the same mod­els a plain fac­tual ques­tion and the pic­ture flips. On SimpleQA, a bench­mark of fac­tual re­call with no tools al­lowed, the cur­rent leader is Gemini 2.5 Pro at 53%, so the best re­call money can buy still misses half the ques­tions. The small mod­els barely reg­is­ter. Artificial Analysis mea­sures Qwen3.5 4B and 9B at hal­lu­ci­na­tion rates of 80 to 82% on its knowl­edge bench­mark, which means that when they don’t know a fact, which is most of the time, they make one up. Ask the 9B for the birth year of a mi­nor 19th-century math­e­mati­cian and you get a con­fi­dent, plau­si­ble, wrong an­swer. The pa­ra­me­ter count did­n’t drop for free. Labs are trad­ing world knowl­edge for rea­son­ing skill, and the trade is de­lib­er­ate.

What the pa­ra­me­ters were for

Facts take space. Research on knowl­edge ca­pac­ity (the Physics of Language Models” se­ries has the clean­est mea­sure­ments) puts it on the or­der of two bits of fac­tual knowl­edge per pa­ra­me­ter. If you want a model that knows the birth year of every mi­nor Wikipedia fig­ure, the pop­u­la­tion of every Dutch mu­nic­i­pal­ity, and the ar­gu­ment or­der of every func­tion in every npm pack­age, you pay for that in weights, and it’s a big part of why fron­tier mod­els grew to tril­lions of pa­ra­me­ters.

Reasoning com­presses much bet­ter than facts do, be­cause it’s a rel­a­tively small set of pro­ce­dures ap­plied over and over: break the prob­lem into parts, track in­ter­me­di­ate state, check your own work, back­track when a step fails. Distillation and re­in­force­ment learn­ing on ver­i­fi­able tasks turn out to trans­fer those pro­ce­dures into small mod­els re­mark­ably well. Phi-4 is 14 bil­lion pa­ra­me­ters, trained heav­ily on syn­thetic text­book-style data, and it’s good at math and bad at trivia, which tells you ex­actly what its train­ing data con­tained. That mix used to look like a lim­i­ta­tion of the syn­thetic-data ap­proach. It now looks like the de­sign goal.

The knowl­edge that sur­vives the trade has a shape. These mod­els are gen­er­al­ists: they know a lit­tle about nearly every­thing and al­most noth­ing in depth. Ask one about PostgreSQL and it knows what it is, what it’s good at, and roughly how MVCC works, but ask which ver­sion added a spe­cific plan­ner fea­ture and you’re back to in­vented facts. That’s the right layer to keep in weights, be­cause breadth is what lets a model un­der­stand what a ques­tion is about, know what to look up, and judge whether a source is plau­si­ble. The depth is cheap to re­trieve and ex­pen­sive to store, so it’s the part that goes.

Facts rot, pro­ce­dures don’t

A fron­tier train­ing run takes months and costs hun­dreds of mil­lions of dol­lars, and the mo­ment it fin­ishes, the facts in­side it start go­ing stale. Library APIs change, prices change, peo­ple change jobs, and half of what a 2024 model be­lieved about the JavaScript ecosys­tem was out­dated be­fore the model shipped. Every fact you bake into weights has a shelf life, and the only way to re­fresh it is an­other train­ing run.

The pro­ce­dures don’t rot. Algebra worked the same way in 1970 as it does now, and so does break­ing a prob­lem down or spot­ting a con­tra­dic­tion be­tween two sources. A model that’s mostly pro­ce­dure and only lightly loaded with facts does­n’t age the way a knowl­edge-heavy model does. Its train­ing cut­off mat­ters much less, be­cause the cur­rent state of the world was never sup­posed to live in the weights in the first place. I think this is the best ar­gu­ment for the whole ap­proach: it de­cou­ples the ex­pen­sive, slow ar­ti­fact (the trained model) from the thing that changes daily (what’s true).

The har­ness car­ries the knowl­edge

If the model does­n’t know things, some­thing else has to, and that some­thing is the har­ness: re­trieval over a knowl­edge base, tool calls, web search, a filesys­tem full of docs. I wrote ear­lier that Rust is a har­ness for agents, a source of cheap ma­chine-check­able feed­back. This is the same shape from the other side. The model con­tributes rea­son­ing, and every­thing it rea­sons about gets sup­plied at run­time.

You can al­ready watch agents work this way. A cod­ing agent does­n’t need to have mem­o­rized your de­pen­den­cy’s API sur­face, be­cause it greps node_­mod­ules or reads the docs be­fore call­ing any­thing, and its an­swer is grounded in the ver­sion you ac­tu­ally have in­stalled rather than whichever ver­sion dom­i­nated the train­ing data. The re­call that used to be a fixed cost in every for­ward pass be­came an on-de­mand lookup.

A fron­tier model on your GPU

Follow the trend a cou­ple of years out and I think we get a model with fron­tier-qual­ity rea­son­ing, Fable-quality, that runs on a sin­gle con­sumer GPU. The com­pute half is nearly there. DeepSeek V4-Flash rea­sons with about 13 bil­lion ac­tive pa­ra­me­ters per to­ken, well within con­sumer-GPU range. What does­n’t fit is the other 271 bil­lion pa­ra­me­ters sit­ting in its ex­perts, and ex­pert lay­ers are mostly fact stor­age. That’s the part this whole trade makes op­tional. Strip the knowl­edge out and to­tal size shrinks to­ward ac­tive size, and a 20 to 40B model at 4-bit quan­ti­za­tion fits on the 24GB card that’s been sit­ting in gam­ing PCs since 2022.

The catch is that it won’t know much. Ask it a bare fac­tual ques­tion with no tools at­tached and the right be­hav­ior is to say it does­n’t know and go look it up. Paired with a de­cent har­ness, that’s most of what I use a fron­tier model for to­day, run­ning lo­cally with no per-to­ken bill and no data leav­ing the ma­chine.

This mostly solves hal­lu­ci­na­tion

The part I find most promis­ing is what this does to hal­lu­ci­na­tion. When a fact lives in weights, a wrong fact is un­find­able and un­fix­able. You can’t grep the weights, you can’t diff them against last month, and cor­rect­ing one er­ror means a fine-tune that might break who knows what else. The model states the wrong fact with the same flu­ent con­fi­dence as a right one, and there’s no ar­ti­fact to check it against.

When the fact lives out­side the model, a wrong an­swer has an ad­dress. The model cites a doc­u­ment, so you can open the doc­u­ment. If the doc­u­ment is wrong, you edit the doc­u­ment, and every fu­ture query gets the cor­rec­tion, which beats wait­ing for the next train­ing run by roughly a year. Retrieval does­n’t get you to zero, since a model can still mis­read a source or stitch two of them to­gether wrong, but a claim with a source is check­able and a claim from weights is­n’t. A wrong fact in a knowl­edge base is an or­di­nary data bug, the kind we al­ready know how to trace, fix, and write a re­gres­sion test for.

There’s a ver­sion of this fu­ture where the model card stops list­ing a knowl­edge cut­off at all, be­cause what’s left in the weights goes stale on a scale of years in­stead of weeks. The model just gets handed the world’s cur­rent state at run­time, the same way a CPU gets handed a pro­gram.

Bloomberg - Are you a robot?

www.bloomberg.com

We’ve de­tected un­usual ac­tiv­ity from your com­puter net­work

To con­tinue, please click the box be­low to let us know you’re not a ro­bot.

Why did this hap­pen?

Please make sure your browser sup­ports JavaScript and cook­ies and that you are not block­ing them from load­ing. For more in­for­ma­tion you can re­view our Terms of Service and Cookie Policy.

Need Help?

For in­quiries re­lated to this mes­sage please con­tact our sup­port team and pro­vide the ref­er­ence ID be­low.

Block ref­er­ence ID:3ca9eb24 – 99fb-11f1-bb7a-b575ea612eea

Get the most im­por­tant global mar­kets news at your fin­ger­tips with a Bloomberg.com sub­scrip­tion.

Who Are the Token Brokers?

vectoral.com

threat-re­search llm-se­cu­rity

August 10, 2026 Matt Lenhard 5 min read

Share

Where This Started

This is a fol­low-up ar­ti­cle to a piece I re­cently wrote about the to­ken re­lay mar­ket. Noticeably ab­sent from that piece was a men­tion of the rise of token bro­kers” — peo­ple who buy un­used cred­its from star­tups and then re­sell them.

I first heard about to­ken bro­kers while chat­ting with a good friend of mine who was re­ceiv­ing of­fers for Anthropic to­kens at steep dis­counts.

It was­n’t just him, though. As I started talk­ing to more founders about what I was build­ing, they said the same thing: they were get­ting a lot of in­bound email from peo­ple look­ing to buy or sell off-mar­ket in­fer­ence.

Startups swap­ping cred­its is noth­ing new, and I knew this was hap­pen­ing in sev­eral startup fo­rums and groups, but this was when I re­al­ized that the mar­ket was be­ing com­mer­cial­ized.

So I did what any nor­mal per­son would do. I got the bro­kers’ email ad­dresses and started email­ing them to learn more.

Before my own out­reach, it’s worth see­ing what founders are ac­tu­ally re­ceiv­ing. Both of these were for­warded to me by friends.

I started by sourc­ing a few email ad­dresses from friends. The first two emails I sent bounced, but the third was a hit. Here’s a screen­shot of that con­ver­sa­tion:

What’s in­ter­est­ing is the amount of sup­ply. The seller was of­fer­ing $100k in spend per day.

They aren’t hand­ing out the provider keys di­rectly; in­stead, they act as a proxy that prob­a­bly picks from a pool of keys and for­wards the re­quest.

The Listings

Credit Marketplaces

There are a few web­sites pro­mot­ing credit bro­ker­ing as well. One of them, AI Credits, bills it­self as a credit mar­ket­place. For an­other fla­vor of the pure-play credit re­seller mar­ket­places, take a look at AICreditMart.

These sites of­fer cred­its at most of the ma­jor cloud and in­fer­ence providers.

AI Credits’ on­board­ing process is pretty straight­for­ward, and you can even se­lect your pre­ferred de­liv­ery method as the seller.

I went ahead and listed my cred­its, which are still pend­ing ap­proval.

Bulk Discounts

Another site that I found through a friend was CheapCredits. This site po­si­tions it­self as a router that is able to achieve its dis­counts through bulk pric­ing.”

I no­ticed that this was a trend with a num­ber of sites that I be­lieve are act­ing as credit bro­kers. They pre­sent them­selves as be­ing able to of­fer dis­counts based on bulk pur­chases. Some other ex­am­ples in­clude Tokvana and Neokens.

Having spent time in the in­dus­try, I’d say that a 40% dis­count is very un­likely un­less you are one of the provider’s top cus­tomers. My hunch is that CheapCredits is ac­quir­ing the sup­ply in other ways.

CheapCredits even has a Data Processing Agreement for any­one look­ing to stay GDPR com­pli­ant.

The Message Boards

I checked where you’d ex­pect to find un­der­ground mar­ket­places.

Telegram had a few chan­nels, with one be­ing rel­a­tively ac­tive.

There are also spo­radic Reddit posts.

If you’ve been hang­ing out in any of the closed-off startup groups, I’m sure you’ve seen a num­ber of these posts as well.

So How Big Is This Market?

My rough es­ti­mate is that, across the sites, fo­rums, and re­sellers I looked at, there are prob­a­bly tens of mil­lions of these cred­its be­ing of­fered.

Unfortunately, when you try to of­fer nice things, abuse is­n’t far be­hind. Tokens have be­come a pseudo-cur­rency, and there is enough liq­uid­ity in the mar­ket to al­low for a lot of abuse. As we see the mar­ket turn and com­pa­nies be­come more aware of costs, crack­downs on this type of abuse prob­a­bly aren’t far be­hind.

Sources

Company and site names be­low are as they pre­sent them­selves pub­licly. Screenshots are from my own out­reach and from brows­ing the sites as a prospec­tive buyer and seller.

Previous piece: An Inside Look at the Relay Market Powering Token Resellers and Fraud

Credit mar­ket­places: AI Credits, AICreditMart

Bulk-discount routers: CheapCredits (cheapcredits.ai), Tokvana, Neokens

Direct out­reach: email ex­change with a bro­ker of­fer­ing $100k/day in spend

Share

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

simonwillison.net

16th August 2026

Friday’s big re­lease was Qwen 3.8 27B, an Apache 2 li­censed 27B pa­ra­me­ter vi­sion-ca­pa­ble LLM from Alibaba’s Qwen re­search lab. I’ve been look­ing for­ward to this one: 27B is an ex­cel­lent size for run­ning a model on a rea­son­ably specced lap­top, and its pre­de­ces­sor Qwen 3.6 27B was im­pres­sive.

Qwen’s self-re­ported bench­marks for this model are eye-open­ing. They show a boost from both Qwen 3.6 27B and the closed-weight Qwen 3.7-Plus, which was one of Qwen’s strongest mod­els of any size as re­cently as May this year. It will be in­ter­est­ing to hear what in­de­pen­dent bench­marks have to say about the model.

I’ve been run­ning the model on two dif­fer­ent ma­chines: my 128GB M5 Max MacBook Pro, and an NVIDIA DGX Spark. On both ma­chines I’m run­ning LM Studio and their 17GB Q4_K_M quan­tized build. I also tried us­ing llama-server di­rectly on the Spark.

Qwen’s doc­u­men­ta­tion de­scribes the model as de­fault­ing to xhigh for the rea­son­ing ef­fort, and the LM Studio GGUF I’ve been try­ing pre­serves that de­fault:

Qwen3.8 comes with of­fi­cial sup­port for rea­son­ing_­ef­fort, which can be used to ad­just rea­son­ing depth and con­trol cost:

xhigh (default): for com­plex tasks de­mand­ing thor­ough analy­sis

medium: bal­anc­ing ac­cu­racy and speed

low: ef­fi­cient rea­son­ing op­ti­miz­ing for speed and cost

Qwen3.8 comes with of­fi­cial sup­port for rea­son­ing_­ef­fort, which can be used to ad­just rea­son­ing depth and con­trol cost:

xhigh (default): for com­plex tasks de­mand­ing thor­ough analy­sis

medium: bal­anc­ing ac­cu­racy and speed

low: ef­fi­cient rea­son­ing op­ti­miz­ing for speed and cost

This is a hi­lar­i­ous de­fault. It’s ab­solutely not a good way to run the model, es­pe­cially on con­sumer hard­ware. I’ve been find­ing the re­sults ex­tremely en­ter­tain­ing.

I quickly ran into prob­lems with LM Studio’s de­fault con­text limit of 8,192 to­kens—Qwen was us­ing them all up think­ing about even the most mun­dane of prob­lems. I loaded the model with the full 262,144 max­i­mum con­text length and that prob­lem went away.

Here’s the pel­i­can rid­ing a bi­cy­cle SVG I got from my first at­tempt with that in­creased con­text length. It took 21 min­utes to gen­er­ate, us­ing 22,276 rea­son­ing to­kens to pro­duce 3,223 to­kens of out­put. You can read the rea­son­ing trace here.

This is by far the best pel­i­can SVG I’ve been able to gen­er­ate with a model that runs on a lo­cal ma­chine—and this Qwen is pretty small, just a 17GB file on disk. There’s a lot to like about this:

The bi­cy­cle frame is the right shape

It has legs on each side of the bike—that’s very rare

Good, clear pel­i­can pouch

The wings ex­tend to touch the han­dle­bars!

The mo­tion lines are be­hind, not in front

It has a taste­ful back­ground—nice sun, clouds, hill, flow­ers and grass.

Was that worth wait­ing 21 min­utes for? Absolutely not.

Here’s that same prompt run with rea­son­ing turned off—tran­script here. This one pro­duced 3,715 to­kens and took 137s—just over two min­utes.

And just for com­plete­ness, I used OpenRouter to run the same prompt through the much larger Qwen 3.8 2.4T-A95B (released last week) and got this snazzy an­i­mated SVG:

Your browser does not sup­port HTML5 video.

I said Qwen at xhigh has a ten­dency to over-think things, but how bad re­ally is it?

I tried a much sim­pler prompt, again with that de­fault ex­tra high set­ting:

draw an svg of a cir­cle

draw an svg of a cir­cle

Qwen’s rea­son­ing trace started like this:

The user is ask­ing for an SVG draw­ing of a cir­cle. Simple re­quest — but I want it to be a care­fully crafted piece. Let me make some­thing that goes be­yond just <circle>: a sin­gle self-con­tained SVG file with char­ac­ter — maybe a geo­met­ric circle study,” with sub­tle an­i­ma­tion, lay­ered rings, and a dis­tinc­tive palette. Keep the scope right: they asked for an SVG of a cir­cle. So the core is a cir­cle. But I can add craft: con­cen­tric guide cir­cles (like a com­pass/​geom­e­try draw­ing), tick marks, a soft gra­di­ent fill on the main cir­cle, re­strained am­bi­ent mo­tion (a slowly ro­tat­ing dashed ring, puls­ing glow). Respect prefers-re­duced-mo­tion? That’s more for CSS; SVG SMIL or CSS in­side SVG will do. Palette op­tions: deep teal ink on warm pa­per? Or bold ver­mil­ion cir­cle on off-white with navy con­struc­tion lines — Bauhaus/compass-drawing vibe. Let me go with a geometric study” look: cool slate back­ground, or bright pa­per white? Paper white is fine if it’s not the cream-and-ter­ra­cotta combo. […]

The user is ask­ing for an SVG draw­ing of a cir­cle. Simple re­quest — but I want it to be a care­fully crafted piece. Let me make some­thing that goes be­yond just <circle>: a sin­gle self-con­tained SVG file with char­ac­ter — maybe a geo­met­ric circle study,” with sub­tle an­i­ma­tion, lay­ered rings, and a dis­tinc­tive palette.

Keep the scope right: they asked for an SVG of a cir­cle. So the core is a cir­cle. But I can add craft: con­cen­tric guide cir­cles (like a com­pass/​geom­e­try draw­ing), tick marks, a soft gra­di­ent fill on the main cir­cle, re­strained am­bi­ent mo­tion (a slowly ro­tat­ing dashed ring, puls­ing glow). Respect prefers-re­duced-mo­tion? That’s more for CSS; SVG SMIL or CSS in­side SVG will do.

Palette op­tions: deep teal ink on warm pa­per? Or bold ver­mil­ion cir­cle on off-white with navy con­struc­tion lines — Bauhaus/compass-drawing vibe. Let me go with a geometric study” look: cool slate back­ground, or bright pa­per white? Paper white is fine if it’s not the cream-and-ter­ra­cotta combo. […]

Several min­utes later it pro­duced this ab­solutely beau­ti­ful an­i­mated cir­cle, which was en­tirely not what I had asked for!

Your browser does not sup­port HTML5 video.

My strong rec­om­men­da­tion: ig­nore that de­fault. Run Qwen 3.8 27B on low or even no rea­son­ing lev­els at first. It’s a great model, but wow that de­fault set­ting is a bad place to start.

It’s very good at bound­ing boxes

A fun way to test a vi­sion model is to see how well it can re­turn bound­ing boxes around items in a pho­to­graph. I’ve seen pre­vi­ous Qwen mod­els deal well with this, so I de­cided to put it to the test draw­ing bound­ing boxes around some pel­i­cans.

I’ve seen ask­ing for 0 – 1000 scale pro­duce good re­sults in the past. I tried this:

llm -a https://​sta­tic.inat­u­ral­ist.org/​pho­tos/​714731804/​large.jpg \ -m lm­stu­dio/​qwen/​qwen3.8 – 27b \ Return JSON bound­ing boxes for the pel­i­cans in this photo, 0 – 1000 scale for each di­men­sion’

Here’s the rea­son­ing trace, which pro­duced this:

[ {“bbox_2d”: [195, 290, 370, 780], label”: pelicans”}, {“bbox_2d”: [445, 320, 675, 850], label”: pelicans”} ]

This is such a good match. Here are those boxes ren­dered on top of the photo:

Building a tool to la­bel bound­ing boxes

That vi­su­al­iza­tion of the bound­ing boxes was taken us­ing a new cus­tom tool that I had Qwen 3.8 27B build for me, run­ning of­fline on my lap­top.

I for­got to dial down the think­ing ef­fort so it was mas­sively over-en­gi­neered, but it did man­age to pro­duce this full in­ter­face from this sin­gle prompt:

[ {“bbox_2d”: [195, 290, 370, 780], label”: pelicans”}, {“bbox_2d”: [445, 320, 675, 850], label”: pelicans”} ]

Build an HTML page which has an in­put box for ac­cept­ing the URL to an im­age and a textarea for ac­cept­ing the above style of JSON. It ap­pends the im­age to the page, mea­sures its width and height, then treats the co­ords in the bbox_2d as scaled from 0 – 1000 and scales them against the ac­tual width and height, then it ren­ders la­belled boxes over the im­age.

[ {“bbox_2d”: [195, 290, 370, 780], label”: pelicans”}, {“bbox_2d”: [445, 320, 675, 850], label”: pelicans”} ]

Build an HTML page which has an in­put box for ac­cept­ing the URL to an im­age and a textarea for ac­cept­ing the above style of JSON.

It ap­pends the im­age to the page, mea­sures its width and height, then treats the co­ords in the bbox_2d as scaled from 0 – 1000 and scales them against the ac­tual width and height, then it ren­ders la­belled boxes over the im­age.

This screen­shot shows one of the fea­tures I did not ask for—a demo scene, for if you don’t have a pho­to­graph to test the tool with:

Here’s the rel­e­vant seg­ment of the think­ing trace, where it de­cided to draw its own pel­i­cans purely be­cause I had used the la­bel pelicans” in the ex­am­ple JSON I gave it in the prompt:

Also a load sam­ple” that uses a known im­age? Can’t de­pend on ex­ter­nal im­ages, but… the im­age URL in­put is user-pro­vided; I could add a try with sam­ple” but­ton […] Hmm, I can draw a sim­ple scene on can­vas, ex­port it as a data URL, and load it into the im­age — that’s self-con­tained and demo-able! […] But the user’s co­ords are for an ac­tual pel­i­can im­age; a gen­er­ated place­holder can still demo the scal­ing. Generate a 1000x1000 place­holder: gra­di­ent wa­ter + two blob-like pelican” sil­hou­ettes placed at the given bboxes (using the same scale — cute: sil­hou­ettes at the ex­act 0 – 1000 po­si­tions, show­ing the boxes align). This makes for a fun, self-con­tained demo. Keep it sim­ple: sky gra­di­ent, sun, wa­ter, two pel­i­can-ish shapes (ellipse body, cir­cle head, beak). Place at bbox cen­ters.

Also a load sam­ple” that uses a known im­age? Can’t de­pend on ex­ter­nal im­ages, but… the im­age URL in­put is user-pro­vided; I could add a try with sam­ple” but­ton […] Hmm, I can draw a sim­ple scene on can­vas, ex­port it as a data URL, and load it into the im­age — that’s self-con­tained and demo-able! […] But the user’s co­ords are for an ac­tual pel­i­can im­age; a gen­er­ated place­holder can still demo the scal­ing. Generate a 1000x1000 place­holder: gra­di­ent wa­ter + two blob-like pelican” sil­hou­ettes placed at the given bboxes (using the same scale — cute: sil­hou­ettes at the ex­act 0 – 1000 po­si­tions, show­ing the boxes align). This makes for a fun, self-con­tained demo. Keep it sim­ple: sky gra­di­ent, sun, wa­ter, two pel­i­can-ish shapes (ellipse body, cir­cle head, beak). Place at bbox cen­ters.

(I’m slightly ner­vous that mod­els around the world might have a bias to­wards draw­ing pel­i­cans at any chance they can get, brought on by nearly two years of ex­po­sure to my own stu­pid bench­mark.)

Is all that over-think­ing nec­es­sary? Maybe it is, at least a bit. I tried with rea­son­ing turned off and got this ver­sion, (transcript here), which nearly works but shows the boxes in the wrong place:

So with­out rea­son­ing it did­n’t quite one-shot a work­ing tool. I’m sure it could get there with some fol­low-up prompts, but this is a good ex­am­ple of how rea­son­ing can make a dif­fer­ence.

Yes, it can drive cod­ing agents

One of the biggest ques­tions around lo­cal mod­els is whether or not they have enough horse­power to suc­cess­fully run a cod­ing agent loop. Coding agents re­quire long con­text, strong code gen­er­a­tion sup­port and re­li­able tool-call­ing. On pa­per Qwen 3.8 27B has all three of these, so is it up to the task?

My ini­tial ex­per­i­ments with Pi have been very promis­ing. I chose Pi be­cause it has a shorter sys­tem prompt than most other op­tions, mak­ing it a bet­ter fit for try­ing out smaller mod­els.

I con­fig­ured Pi to use Qwen 3.8 27B run­ning in LM Studio on the Spark (shared via tailscale serve) by adding this to ~/.pi/agent/models.json:

{ providers”: { spark”: { baseUrl”: https://​spark-18b3.tail68a31.ts.net/​v1, api”: openai-responses”, apiKey”: dummy”, models”: [ { id”: qwen3.8 – 27b”, reasoning”: true } ] } } }

Then ran pi –provider spark –model qwen3.8 – 27b in my ~/dev/datasette folder and prompted:

how does auth work?

how does auth work?

After a se­quence of rea­son­ing and tool calls that ac­cessed a bunch of dif­fer­ent files it pro­duced this re­ply, which is very solid.

Just one prob­lem: I wanted to share that tran­script. So I pointed Pi and Qwen 3.8 27B at the JSONL tran­script file in ~/.pi/agent/sessions/–Users-simon-Dropbox-dev-datasette– and prompted:

Write Python code to con­vert this jsonl to mark­down

Write Python code to con­vert this jsonl to mark­down

And it built and tested this pi_j­son­l_­to_md.py, which did ex­actly what I needed. Here’s that ses­sion tran­script, pub­lished us­ing the tool that it cre­ated.

The quest for speed

So far this is all look­ing very promis­ing. We have a 17GB model that runs on high-end con­sumer hard­ware and can write code, drive tools, an­no­tate im­ages and gen­er­ally do every­thing that I need from an LLM for get­ting real work done.

There’s one very sig­nif­i­cant catch: it feels slow—es­pe­cially when it starts over-think­ing, but even with­out that it’s not par­tic­u­larly sprightly.

I’ve been get­ting around 15 – 30 to­kens a sec­ond from LM Studio. That’s not ter­ri­ble, but it’s slow enough that it’s go­ing to be hard to win me away from hosted API mod­els, which can re­turn re­sults a whole lot faster. Artificial Analysis track to­ken speed and show OpenAI 5.6 Sol at 74 to­kens/​sec­ond and 5.6 Luna at an im­pres­sive 184/second.

The good news is that the com­mu­nity have been ex­plor­ing ways to speed things up since the model was first re­leased two days ago.

One of the most promis­ing op­ti­miza­tions is baked into the model it­self. Qwen sup­ports Multi-Token Prediction, an ar­chi­tec­ture trick where a cheaper mech­a­nism guesses sev­eral to­kens ahead and the main model can then quickly ver­ify if the guesses were cor­rect. This can have quite a dra­matic ef­fect on in­fer­ence per­for­mance.

Based on this tweet from llama.cpp cre­ator Georgi Gerganov I tried run­ning the model with MTP like this on the Spark:

llama serve \ -hf ggml-org/​Qwen3.8 – 27B-GGUF:Q4_K_M \ -hfd ggml-org/​Qwen3.8 – 27B-GGUF:Q4_0 \ –spec-default \ –spec-type draft-mtp \ –reasoning-preserve

And sure enough, this gave me a sig­nif­i­cant boost. I had GPT-5.6 in Codex run a com­par­a­tive bench­mark on the Spark and the –spec-type draft-mtp server out­per­formed the LM Studio de­fault GGUF by around 72%.

I ex­pect we’ll see a whole lot more in­no­va­tion around serv­ing this model faster over the next few weeks. The MLX com­mu­nity likely have some tricks brew­ing as well.

Some ob­ser­va­tions

The fact that a 17GB file can do all of this stuff on my home ma­chines is a mir­a­cle. Once again, I’m de­lighted and amazed at how much progress lo­cal mod­els have made this year. A year ago this would have been com­pet­i­tive with the best and most ex­pen­sive of the pro­pri­etary mod­els—to­day it can run on a ca­pa­ble lap­top.

The only thing hold­ing this back from be­ing a daily dri­ver is per­for­mance. It feels pretty slow on both the M5 Mac and the DGX Spark. That’s the catch with these dense (non-Mixture-of-Experts) mod­els—they re­quire a whole lot of mem­ory band­width to per­form well, and nei­ther of the ma­chines I have ac­cess to are top per­form­ers in that re­gard.

The most im­por­tant thing about Qwen 3.8 27B is what it demon­strates. We can have an open weights gen­eral pur­pose model with a long con­text, ef­fec­tive tool call­ing, strong vi­sion abil­ity, and com­pe­tent code gen­er­a­tion, and we can fit the whole thing in just a 17GB file.

The mod­els at this size con­tinue to get bet­ter at an im­pres­sive rate. We don’t need to spend half a mil­lion dol­lars on dat­a­cen­ter-class hard­ware just to run a com­pe­tent model.

Language Models Under Pedagogically-Controlled Knowledge Exposure

littlelearner-ll.github.io

Talk to LittleLearner

The hosted 5B model, live in your browser. Open in a new tab ↗ if the chat does­n’t load be­low.

A con­trolled sand­box for study­ing how mod­els ac­quire knowl­edge

Modern LMs are trained on every­thing at once, so it is hard to tell whether a new skill was learned or merely elicited. We con­strain the train­ing dis­tri­b­u­tion it­self: an 88B-token cor­pus fil­tered to the U.S. el­e­men­tary-school cur­ricu­lum, with mod­els trained from scratch on it and matched un­fil­tered con­trols.

Dataset

LittleCurriculum

An 88B-token cor­pus dis­tilled from FineWeb-Edu through a five-stage fil­ter­ing pipeline aligned with Common Core stan­dards (K–5). Concepts, facts, and vo­cab­u­lary taught above Grade 5 are ex­plic­itly ex­cluded.

Models

LittleLearner

Three scales (0.6B / 1.3B / 5B) trained from scratch on LittleCurriculum: chat­table mod­els with an in­ter­pretable knowl­edge bound­ary. Each ships with a matched Unfiltered con­trol for clean com­par­i­son.

Findings

Elicitation, not ac­qui­si­tion

In our ex­per­i­ments, scal­ing, SFT+GRPO post-train­ing, and in-con­text learn­ing am­plify what the cur­ricu­lum taught, but none mean­ing­fully im­proves out-of-scope per­for­mance, in­di­cat­ing that the pre­train­ing fil­ter sets the ef­fec­tive ca­pa­bil­ity ceil­ing.

Model check­points

LittleLearner at three scales (0.6B / 1.3B / 5B), each with a matched Unfiltered con­trol shar­ing its ar­chi­tec­ture, to­kens, and recipe.

Base: the pre­trained model. GRPO: math spe­cial­ists post-trained on MathCAMPS; re­sponses may ex­hibit a ten­dency to­ward math-ori­ented out­put. Chatty: vari­ants tuned for gen­eral chat be­hav­ior.

Capability stays in­side the cur­ricu­lum

Can stan­dard in­ter­ven­tions push a model past what its pre­train­ing data taught it? With the bound­ary un­der ex­per­i­men­tal con­trol, we can ask cleanly. In our ex­per­i­ments, each in­ter­ven­tion am­pli­fies in-scope abil­ity; none of them mean­ing­fully im­proves out-of-scope per­for­mance.

Scaling

Scaling model size im­proves per­for­mance within the mod­el’s con­trolled knowl­edge ex­po­sure and ex­tends mod­estly to prob­lems along the same learn­ing tra­jec­tory, but yields lit­tle im­prove­ment on prob­lems re­quir­ing more ad­vanced ca­pa­bil­i­ties out­side the ex­po­sure.

MathCAMPS ac­cu­racy by grade, across model size

Post-training

Post-training through GRPO sig­nif­i­cantly boosts in-scope K–5 ca­pa­bil­i­ties, but fails to re­cover out-of-scope be­yond-K–5 ca­pa­bil­i­ties, even when train­ing with out-of-scope data.

Post-training am­pli­fies K–5, not the be­yond-K–5 gap

In-context learn­ing

In-context learn­ing with the prompts we test does not un­lock new rea­son­ing ca­pa­bil­i­ties in be­yond-K–5 for our trained 5B LittleLearner.

Accuracy by prompt­ing con­di­tion

What will you teach it?

Because LittleLearner’s train­ing ex­po­sure is ex­plic­itly spec­i­fied, be­hav­ioral and rep­re­sen­ta­tional changes can be re­lated di­rectly to the con­cepts you in­tro­duce. Three di­rec­tions we’re ex­cited about:

01

RL & dis­cov­ery

Can RL cre­ate ca­pa­bil­ity?

The prior is re­stricted to K–5, so ca­pa­bil­i­ties that emerge un­der RL can be at­trib­uted to the RL process it­self. A tractable proxy for re­ward-dri­ven dis­cov­ery.

02

Continual learn­ing

Watch a con­cept be­ing learned

Introduce neg­a­tive num­bers and mea­sure sam­ple ef­fi­ciency, re­ten­tion, and in­ter­fer­ence. Or probe be­hav­ior near the bound­ary: does it an­swer, ab­stain, or hal­lu­ci­nate?

03

Educational sci­ence

Machine vs. child learn­ers

Specified ex­po­sure en­ables con­trolled hu­man-model com­par­i­son. Do mod­els and chil­dren need sim­i­lar ex­po­sure to learn frac­tions, or make sim­i­lar er­rors on word prob­lems?

Your turn

Bring your own ques­tion

A known bound­ary turns your idea into a clean ex­per­i­ment!

If you find this work use­ful

Please cite our pa­per:

@misc{littlelearner2026, ti­tle={Lit­tle­Learner: Language Models Under Pedagogically-Controlled Knowledge Exposure}, au­thor={Fan­fei Li and Jana Zeller and Manuel Prada-Corral and Thaddäus Wiedemer and Prasanna Mayilvahanan and Ryan Cotterell and Wieland Brendel}, year={2026}, eprint={2608.13545}, archivePre­fix={arXiv}, pri­ma­ryClass={cs.CL}, url={https://​arxiv.org/​abs/​2608.13545} }

The weekend is 100 years old – but have ‘skiveday Fridays’ and hybrid working ruined it for everyone?

www.theguardian.com

For 11 years, from 1929 to 1940, the Soviet Union did not have week­ends. Instead, to in­crease pro­duc­tiv­ity cit­i­zens were al­lo­cated a day off in every seven at ran­dom. With 80% of the pop­u­la­tion at work on any given day, fac­to­ries never had to power down.

A let­ter from a dis­grun­tled worker to Pravda news­pa­per, pub­lished soon af­ter the im­ple­men­ta­tion of this new work­ing cal­en­dar, out­lined the prob­lems: What is there for us to do at home if our wives are in the fac­tory, our chil­dren at school, and no­body can visit us?” the let­ter-writer asked. It is no hol­i­day if you have to have it alone.” Parents found them­selves at home while their chil­dren were at school, or with kids un­su­per­vised while they were on shift. With no shared day off, ex­tended fam­ily gath­er­ings be­came im­pos­si­ble. The work­force be­came de­mor­alised and list­less, the pro­jected spike in pro­duc­tiv­ity never ma­te­ri­alised, and the pol­icy was first mod­i­fied, then aban­doned al­to­gether.

Eleven years, though. Is that not ab­solutely wild? A world with­out week­ends feels im­pos­si­ble. A world with­out Saturday Night Fever, with­out Manic Monday. We may no longer go to church on Sunday, but we still wor­ship the week­end. The week­end looms large be­cause it rep­re­sents the tri­umph of col­lec­tive time over mar­ket time,” says Brad Beaven, pro­fes­sor of so­cial and cul­tural his­tory at the University of Portsmouth. It is not just about rest, but about re­claim­ing au­ton­omy from the in­dus­trial clock.”

Saturday and Sunday, sa­cred to the dig­nity and hu­man­ity of work­ing peo­ple, are laden with mythol­ogy and cer­e­mony. From Cilla Black to Gary Lineker, the main char­ac­ters of our week­ends be­come gi­ants of the cul­ture. We ro­man­ti­cise the week­end, even the pro­saic bits — the dis­tant roar of a lawn­mower, the rat­tle of clas­si­fied foot­ball re­sults on the ra­dio. Everyone knows what a week­end means. The snap of a lap­top cover at 4.58pm on a Friday. The clink of the first pint glass and the smell of chips on the way home. The slow-mo­tion pace of a week­end pave­ment, the gear change from hus­tle to me­an­der. The rit­u­als have changed with the times, of course, and peo­ple-watch­ing at brunch on Saturday is now as much of a tra­di­tion as cook­ing a roast at home on Sunday. A week­end changes shape as you move through life stages, but in each it­er­a­tion it re­mains a shared ex­pe­ri­ence be­tween you and your peers.

The week­end that I had when I was 20 was very dif­fer­ent from the week­end that I have now, at 40 and with kids,” laughs Pedro Gomes, pro­fes­sor of eco­nom­ics at Birkbeck, University of London and au­thor of the book Friday is the New Saturday. When you are young, you are bond­ing with your friends, and then when you get older, you might be with your fam­ily. Eventually, you might be with grand­chil­dren. We move through dif­fer­ent man­i­fes­ta­tions of the week­end, in our life­times.”

But it can be hard to pin down, these days, where a week­end be­gins and ends. Laundry gets done on a work-from-home Friday, but emails are an­swered on Sunday. A shop­ping splurge is as likely to be a cheer-up treat on your phone af­ter a tough Wednesday as a Saturday out­ing. The tra­di­tional Saturday 3pm foot­ball kick-off has been stretched across the tele­vi­sion sched­ules all the way to Monday evening. The four-day week — first pre­dicted by Richard Nixon, of all peo­ple, in 1956 — has be­come a re­al­ity in the Netherlands, with peo­ple work­ing an av­er­age of 32.1 hours. Friday is forg­ing ahead with a quiet se­ces­sion from the work­ing week with­out any­one sign­ing off the pa­per­work. It raises the ques­tion: what even is the week­end” in 2026?

Remarkably, the con­cept of the week­end as we know it is only 100 years old. In 1926, Henry Ford changed the shape of the week, an­nounc­ing that the work­ers at his fac­to­ries would now do five eight-hour days in­stead of six, with no cut in pay. Ford did not in­vent the week­end — the idea had been bub­bling un­der for a cen­tury, in cam­paigns by trade unions, re­li­gious groups and pro­gres­sive em­ploy­ers — but by putting his con­sid­er­able in­dus­trial weight be­hind it, the two-day chunk of free­dom was born.

It is high time to rid our­selves of the no­tion that leisure for work­men is ei­ther lost time or a class priv­i­lege,” Ford wrote in his com­pa­ny’s Ford News in October 1926. Ford’s in­no­va­tion was in part a re­sponse to his ear­lier in­ven­tion: the as­sem­bly line, which had in­creased pro­duc­tiv­ity but ex­hausted work­ers, with ab­sen­teeism up to 10% in fac­to­ries; in his ar­ti­cle he did not dis­guise that there was self-in­ter­est in­volved. People who have more leisure re­quire more trans­porta­tion in ve­hi­cles,” he con­tin­ued. The new fash­ion for day trips was an ef­fec­tive mar­ket­ing de­vice to sell cars. Meanwhile, his fac­to­ries main­tained a steady level of pro­duc­tiv­ity de­spite the re­duced hours. In 1938, faced with ris­ing un­em­ploy­ment lev­els in the Great Depression, the five-day week was of­fi­cially adopted across the US.

On this side of the Atlantic, Boots the Chemist was the pi­o­neer. In 1933, the com­pany opened a new fac­tory in Nottingham, which proved so ef­fi­cient that there was soon a sur­plus of stock. Reluctant to lay his staff off with un­em­ploy­ment run­ning at 25%, John Boot ended the Saturday morn­ing shift, re­duc­ing hours for the 5,000-strong work­force with­out cut­ting wages. It was com­mer­cially suc­cess­ful, re­sult­ing in lower rates of ab­sen­teeism, and was adopted as Boots pol­icy in 1934. A gov­ern­ment in­quiry led by Richard Redmayne — great-grand­fa­ther, fun fact, of ac­tor Eddie — pub­lished a review of the ex­per­i­men­tal work­ing of the five days week”, which noted an im­prove­ment in sta­mina and an­i­ma­tion of the em­ploy­ees aris­ing from the phys­i­o­log­i­cal and psy­cho­log­i­cal ef­fects of a long week­end’s rest and re­lax­ation”. Slowly, the week­end gath­ered mo­men­tum. In the UK, the early 20th-century ver­sion of the week­end was gen­er­ally recog­nised as a half-day Saturday and Sunday off,” says Beaven. The full two-day week­end only re­ally be­came widely adopted af­ter the sec­ond world war.”

Before in­dus­tri­al­i­sa­tion, there was lit­tle con­cept of con­sec­u­tive days of leisure, be­cause an­i­mals and crops could not be so long ne­glected. Work was dic­tated by the weather and the sea­sons, the clock mat­ter­ing less than the sun. Factories, with their whis­tles and watches, changed the way time worked. A mech­a­nised drum­beat of shifts and pay­days drowned out the old rhythms of sea­sons and saints days.

As work be­came more rigidly or­gan­ised, the pos­si­bil­ity emerged that leisure could be, too. Time off was­n’t merely a con­ces­sion to work­ers, but also an en­gine of con­sumer cap­i­tal­ism. People with week­ends would go to sports sta­di­ums, buy pic­nic bas­kets and new clothes, need cin­ema tick­ets. By the late 20th cen­tury, with the ar­rival of cheap flights, this had evolved into the mini­break: a minia­ture hol­i­day, de­signed to fit into a week­end. Leisure time, once the op­po­site of the econ­omy, be­came part of it.

For gen­er­a­tions, the British week­end re­volved around one im­mov­able ap­point­ment: the 3pm Saturday kick-off. The week­end is about rest, but it is also about pas­sion,” says Gomes. Most of us are not lucky enough to be pas­sion­ate about our jobs, but at the week­end we can fol­low our pas­sions.” The link be­tween foot­ball and the best day of the week is in­trin­sic to the na­tional love af­fair with the sport. It is prob­a­bly partly be­cause foot­ball sym­bol­ised the best bit of the week­end that we ended up so ob­sessed with it.

Morals — and the ab­sence of them — have al­ways been a theme of the week­end. When Sunday, a time of wor­ship, was the only day off, skilled work­ers de­vel­oped a habit of ex­tend­ing their free time into Saint Monday”, by not turn­ing up for work af­ter a par­tic­u­larly en­thu­si­as­tic day of drink­ing. Saturday af­ter­noons off — and then the whole day — were granted by em­ploy­ers partly in the hope of bring­ing the hang­overs for­ward by a day. Victorian re­form­ers, ob­sessed with drunk­en­ness, hoped that free Saturdays would en­cour­age re­spectable recre­ation: or­gan­ised sport, gar­den­ing, fam­ily out­ings. Ford, a ve­he­ment sup­porter of Prohibition, be­lieved that the il­le­gal­ity of al­co­hol made his move to­wards a two-day week­end safe. (“A day off is no longer a day drunk,” he said.)

The re­al­ity has never been quite so clean-cut. Weekends are naughty and nice, both bad be­hav­iour and Sunday best. These are the days for shop­ping splurges and drink­ing sprees and hook-ups, but also for penance, whether by parkrun, DIY or ac­tual prayer. Though per­haps less so the prayer bit: around one in three Britons at­tended church reg­u­larly in 1900, ac­cord­ing to the National Centre for Social Research; now, this is num­ber is around one in 20. Two mo­ments stand out in the story of how Sunday lost its spe­cial place as a day, if not of wor­ship, then of higher pur­pose. In 1994, the Sunday Trading Act al­lowed large shops in England and Wales to open on Sundays. Then, a quar­ter of a cen­tury later, pan­demic lock­downs broke the now-frag­ile bonds be­tween churches and their com­mu­ni­ties, and left many older parish­ioners with a wari­ness of gath­er­ing in ill-ven­ti­lated churches that might be bad for the health, even if good for the soul.

The late, great Maggie Smith had, as she so of­ten did, the best line. What is a week­end?” she asked, in Downton Abbey, with the en­ti­tled be­wil­der­ment of a dowa­ger count­ess for whom in­come is spoon­fed in sil­ver from birth, not doled out in a brown en­ve­lope on a Friday night. The week­end is time carved out of, and in ten­sion with, some­one else’s own­er­ship of your time. Not to men­tion that for aris­to­crats, who had ser­vants for every­thing from lay­ing fires to but­ton­ing their dresses, every day was a day of leisure. The week­end feels like a cor­ner­stone of civil­i­sa­tion, of democ­racy, be­cause it mat­ters most to those who spend the ma­jor­ity of the week fol­low­ing or­ders in­stead of giv­ing them.

af­ter newslet­ter pro­mo­tion

These days, the up­stairs-down­stairs di­vi­sion is be­tween the hy­brid work­ers and those whose work can never be done re­motely. An age-old di­vi­sion be­tween shift work and the rel­a­tive flex­i­bil­ity of white-col­lar jobs has deep­ened. Thinkers like Liselotte Lyngsø, found­ing part­ner of the Copenhagen-based con­sul­tancy Future Navigator, have ar­gued that the work­force is split­ting into time own­ers” — knowl­edge work­ers who in­creas­ingly choose where and when they work — and time slaves”, whose jobs, like health­care, man­ual shift work or gig econ­omy roles like Amazon or Deliveroo dri­ving, re­mains stub­bornly tied to the clock.

For the time own­ers”, Friday is rapidly turn­ing into the mod­ern Saint Monday. It is an in­creas­ingly open se­cret that the last day of the work­ing week has be­come, for re­mote work­ers, an un­of­fi­cial half shift: cal­en­dar tech­ni­cally open, but both brain and lap­top mostly on standby. When Gomes ran a six-month trial of four-day-week work­ing in Portugal in 2023, with more than 41 pub­lic and pri­vate firms in­volved, organisations that can’t re­duce hours — a nurs­ery, for ex­am­ple — worked in shifts with a dif­fer­ent day off within each week. Other firms made the de­ci­sion to cut Friday out. Whenever their week­day off was, we found that many em­ploy­ees ap­proached that day a lit­tle dif­fer­ently, us­ing it to get life ad­min’ done so that they could have their Saturdays and Sundays free. The week­end is­n’t just about the num­ber of hours, it is also a co­or­di­na­tion de­vice for com­mu­ni­ties to con­nect.”

As with Saint Monday, Skiveday Friday” has proved in the UK to be stub­bornly, if silently, ad­hered to. In 2024, Transport for London ran a three-month trial scrap­ping peak fares on Fridays ex­plic­itly in or­der to lure com­muters back into cen­tral London. It failed, Fridays re­main­ing quiet, de­spite the dis­count. For the lap­topped-classes, there has been a shift in what Friday is for. Several ma­jor rail com­pa­nies, in­clud­ing LNER and Avanti, have now made Fridays an en­tirely off-peak day like Saturday and Sunday, ac­knowl­edg­ing that the week­end now be­gins ear­lier for many.

Though it should be noted that, for those same work­ers, Saturday and Sunday them­selves are un­der con­stant siege from the mis­sion creep of work emails and con­tactabil­ity out­side work­ing hours in the age of the smart­phone. These are two sides of the same coin,” says Gomes. Work in­ten­si­fies, and the week­end ex­pands in or­der to ab­sorb the pres­sure. The speed of com­mu­ni­ca­tion means that we now live with con­stant in­ter­rup­tions, and it is hard to find space ei­ther for deep work or for real rest. And yet we con­tinue to struc­ture the work week in much the same way as we did 100 years ago. It is no sur­prise that this is­n’t work­ing.”

Hence why a three-day week­end is be­ing sug­gested as a way to bal­ance the books. But is it all woke non­sense? A pie in the sky idea, dreamed up by a lazy work­force which no longer knows the mean­ing of hard work? The Green party sup­ports a move to­wards” a four-day week, but other politi­cians are wary. In 2025, the Liberal Democrat-led South Cambridgeshire coun­cil, which had been op­er­at­ing a four-day week for two years, was at­tacked by the then lo­cal gov­ern­ment sec­re­tary, Steve Reed, who said that lo­cal gov­ern­ments should not be pay­ing full-time wages for part-time work”. In April this year, James Cleverly an­nounced that a fu­ture Conservative gov­ern­ment would look to ban four-day weeks for coun­cil staff, de­nounc­ing the push from the left of cen­tre in British pol­i­tics” to­wards a four-day week as completely wrong”.

The 4 Day Week Foundation, which is lead­ing the drive for change in Britain, says that 56 out of the 61 com­pa­nies who signed up for their four-day week pi­lot de­cided to stick with it af­ter the scheme ended. They cite an av­er­age of 35% in­creased rev­enue, and 57% de­cline in staff leav­ing rates, dur­ing the trial. There are now 260 com­pa­nies in the UK of­fi­cially signed up to the four-day week.

The di­vide — be­tween time shar­ers and time slaves, those work­ing five-day weeks, four-day weeks or more ad-hoc, less week­end-friendly shift work — is prob­lem­atic be­cause the week­end is de­signed to be shared. It’s where British cul­ture learned to gather, a col­lec­tive ex­pe­ri­ence. Every gen­er­a­tion has its own shared rit­u­als. Boomers love a Saturday night movie, gen­er­a­tion X are ob­sessed with Sunday lunch, mil­len­ni­als love to flock to a farm­ers’ mar­ket, while gen Z, for some rea­son, get their kicks stand­ing in line for baked goods that have gone vi­ral on TikTok. Even the Sunday scaries are made man­age­able by the knowl­edge that every­one out there is feel­ing the same. As a gulf widens — be­tween the time own­ers and time slaves, be­tween in­her­ited wealth and the in­creas­ingly weedy salary pipeline — the week­end starts to feel less like an ex­pe­ri­ence, and more like a nos­tal­gic mem­ory.

It is funny to think that, at first, tech­nol­ogy ex­panded our hori­zons. Railways gave or­di­nary peo­ple ac­cess to the coast; mass pro­duc­tion made bi­cy­cles and cars af­ford­able. It is only re­cently that tech­nol­ogy has mu­tated into the en­ergy vam­pire it is to­day, suck­ing the life from the week­end.

Only in the last decade has on­line au­to­bi­og­ra­phy be­come every­body’s un­paid side hus­tle, so that your friends have al­ready seen your hol­i­day pho­tos be­fore you meet them for brunch. It is a very mod­ern phe­nom­e­non that our friend­ships have mi­grated on­line, this week’s boy drama or fam­ily quar­rel al­ready de­bated at length in text bub­bles and voice notes with­out the need to meet in the pub. Add to this the post-pan­demic nor­mal­i­sa­tion of the “soft com­mit­ment”, in which mak­ing plans and can­celling them have be­come two parts of the same so­cial rit­ual and I’ll see how I feel” has be­come an ac­cept­able RSVP, and we have be­come less good at a core el­e­ment of the week­end, which is ac­tu­ally see­ing other peo­ple.

Perhaps we should­n’t fret so much. Ever since we in­vented the week­end, we have been anx­iously tak­ing its pulse. The Victorians wor­ried that work­ers would waste their pre­cious leisure days drink­ing. Mid-century Britain feared tele­vi­sion was keep­ing fam­i­lies in­doors. Today’s con­cern is that we spend Saturday scrolling Instagram and Sunday an­swer­ing emails. The rit­u­als change; the sus­pi­cion that the week­end is be­ing ru­ined re­mains. Yet the week­end sur­vives, per­haps be­cause it has never de­manded per­fec­tion. The lie-in and the late night are equally valid forms of re­sis­tance. Its holi­est rites — the greasy fry-up, the foot­ball ter­race, Saturday-night shiny-floor tele­vi­sion — are glo­ri­ously un­re­fined. The week­end is where a good time still trumps good taste.

The Soviets dis­cov­ered that a day off is not much use un­less other peo­ple are off, too; Ford un­der­stood that work­ers needed time in which to be­come them­selves again, even if he hoped they would spend it buy­ing cars. At 100, the two-day week­end is fray­ing at the edges, leak­ing into Friday and nib­bled away by Sunday night emails. The week­end will keep chang­ing, be­cause work keeps chang­ing. But it will still al­ways be the best thing work ever in­vented.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.