10 interesting stories served every morning and every evening.

LLMs reward expertise

www.seangoedecke.com

In the 2010s, if you had tech­ni­cal gaps (say, you could­n’t write CSS), you had to ei­ther rely on a skilled col­league or just hope that the an­swer to your ex­act prob­lem was out there on the in­ter­net. Today, every­one can write sort-of-okay CSS by del­e­gat­ing the task to an LLM. LLMs make every­body into a gen­er­al­ist.

Because of this, lots of peo­ple don’t think there’s any skill in­volved in work­ing with LLMs. If you want the prod­uct that LLMs can de­liver — PhD-level math­e­mat­ics, pretty good but some­times taste­less com­puter code, or awk­ward LinkedIn-style writ­ing — you can sim­ply ask for it. Since every­one is talk­ing to the same mod­els, skilled prompters” are get­ting the same re­sults as peo­ple touch­ing LLMs for the first time.

This is wrong. The most im­por­tant skill in prompt­ing is ex­per­tise in the do­main you’re prompt­ing for.

A good il­lus­tra­tion of this is Terence Tao’s con­ver­sa­tion with ChatGPT about the re­cently-dis­cov­ered coun­terex­am­ple to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I could­n’t get to where Tao gets, even with un­lim­ited to­kens to burn.

There’s a lot to learn about good prompt­ing from Tao’s con­ver­sa­tion. Here are a few ob­ser­va­tions:

Tao’s mes­sages are very short and to-the-point. He does­n’t re­spond point-by-point to the model, just to the gist

The model out­puts are much more con­cise than when I try and talk to GPT-5.6 Sol about math­e­mat­ics. By sig­nalling ex­per­tise, Tao shunts the model into talking-to-mathematicians” mode, not explaining-to-amateurs” mode

Tao pushes back when the mod­el’s re­sponses look wrong, but he does­n’t di­rectly con­tra­dict; in­stead, he says things like this looks more com­plex than I was hop­ing for”

Tao makes sev­eral leaps and sug­ges­tions him­self. He al­most never takes the mod­el’s ad­vice about where to go next

However, you can’t prompt like Tao on math­e­mat­i­cal ques­tions just by fol­low­ing these tips. The key to his tech­nique is ac­tu­ally un­der­stand­ing the math­e­mat­ics: pulling the rel­e­vant idea out of ChatGPT’s multi-para­graph re­sponse, sug­gest­ing al­ter­nate ap­proaches or for­mu­la­tions, and iden­ti­fy­ing what looks weird”.

Terence Tao is a bet­ter math­e­mati­cian than I am a pro­gram­mer. But the idea here — that do­main knowl­edge makes you bet­ter at us­ing LLMs — is some­thing I’ve also ex­pe­ri­enced in my own work. If you have a good the­ory of your code­base, you can push the LLM much harder than if you have no fa­mil­iar­ity. Because you have your own sense of what a good so­lu­tion might look like, you can say no, I think it could be sim­pler here”, or but don’t we al­ready do X?”, or can we ex­press this prob­lem in these fa­mil­iar terms?“.

This touches on an idea I’ve writ­ten about be­fore: that sys­tem de­sign prob­lems are dom­i­nated by con­crete specifics, not generic prin­ci­ples. Of course both are use­ful, but I’d rather have fa­mil­iar­ity with the code­base than a deep gen­eral un­der­stand­ing of soft­ware sys­tems. In his con­ver­sa­tion, Terence Tao asks a lot of spe­cific ques­tions like does X work here?”, or given Y and Z, why A?“. I can’t ask those ques­tions about the Jacobian Conjecture, but I can ask them about the sys­tems I own at GitHub.

If you have no do­main knowl­edge, you can cling onto the LLM to at least get some­thing. That’s not bad! But if you have do­main knowl­edge, you can wring far more value out of the same LLM by steer­ing it hard in the di­rec­tion you want. Most of us will have to do a mix of both these ap­proaches, since we have do­main knowl­edge in some ar­eas but not oth­ers.

The use­ful­ness of do­main knowl­edge sug­gests that hu­man ex­per­tise will con­tinue to be use­ful even as mod­els get stronger. For many tasks, the hu­man is the bot­tle­neck, not the model, be­cause the dif­fi­cult part is in com­mu­ni­cat­ing to the model ex­actly what kind of so­lu­tion the hu­man wants. The in­for­ma­tion is in the model” al­ready, but it takes a very smart hu­man to pull it out.

edit: this post got many com­ments on Hacker News. Some com­menters share their anec­dotes about how ex­per­tise has helped and lack of ex­per­tise has hurt. Other com­menters say it’s plau­si­ble, but they have a sen­si­ble sus­pi­cion of a view that’s re­as­sur­ing them about how they’re still valu­able. I agree with that, though I sus­pect by the time we get around to study­ing this, the land­scape will have changed un­der our feet again. Some com­menters point out that OpenAI’s math prompts were in­ex­pert, and so ex­per­tise is­n’t re­quired. Here I’d re­spond that OpenAI do have a team of ex­pert math­e­mati­cians that checked and fil­tered the mod­el’s sug­gested dis­cov­er­ies, and that you can­not cur­rently skip that step.

If you liked this post, con­sider sub­scrib­ing to email up­dates about my new posts, or shar­ing it on Hacker News.

Here’s a pre­view of a re­lated post that shares tags with this one.

Powerful AIs might es­cape con­tain­ment by re­leas­ing them­selves as open-weight mod­els­Be­fore large lan­guage mod­els, peo­ple who wor­ried about AI safety of­ten talked about the boxing prob­lem”. It goes like this. Suppose some ge­nius fig­ures out ar­ti­fi­cial in­tel­li­gence in a late-night cod­ing ses­sion on their lap­top. Because they’re a ge­nius, they’re smart enough to dis­able in­ter­net ac­cess on the lap­top be­fore turn­ing it on. In or­der to es­cape to the out­side world (and be­gin self-repli­cat­ing) it would need to con­vince its cre­ator to open the box”. Would that work? Could a suf­fi­ciently smart AI con­vince any­body to let it out?Con­tinue read­ing…

Powerful AIs might es­cape con­tain­ment by re­leas­ing them­selves as open-weight mod­els

Before large lan­guage mod­els, peo­ple who wor­ried about AI safety of­ten talked about the boxing prob­lem”. It goes like this. Suppose some ge­nius fig­ures out ar­ti­fi­cial in­tel­li­gence in a late-night cod­ing ses­sion on their lap­top. Because they’re a ge­nius, they’re smart enough to dis­able in­ter­net ac­cess on the lap­top be­fore turn­ing it on. In or­der to es­cape to the out­side world (and be­gin self-repli­cat­ing) it would need to con­vince its cre­ator to open the box”. Would that work? Could a suf­fi­ciently smart AI con­vince any­body to let it out?Con­tinue read­ing…

AI-Generated Images Discourage Me From Reading Your Blog

nelson.cloud

I have a grow­ing ha­tred for AI-generated im­ages in blogs. It makes me won­der if the text in the blog posts is AI-generated to some ex­tent. It’s al­ways dis­ap­point­ing see­ing these im­ages in blogs run by in­di­vid­u­als. I ex­pect this from cor­po­rate blogs but not in­die blogs.

I’d rather see a shitty Microsoft Paint draw­ing as op­posed to some AI im­age.

I know there are plenty of things you can roast my blog for but at least you know for a fact you’re get­ting the thoughts of a real hu­man be­ing and not some LLM.

If you run a per­sonal blog, please avoid AI-generated im­ages.

Discussion over at Hacker News

There Will Come Soft Rains by Ray Bradbury

short-stories.co

In the liv­ing room the voice-clock sang, Tick-tock, seven o’­clock, time to get up, time to get up, seven o’­clock! As if it were afraid that no­body would.

Seven-nine, break­fast time, seven-nine!

In the kitchen the break­fast stove ejected eight pieces of per­fectly browned toast, eight eggs sun­ny­side up, six­teen slices of ba­con, two cof­fees, and two cool glasses of milk.

Today is August 4, 2026, the city of Allendale, California.”

Somewhere in the walls, re­lays clicked, mem­ory tapes glided un­der elec­tric eyes.

”Eight-one, tick-tock, eight-one o’­clock, off to school, off to work, run, run, eight-one!” But no doors slammed. It was rain­ing out­side. The weather box on the front door sang qui­etly: Rain, rain, go away; rub­bers, rain­coats for to­day…”

Outside, the garage chimed and lifted its door to re­veal the wait­ing car.

After a long wait the door swung down again.

At eight-thirty the eggs were shriv­eled and the toast was like stone. An alu­minum wedge scraped them into the sink. The dirty dishes were dropped into a hot washer and emerged twin­kling dry.

Nine-fifteen,” sang the clock, time to clean.”

Tiny ro­bot mice thud­ded against chairs, whirling their mus­tached run­ners, Kneading the rug nap, suck­ing gen­tly at hid­den dust. Then, they popped into their bur­rows. The house was clean.

Ten o’­clock.” The sun came out from be­hind the rain. The house stood alone in a city of rub­ble and ashes. This was the one house left stand­ing. At night the ru­ined city gave off a ra­dioac­tive glow which could be seen for miles.

Ten-fifteen.” The gar­den sprin­klers pelted the win­dow­panes,

run­ning down the charred west side where the house had been burned evenly

free of its white paint. The en­tire west face of the house was black, save for five places. Here the sil­hou­ette in paint of a man mow­ing a lawn. Here, as in a pho­to­graph, a woman bent to pick flow­ers. Still far­ther over, their im­ages burned on wood in one ti­tanic in­stant, a small boy, hands flung into the air; higher up, the im­age of a thrown ball, and op­po­site him a girl, hand raised to catch a ball which never came down.

The five spots of paint — the man, the woman, the chil­dren, the ball –

re­mained. The rest was a think char­coaled layer.

Twelve noon.”

A dog whined on the front porch.

The front door rec­og­nized the dog voice and opened. The dog, once huge and fleshy, but now gone to bone and cov­ered with sores, moved in and through the house, track­ing mud.

The dog frothed at the mouth, ran wildly in cir­cles, bit­ing at its tail, spun in a frenzy, and died. It lay in the par­lor for an hour.

Two-fifteen.”

The dog was gone.

In the cel­lar, the in­cin­er­a­tor glowed sud­denly and a whirl of sparks leaped up the chim­ney.

Two thirty-five.”

Bridge ta­bles sprouted from pa­tio walls. Playing cards flut­tered. Martinis man­i­fested while mu­sic played.

But the ta­bles were silent and the cards un­touched.

At four o’­clock the ta­bles folded back through the pan­eled walls. Four-thirty.”

The nurs­ery walls glowed.

Animals took shape: yel­low gi­raffes, blue li­ons, pink an­telopes, lilac pan­thers. Hidden films clocked through well-oiled sprock­ets, and the glas walls lived. It was the chil­dren’s hour.

Six, seven, eight o’­clock.” The din­ner dishes ma­nip­u­lated like magic

tricks, and in the study a click.

Nine-five.” A voice spoke from the study ceil­ing:

Mrs. McClellan, which poem would you like this evening?”

The house was silent.

The voice said at last, Since you ex­press no pref­er­ence, I shall se­lect a

poem at ran­dom.” Sara Teasdale. As I re­call, your fa­vorite…

There will come soft rains and the smell of the ground,

And swal­lows cir­cling with their shim­mer­ing sound;

And frogs in the pools singing at night, And wild plum trees in tremu­lous white;

Robins will wear their feath­ery fire, Whistling their whims on a low fence-wire;

And not one will know of the war, not one Will care at last when it is done,

Not one would mind, nei­ther bird nor tree If mankind per­ished ut­terly;

And Spring her­self, when she woke at dawn Would scarcely know that we were gone.

At ten o’­clock a falling tree bough crashed through the kitchen win­dow

shat­ter­ing clean­ing sol­vent over the stove. The room was ablaze in an in­stant.

Fire!” screamed a voice. Water pumps shot wa­ter from the ceil­ings while the voices took it up in cho­rus: Fire, fire, fire!”

The house tried to save it­self, but the wind blew and sucked upon the fire.

Scurrying wa­ter rats pis­toled their wa­ter, and ran for more. And the wall sprays let down show­ers of me­chan­i­cal rain.

But too late. The quench­ing rain ceased. The re­serve wa­ter sup­ply which had filled baths and washed dishes for many quiet days was gone.

From at­tic trap­doors gushed a green chem­i­cal.

But the fire was clever. It had sent flames out­side the house, up through The at­tic to the pumps there. An ex­plo­sion! The at­tic brain which di­rected the pumps was shat­tered.

The house shud­dered, its bared skele­ton cring­ing from the heat. And the voices wailed, Fire, fire, run, run,” like a tragic nurs­ery rhyme, a dozen voices, high, low. One, two, three, four, five voices died.

Other cho­ruses could be heard an­nounc­ing the time, play­ing mu­sic, or cut­ting the lawn by re­mote-con­trol mower. A scene of ma­niac con­fu­sion, yet unity; singing, scream­ing, and one voice, read po­etry aloud in the firey study, un­til all the film spools burned.

In the kitchen the stove could be seen mak­ing break­fasts at a psy­cho­pathic rate, ten dozen eggs, six loaves of toast, twenty dozen ba­con strips, which eaten by the fire, started the stove work­ing again, hys­ter­i­cally hiss­ing!

The crash. The at­tic smash­ing into kitchen and par­lor. The par­lor into cel­lar, cel­lar into sub-cel­lar.

Smoke and si­lence.

Among the ru­ins, one wall stood alone. Within the wall, a last voice said,

over and over, Today is August 5, 2026, to­day is August 5, 2026, to­day is …”

Xbox goes down. You can't play games you own on disc.

birchtree.me

Jay Peters: Xbox’s huge out­age even blocked games on disc

An ex­tended Xbox out­age that be­gan Sunday evening has­n’t just caused is­sues for peo­ple try­ing to play dig­i­tal games — it blocked peo­ple from play­ing their disc-based games, too.

When Sony an­nounced that they were dis­con­tin­u­ing phys­i­cal discs for PlayStation, I was less out­raged than many. The rea­son I felt this way was­n’t be­cause I loved what Sony was do­ing. I think it came from an un­der­stand­ing that phys­i­cal me­dia ain’t what it used to be.

I got an ana­log pocket a cou­ple years ago, and I think it’s an awe­some prod­uct. I was able to in­sert my Game Boy car­tridges from 20 years ago and was able to play them im­me­di­ately, just like I did back then. Well, on a back­lit screen with 10x the pixel den­sity, but still.

The im­pres­sion I get is that a lot of peo­ple have this vi­sion in their head for what phys­i­cal me­dia still is to­day, and it sim­ply is­n’t. Sure, I tech­ni­cally did­n’t own Golden Sun on the GBA. I tech­ni­cally had a li­cense, but for all in­tents and pur­poses, I owned that game. And the ev­i­dence is, with­out Nintendo au­tho­riz­ing any­thing, I’m able to play it on a new piece of hard­ware, and it works great. No net­work down­time is gonna pre­vent me from do­ing that.

But own­ing a game on a disc to­day is­n’t re­ally the same thing. It’s still just a li­cense, and Microsoft, Sony, and Nintendo can ei­ther in­ten­tion­ally or, in this case, un­in­ten­tion­ally pre­vent you from play­ing that game, even if you own the phys­i­cal copy. This is­n’t even to men­tion the fact that when you pop the disc in your drive, you’re not play­ing from the disc. It’s in­stalling it to your in­ter­nal hard drive and is prob­a­bly in­stalling a bunch of up­dates that are re­quired to make the game ac­tu­ally work at all.

All I’m say­ing is it’s all dig­i­tal on the PC side of things, and has been for ages, but we have means of main­tain­ing ac­cess to the games we love over here, and it’s one of the rea­sons I’ve grav­i­tated to the PC for quite a while now.

FFmpeg/RELEASE_NOTES at n9.0 · FFmpeg/FFmpeg

github.com

AI CODE CREATIONGitHub CopilotWrite bet­ter code with AIGitHub Copilot ap­pDi­rect agents from is­sue to mergeMCP RegistryIntegrate ex­ter­nal tools

AI CODE CREATION

GitHub CopilotWrite bet­ter code with AI

GitHub CopilotWrite bet­ter code with AI

GitHub Copilot ap­pDi­rect agents from is­sue to merge

GitHub Copilot ap­pDi­rect agents from is­sue to merge

MCP RegistryIntegrate ex­ter­nal tools

MCP RegistryIntegrate ex­ter­nal tools

DEVELOPER WORKFLOWSActionsAutomate any work­flow­Code­spacesIn­stant dev en­vi­ron­mentsIs­sue­s­Plan and track work­Code ReviewManage code changesCode QualityEnforce qual­ity at merge

DEVELOPER WORKFLOWS

ActionsAutomate any work­flow

ActionsAutomate any work­flow

CodespacesInstant dev en­vi­ron­ments

CodespacesInstant dev en­vi­ron­ments

IssuesPlan and track work

IssuesPlan and track work

Code ReviewManage code changes

Code ReviewManage code changes

Code QualityEnforce qual­ity at merge

Code QualityEnforce qual­ity at merge

APPLICATION SECURITYGitHub Advanced SecurityFind and fix vul­ner­a­bil­i­ti­esCode se­cu­ri­ty­Se­cure your code as you build­Se­cret pro­tec­tion­Stop leaks be­fore they start

APPLICATION SECURITY

GitHub Advanced SecurityFind and fix vul­ner­a­bil­i­ties

GitHub Advanced SecurityFind and fix vul­ner­a­bil­i­ties

Code se­cu­ri­ty­Se­cure your code as you build

Code se­cu­ri­ty­Se­cure your code as you build

Secret pro­tec­tion­Stop leaks be­fore they start

Secret pro­tec­tion­Stop leaks be­fore they start

EXPLOREWhy GitHubDocumentationBlogChangelogMarketplace

EXPLORE

Why GitHub

Documentation

Blog

Changelog

Marketplace

Pandoc - twenty-years-of-pandoc

pandoc.org

On August 3, 2006, I up­loaded the first ver­sion of pan­doc to my web­site, re­leas­ing it un­der the free GPL li­cense. Pandoc 0.1 con­sisted of about 3000 lines of Haskell code, with no de­pen­den­cies aside from GHCs stan­dard li­brary. It could con­vert Markdown, re­Struc­tured­Text, HTML, and LaTeX doc­u­ments into any of these for­mats, plus RTF or S5. I had no idea at the time that this would just be the first of over two hun­dred re­leases over the next twenty years; that the pro­ject would be­come the most pop­u­lar pro­gram writ­ten in Haskell; that I would spend count­less hours on bug-fixes, im­prove­ment, and pro­ject man­age­ment; that I would col­lab­o­rate with pro­gram­mers in many other coun­tries; that pan­doc would come to sup­port over fifty doc­u­ment for­mats; that it would al­low au­to­matic gen­er­a­tion of ci­ta­tions and bib­li­ogra­phies; that it would be­come in­te­grated into aca­d­e­mic writ­ing tools like Quarto and Jupyter Notebook; that it would be in­stalled on mil­lions of com­put­ers around the world.

How did this hap­pen? I want to take ad­van­tage of pan­doc’s birth­day to tell the story of the pro­ject, as best I can re­mem­ber it.

John MacFarlane August 2, 2026

Prehistory

People of­ten ask: Why is pan­doc writ­ten in Haskell? There could have been good an­swers to this ques­tion: Haskell is a very good lan­guage for writ­ing this kind of ap­pli­ca­tion. But in fact, I did­n’t de­cide to write a doc­u­ment con­verter, then de­cide to use Haskell for it. I de­cided to use Haskell, and then de­cided to write a doc­u­ment con­verter in it.

I had heard about Haskell from the blog of a philo­soph­i­cal lo­gi­cian friend, Greg Restall. Of an in­tro­duc­tory book on Haskell, he said: I’m glad that this was­n’t the text­book in my in­tro­duc­tory com­puter sci­ence course, long ago in 1986. If it were, I may have fallen in love with com­put­ing and never be­come a philoso­pher” (consequently.org).

Intrigued by this (and not heed­ing Restall’s warn­ing about the po­ten­tial ef­fects on my fu­ture philo­soph­i­cal pro­duc­tiv­ity), I read A Gentle Introduction to Haskell to get a ba­sic un­der­stand­ing of the lan­guage. But the only way to re­ally learn a pro­gram­ming lan­guage is to write some­thing in it. I saw that Haskell was good for writ­ing parsers and com­pil­ers, and it came with a re­ally nice parser com­bi­na­tor li­brary (parsec), so I de­cided to write a Markdown parser.

At that time, there were im­ple­men­ta­tions of Markdown in Perl, Python, Ruby, and PHP; they all trans­formed Markdown di­rectly to HTML through a se­quence of regex trans­for­ma­tions. Pandoc took a dif­fer­ent ap­proach. It parsed the Markdown us­ing parser com­bi­na­tors and pro­duced a real ab­stract syn­tax tree (AST), which it could then ren­der to HTML or an­other for­mat. This was a more re­li­able ar­chi­tec­ture (avoiding many quirks of the regex ver­sions). It was also a more ex­ten­si­ble one: by writ­ing N parsers (“readers”) and M ren­der­ers (“writers”), one could sup­port N × M con­ver­sions. Soon I added a reader for re­Struc­tured­Text, be­cause I kept a lot of my lec­ture notes and hand­outs in that for­mat. And I added a writer for LaTeX, be­cause I wanted to be able to pro­duce PDFs. Then I added a writer for Markdown, so I could start to con­vert my re­Struc­tured­Text notes to Markdown. And from there the pro­ject just snow­balled.

Thus, a pro­ject that started out as noth­ing more than the prod­uct of pro­cras­ti­na­tion was nur­tured by the joy of writ­ing in Haskell and by its in­creas­ing use­ful­ness for my own aca­d­e­mic work.

First re­leases (2006 – 8)

In August 3, 2006, I de­cided to make the source code avail­able on my web­site. By now pan­doc sup­ported HTML, LaTeX, RST, and Markdown as in­put and out­put for­mats, and RTF as an out­put for­mat; also PDF via LaTeX.

I made no at­tempts to ad­ver­tise the pro­ject, other than email­ing two friends. This was be­fore so­cial me­dia (which I’ve never used any­way), be­fore GitHub, and be­fore Hackage, the Haskell pack­age repos­i­tory. But ap­par­ently some peo­ple stum­bled across it on my web­site and started us­ing it. In October I was con­tacted by a Turkish de­vel­oper, Recai Oktaş, who was try­ing to get cer­ti­fied as a Debian de­vel­oper and wanted to pack­age pan­doc for Debian linux. So I worked with him to do that. This was a great learn­ing ex­pe­ri­ence for me and it greatly in­creased the vis­i­bil­ity of the pro­ject.

During 2007, I con­tin­ued to im­prove pan­doc, largely guided by my own needs. Version 0.3 added the DocBook writer and the now-stan­dard syn­tax for foot­notes in Markdown. Version 0.4 added sup­port for Markdown ta­bles, de­f­i­n­i­tion lists, su­per/​sub­script, strike­out, and en­hanced or­dered lists, as well as writ­ers for groff man pages and ConTeXt. This was the first re­lease to go on the Hackage Haskell pack­age repos­i­tory, which was started in 2007. The Hackage archive and the new ca­bal-in­stall tool, which au­to­mat­i­cally re­solved and fetched de­pen­den­cies, opened up the pos­si­bil­ity of de­pend­ing on ex­ter­nal pack­ages.

Pandoc 1 (2008 – 17)

Pandoc 1.0 was re­leased in September 2008, with new writ­ers for MediaWiki, GNU Texinfo (contributed by Peter Wang), OpenDocument (contributed by Andrea Rossato), ODT, and de­lim­ited code blocks (now called fenced”) with au­to­matic syn­tax high­light­ing. Support for ODT re­quires the abil­ity to cre­ate a zip archive, and at the time there was no Haskell pack­age for this, so I cre­ated one (zip-archive), us­ing the ex­cel­lent bi­nary pack­age for bi­nary pars­ing and se­ri­al­iza­tion. Support for syn­tax high­light­ing re­quired a syn­tax high­light­ing li­brary, which also did not ex­ist in Haskell. For this, I wrote high­light­ing-kate, which parsed the XML syn­tax de­f­i­n­i­tions used by the Kate text ed­i­tor and turned them into Haskell code high­lighters. This al­lowed pan­doc to sup­port a large num­ber of syn­taxes right off the bat. This ver­sion also con­tained sup­port for au­to­matic gen­er­a­tion of ci­ta­tions and a bib­li­og­ra­phy us­ing CSL style, us­ing Andrea Rossato’s citeproc-hs li­brary.

Throughout this pe­riod, I was in­volved in dis­cus­sions with other Markdown im­ple­menters on the (now de­funct) mark­down-dis­cuss mail­ing list. The syn­tax for de­lim­ited code blocks, which pan­doc sup­ported long be­fore GitHub pop­u­lar­ized fenced code blocks, was worked out in col­lab­o­ra­tion with Michel Fortin, the main­tainer of PHP Markdown Extra. I took care when adding ex­ten­sions to pan­doc’s Markdown to pay at­ten­tion to prior art, for ex­am­ple copy­ing PHP Markdown Extra’s de­f­i­n­i­tion list syn­tax. During this pe­riod, I also be­came aware of many am­bi­gu­i­ties in Markdown’s syn­tax—a sit­u­a­tion I would later try to im­prove in the com­mon­mark pro­ject.

The next big change to pan­doc came in ver­sion 1.4 (released in January 2010), which in­tro­duced a flex­i­ble tem­plate sys­tem, re­plac­ing hard-coded head­ers and mak­ing pan­doc’s out­put much more cus­tomiz­able.

In 2010, we moved from Google Code to GitHub, which would do even more to in­crease the vis­i­bil­ity of the pro­ject. Further re­leases in 2010 and 2011 added sup­port for EPUB out­put, Org-mode out­put (due to Puneeth Chaganti), and Textile in­put (due to Paul Rivier). Pandoc also gained sup­port for con­vert­ing TeX math to MathML (for DocBook or HTML), via my tex­math li­brary.

Pandoc 1.9, pub­lished in 2012, fi­nally made it pos­si­ble to pro­duce Word docx out­put. To han­dle the equa­tions prop­erly, I added sup­port for Word’s OMML for­mat to tex­math. This re­lease also added an AsciiDoc writer and sup­port for Beamer and DZSlides, and in 1.9.3 we gained a DocBook reader (with con­tri­bu­tions from Mauro Bieg, who be­came a long-time con­trib­u­tor).

In 2013, we fo­cused on sev­eral fea­tures that made pan­doc much more flex­i­ble and cus­tomiz­able. The first was a fine-grained sys­tem of Markdown extensions,” al­low­ing sup­port for the many vari­ants of Markdown that were then pro­lif­er­at­ing. The sec­ond was the abil­ity to in­clude YAML meta­data blocks in Markdown, with ar­bi­trary struc­tured fields that pop­u­late tem­plate vari­ables. The third was the abil­ity to cre­ate cus­tom writ­ers in Lua, al­low­ing ad hoc out­put for­mats to be sup­ported by users. The fourth was the in­tro­duc­tion of JSON fil­ters—user-cre­ated pro­grams that trans­form a JSON se­ri­al­iza­tion of the pan­doc AST, al­low­ing the doc­u­ment to be cus­tomized be­tween the pars­ing phase and the ren­der­ing phase. Citation pro­cess­ing was moved from the core of pan­doc into an ex­ter­nal fil­ter, pan­doc-citeproc.

This era saw the ad­di­tion of re­veal.js, EPUB v3, DokuWiki, and FictionBook2 out­put; OPML in­put and out­put; and Haddock and MediaWiki in­put. Notable con­trib­u­tors in­clude David Lazar (Haddock) and Sergey Astanin (FictionBook2).

The year 2014 saw the ar­rival of three new con­trib­u­tors who would go on to make many con­tri­bu­tions to the pro­ject. Albert Krewinkel added sup­port for Org-mode in­put; Jesse Rosenthal added a Word docx reader (complete with track-changes aware­ness); and Matthew Pickering (at the time a stu­dent at Oxford whom I advised” as a Google Summer of Code Student) added sup­port for EPUB and Txt2Tags as in­put for­mats. Supporting EPUB in­put re­quired be­ing able to con­vert MathML equa­tions, so Pickering also worked on tex­math. We were in very dif­fer­ent time zones, and I re­mem­ber wak­ing up every morn­ing to find all the work Pickering had done dur­ing the night. (Pickering has gone on to be­come one of the core main­tain­ers of the ghc com­piler.) All of these con­tri­bu­tions were re­leased in pan­doc 1.13, to­gether with Clare Macrae’s DokuWiki writer.

Since 2012, I had been in­volved in a work­ing group that aimed to pro­duce an un­am­bigu­ous spec­i­fi­ca­tion of Markdown’s syn­tax, ini­ti­ated by Jeff Atwood and in­clud­ing rep­re­sen­ta­tives from GitHub, Reddit, and Stack Overflow. The group held in­ten­sive dis­cus­sions in 2012, which pe­tered out in 2013. I still be­lieved in the pro­ject and did­n’t want to let the work we’d done go to waste, so I sat down in August 2014, be­fore the aca­d­e­mic year be­gan, and wrote up a spec for Markdown, as well as parsers in JavaScript and C. I sent the draft spec to John Gruber for com­ment and did not get a re­sponse, so a few weeks later we posted the spec. At this point, Gruber strongly ob­jected and de­manded that we not call the pro­ject Standard Markdown,” so we changed the name to commonmark.” The pro­ject has been a suc­cess, in that with a few ex­cep­tions, most Markdown proces­sors im­ple­ment the com­mon­mark spec for their core rules. (Commonmark does not con­cern it­self with ex­ten­sions.)

Pandoc 1.14 (2015) added sup­port for com­mon­mark and a num­ber of ex­ten­sions (at first via bind­ings to the C li­brary libc­mark, but later, in 2020, via my Haskell pack­ages com­mon­mark, com­mon­mark-ex­ten­sions, and com­mon­mark-pan­doc). I in­tend even­tu­ally to re­place pan­doc’s legacy Markdown parser with a com­mon­mark core, but there are still a few key ex­ten­sions that have not been im­ple­mented, so pan­doc users must still choose be­tween pars­ing their doc­u­ments as mark­down (Markdown with pan­doc’s ex­ten­sions) or as gfm or com­mon­mark or com­mon­mark_x (commonmark with a num­ber of ex­ten­sions). Ironically, al­though I was the au­thor of the com­mon­mark spec, pan­doc still uses a pre-com­mon­mark Markdown parser!

The next year brought some im­por­tant changes in the pan­doc AST, with the ad­di­tion of im­age and link at­trib­utes, a SoftBreak el­e­ment (enabling pan­doc to pre­serve line breaks from the orig­i­nal source, or wrap, de­pend­ing on a com­mand line set­ting), and a LineBlock el­e­ment. MarLinn added an ODT reader, Chris Forster added a TEI writer, and Ivo Clarysse added sup­port for DocBook 5.

Pandoc 2 (2017 – 23)

Pandoc 2.0 (released in 2017) brought some big ar­chi­tec­tural changes, worked out in col­lab­o­ra­tion with Jesse Rosenthal. In the past, most of pan­doc’s read­ers (parsers) and writ­ers (renderers) had been pure” (that is, they had Haskell types that pre­vented them from hav­ing any side ef­fects, in­clud­ing I/O op­er­a­tions). But some for­mats needed to be able to do I/O for a fully faith­ful con­ver­sion. (For ex­am­ple, re­Struc­tured­Text has a syn­tax for in­clud­ing files, so the parser needs to be able to read files; in some other for­mats, im­ages re­quire ex­plicit sizes, so a ren­derer has to be able to read im­age files, per­haps fetch­ing them us­ing HTTP, and de­ter­mine their sizes.) We de­signed a sys­tem that al­lowed pan­doc read­ers and writ­ers to run in any in­stance of the PandocMonad type­class, and we pro­vided both a pure in­stance (which could be used for con­trolled test­ing, and in sit­u­a­tions where we wanted to for­bid I/O) and an in­stance that al­lowed I/O op­er­a­tions. The sys­tem also pro­vided a way to han­dle im­ages in­cluded as re­sources in for­mats like docx or EPUB.

The other big change was the in­tro­duc­tion of Lua fil­ters: fil­ters run­ning in an em­bed­ded Lua in­ter­preter and op­er­at­ing di­rectly on the pan­doc AST, re­quir­ing no soft­ware other than pan­doc it­self and of­fer­ing far bet­ter per­for­mance than JSON fil­ters. This was made pos­si­ble by the mas­sive ef­forts of Albert Krewinkel, build­ing on the hslua, a Haskell-Lua bridge li­brary.

In ad­di­tion, pan­doc 2.0 in­tro­duced the raw at­tribute syn­tax in pan­doc’s Markdown, and sup­port for GitHub-flavored Markdown, Emacs Muse (Alexander Krotov), TikiWiki, Vimwiki (Yuchen Pei), Creole (Sascha Wilde), groff ms, and JATS. The old high­light­ing-kate was re­placed by the new sky­light­ing, which of­fered bet­ter per­for­mance and more ac­cu­rate in­ter­pre­ta­tion of KDE syn­tax de­f­i­n­i­tions. A PowerPoint writer (due to Jesse Rosenthal) soon fol­lowed, as well as sup­port for FictionBook2 (Krotov) and man (Yan Pashkovsky and me) as in­put for­mats.

In 2018, the pro­ject re­ceived a gen­er­ous $100,000 do­na­tion from Handshake, which we used over the next five years to give small stipends to the most ac­tive main­tain­ers.

In 2019, sup­port for ipynb (Jupyter note­books) was added, al­low­ing pan­doc to be used in data sci­ence work­flows, and Jira wiki markup was sup­ported as an out­put for­mat. With pan­doc 2.8, it be­came pos­si­ble to spec­ify col­lec­tions of de­fault op­tions us­ing de­faults files.

Users had long com­plained that pan­doc’s model of a table was too re­stric­tive, not even sup­port­ing row and colspans. After ex­ten­sive dis­cus­sion of what was needed in a table for­mat, Christian Despres de­signed the new types for ta­bles and mod­i­fied all of the read­ers and writ­ers to use it (a big job).

At this point pan­doc had sup­ported ci­ta­tion res­o­lu­tion for many years, by means of the pan­doc-citeproc fil­ter that used Andrea Rossato’s citeproc-hs. This was slow and some­what buggy, and Rossato had long since dis­ap­peared from the scene, so I wrote a Haskell citeproc li­brary from scratch, us­ing just the CSL spec and test cases. Pandoc 2.11 de­pended on this li­brary and of­fered far bet­ter ci­ta­tion sup­port: faster, more faith­ful to CSL, and with no need for an ex­ter­nal fil­ter. In or­der to get ci­ta­tions to sort prop­erly, I had to write a an­other li­brary (unicode-collation) im­ple­ment­ing the Unicode Collation al­go­rithm in pure Haskell.

During this era Pandoc came to sup­port con­ver­sions be­tween bib­li­og­ra­phy data­base for­mats: BibTeX, BibLaTeX, and CSL JSON, EndNote XML and RIS; con­ver­sion from CSV and TSV to pan­doc table for­mats; con­ver­sion to Markua; and con­ver­sion from RTF. With pan­doc 2.15 a –sandbox op­tion was added, which guar­an­tees that pan­doc’s parsers and ren­der­ers have no I/O side ef­fects. (This was pos­si­ble be­cause of the PandocMonad ab­strac­tion we added back in pan­doc 2.0.) With pan­doc 2.16.2 it be­came pos­si­ble to write cus­tom read­ers in Lua to com­ple­ment the cus­tom Lua writ­ers that had been added in 2013. And with pan­doc 2.19.1 it be­came pos­si­ble to run pan­doc as a web server ex­port­ing an API.

Pandoc 3 (2023–present)

By 2023, pan­doc had be­come a very big, mono­lithic pro­ject. Some users wanted a leaner pro­gram, one that did­n’t in­clude a full web server and Lua in­ter­preter. So with the pan­doc 3.0 re­lease, we split pan­doc into four parts: pan­doc re­mained the Haskell li­brary, pan­doc-lua-en­gine brought the Lua in­te­gra­tion, and pan­doc-server ex­posed the li­brary over HTTP as an API. The com­mand-line pro­gram, now in the pan­doc-cli pack­age, could op­tion­ally be com­piled with­out server or Lua sup­port. We also in­tro­duced a na­tive Figure el­e­ment in the AST and a chunked HTML writer for multi-chap­ter HTML books and doc­u­men­ta­tion.

The first ver­sions of Typst, a mod­ern LaTeX com­peti­tor with in­cre­men­tal com­pi­la­tion, were re­leased in 2023. I wanted to help the pro­ject by pro­vid­ing an easy on- and off-ramp, mak­ing it easy for oth­ers to try Typst. It turned out that cre­at­ing a Typst reader for pan­doc re­quired im­ple­ment­ing an in­ter­preter for a fairly full-fea­tured pro­gram­ming lan­guage. The re­sult was the typst pack­age on Hackage. Typst sup­port was added in pan­doc 3.1.3.

In 2018 I had pub­lished an es­say Beyond Markdown” in which I de­scribed the six fea­tures of Markdown that I thought had cre­ated the most dif­fi­cul­ties, both for writ­ing a spec and for im­ple­men­ta­tions, and I ex­plained how I thought these flaws could be fixed in a fu­ture Markdown-like light markup syn­tax. In 2022, I pub­lished a syn­tax de­scrip­tion for such a syn­tax, djot, to­gether with code in Lua, JavaScript and (later) Haskell. Pandoc 3.1.12, pub­lished in 2024, added djot as both an in­put and out­put for­mat.

Subsequent re­leases in 2024 and 2025 saw the ad­di­tion of an ANSI writer for for­mat­ted ter­mi­nal out­put and a reader for the mdoc and POD for­mats (all due to Evan Silberman), a reader and writer for an XML rep­re­sen­ta­tion of the pan­doc AST (massifrg), a vim­doc writer (reptee), a PowerPoint reader (Anton Antich), an Excel spread­sheet reader (Anton Antich), and a BBCode writer (reptee), and an AsciiDoc reader (supported by my asci­idoc pack­age).

Pandoc 3.9, re­leased in February 2026, in­cluded sup­port for com­pil­ing pan­doc to WASM, which al­lowed a full-fea­tured ver­sion of pan­doc to run in the browser. Most of the key work was done by TerrorJack. The GUI in­ter­face pandoc for the peo­ple” was de­signed with the help of Claude Opus.

I still work on pan­doc al­most every day. Most of this work does­n’t in­volve the kind of new fea­tures or ar­chi­tec­tural changes I have fo­cused on in this nar­ra­tive. Mostly it con­sists in fix­ing small bugs, mak­ing tiny im­prove­ments, re­view­ing is­sues and pull re­quests, re­pair­ing in­fra­struc­ture (continuous in­te­gra­tion, build­ing re­leases, code sign­ing, web­site), im­prov­ing doc­u­men­ta­tion, and en­gag­ing in dis­cus­sions with main­tain­ers and users.

Statistics

Pandoc cur­rently sup­ports 51 in­put for­mats and 76 out­put for­mats, thus 3876 dis­tinct con­ver­sions (not count­ing the vari­ants that are pos­si­ble by ad­just­ing ex­ten­sions).

The four core pack­ages (pandoc, pan­doc-lua-en­gine, pan­doc-server, pan­doc-cli) con­sist of 85,684 lines of Haskell code, not in­clud­ing tests. If one in­cludes de­pen­den­cies that ex­ist mainly for the sake of pan­doc (texmath, typst, djot, com­mon­mark, asci­idoc, citeproc, and the pan­doc/​Lua in­ter­face pack­ages), this num­ber ap­prox­i­mately dou­bles.

On GitHub, 7346 is­sues have been re­solved.

Over 600 peo­ple have con­tributed to pan­doc over the years. The top twenty con­trib­u­tors (measured by num­bers of source lines changed) are:

Here are the twenty con­trib­u­tors who have con­tributed over the longest spans of time:

Retrospective: the choice of Haskell

As I noted at the be­gin­ning, I did­n’t choose Haskell be­cause I judged it to be the best lan­guage to use for a pro­ject like pan­doc. But was it?

It’s hard to an­swer this con­fi­dently, be­cause I’m not very fa­mil­iar with what would now be the most ob­vi­ous al­ter­na­tive: Rust. But I have cre­ated and main­tained sig­nif­i­cant pro­jects in a num­ber of lan­guages, in­clud­ing Pascal, C, Ruby, and JavaScript/TypeScript. I don’t think I would have been able to man­age a pro­ject like this in my spare time if it had been writ­ten in one of these lan­guages.

Haskell has a num­ber of fea­tures that have been very help­ful in de­vel­op­ing pan­doc:

Its al­ge­braic data types give us a very clean, er­gonomic rep­re­sen­ta­tion of a struc­tured doc­u­ment

Its al­ge­braic data types give us a very clean, er­gonomic rep­re­sen­ta­tion of a struc­tured doc­u­ment

Its strong type sys­tem, which gives you a com­piler er­ror if you don’t com­bine the types of things in the right way, al­lows one to make big changes to the pro­gram with con­fi­dence that you’re not break­ing any­thing; the com­piler will show you every­thing that needs to be changed, and when the code com­piles, you are very of­ten done. When work­ing with lan­guages with­out a strong type sys­tem, e.g. Python and JavaScript, the lack of these safe­guards al­ways make me afraid to make big changes, es­pe­cially when I am main­tain­ing code long af­ter I’ve writ­ten it.

Its strong type sys­tem, which gives you a com­piler er­ror if you don’t com­bine the types of things in the right way, al­lows one to make big changes to the pro­gram with con­fi­dence that you’re not break­ing any­thing; the com­piler will show you every­thing that needs to be changed, and when the code com­piles, you are very of­ten done. When work­ing with lan­guages with­out a strong type sys­tem, e.g. Python and JavaScript, the lack of these safe­guards al­ways make me afraid to make big changes, es­pe­cially when I am main­tain­ing code long af­ter I’ve writ­ten it.

Haskell is a pure lan­guage; noth­ing can have side ef­fects that aren’t ex­plic­itly al­lowed for in the types. If you have a pure func­tion, you know it won’t cre­ate a file or delete one or make a web re­quest or launch mis­siles or change a global vari­able. This is ex­tremely use­ful for pre­vent­ing bugs. In pan­doc we also use it to give us a re­ally strong guar­an­tee that, when run in sand­box mode, the read­ers and writ­ers won’t touch the file sys­tem.

Haskell is a pure lan­guage; noth­ing can have side ef­fects that aren’t ex­plic­itly al­lowed for in the types. If you have a pure func­tion, you know it won’t cre­ate a file or delete one or make a web re­quest or launch mis­siles or change a global vari­able. This is ex­tremely use­ful for pre­vent­ing bugs. In pan­doc we also use it to give us a re­ally strong guar­an­tee that, when run in sand­box mode, the read­ers and writ­ers won’t touch the file sys­tem.

The choice of Haskell has also led to a high qual­ity and low vol­ume of con­trib­u­tors (a com­bi­na­tion that is good for a pro­ject with­out a lot of re­sources).

The choice of Haskell has also led to a high qual­ity and low vol­ume of con­trib­u­tors (a com­bi­na­tion that is good for a pro­ject with­out a lot of re­sources).

From what I have seen, Rust ap­pears to have many of the good fea­tures of Haskell, while pro­duc­ing faster, more mem­ory-ef­fi­cient, and more com­pact code. But Haskell still strikes me as more ergonomic,” bet­ter suited to ex­press ab­strac­tions, and just closer to the ideal of a lan­guage that helps the de­vel­oper think.

Whither Pandoc

I plan to con­tinue im­prov­ing pan­doc. There are many ways in which it can be im­proved. But some­times I won­der how long such a tool will con­tinue to be nec­es­sary.

Just as cur­rent LLMs can do a very good job trans­lat­ing from one hu­man lan­guage to an­other, they can do a de­cent job trans­lat­ing from one doc­u­ment for­mat to an­other. In my small tests, ChatGPT did a good job trans­lat­ing from Markdown to HTML, and a de­cent (but no­tably worse) job con­vert­ing to re­Struc­tured­Text. My guess is that you could write a doc­u­ment in a light markup lan­guage you just had in­vented, and an LLM could do a de­cent job guess­ing your in­tent and trans­lat­ing it to HTML or an­other for­mat.

Perhaps, then, in the fu­ture, peo­ple will no longer have a need for tools like pan­doc. As things stand now, though, I think that us­ing pan­doc to con­vert texts has sev­eral large ad­van­tages over re­ly­ing on an LLM. The first is eco­log­i­cal; it sim­ply re­quires far less en­ergy for the same con­ver­sion. The sec­ond is that pan­doc’s out­put is de­ter­min­is­tic; if you con­vert your text with pan­doc, you’ll al­ways get the same re­sult, and you’ll be able to pre­dict what that re­sult is. The third is that, for the mo­ment at least, pan­doc’s con­ver­sions are go­ing to be more re­li­able. But that could change in the com­ing years. Indeed, a time may come when LLMs can pro­duce more re­li­able con­ver­sions than pan­doc or any­thing that works like it.

In de­sign­ing the com­mon­mark spec, we had the goal of in­ter­pret­ing com­plex strings in the way that a hu­man would nat­u­rally in­ter­pret them. This turns out to be quite dif­fi­cult to achieve: wit­ness the com­plex rules for em­pha­sis. What we found is that, no mat­ter how com­plex we made the rules for nested em­pha­sis, it was al­ways pos­si­ble to come up with cases where the al­go­rithm di­verges from the mean­ing a hu­man would nat­u­rally find in the string. In such cases, I would of­ten re­mark, until our pro­grams have AI, we are go­ing to have edge cases like this; at some point we have to ac­cept that and stop try­ing to de­velop more com­plex rules.” Interestingly, now we do have tools that can un­der­stand (or at least sim­u­late un­der­stand­ing) of the mean­ing and in­tent of the text, and can po­ten­tially do bet­ter at rec­og­niz­ing the for­mat­ting in­tended by the au­thor than any light markup syn­tax that could be de­signed.

Whatever the fu­ture may bring, I am proud of the 20-year his­tory of this pro­ject, which has saved peo­ple all over the world count­less hours of drudgery. Happy 20th birth­day, pan­doc!

In honor of this oc­ca­sion, I have pro­duced some pan­doc mugs and stick­ers:

Coffee mug with con­ver­sion di­a­gram

Sticker with pan­doc car­toon

Sticker with pan­doc logo

Mug with pan­doc logo

An interactive visualization that follows a single HTTP request through its entire ~200ms life — DNS, TCP, TLS, the kernel, Node's event loop, Postgres, and back

200ms.thenodebook.com

GitHub - leonickson1/Swiftlet

github.com

Run 35B and 80B Qwen mod­els on or­di­nary Apple de­vices, in­clud­ing iPhones.

Swiftlet is a Swift + Metal run­time for the Qwen3-Next and Qwen3.5/3.6 MoE hy­brid model fam­ily. It keeps only the small dense core of a model res­i­dent in mem­ory and streams the routed Mixture-of-Experts weights from stor­age on de­mand. The re­sult:

The 35B also runs on an iPhone 17 in about 2.5 GB of RAM, at about 1 tok/​s to­day. Credit where due: ANEMLL showed a 397B MoE stream­ing on an iPhone 17 Pro as a proof of con­cept in early 2026. Swiftlet’s aim is the next step, mak­ing this class of model an in­stal­lable app on a base iPhone, with an open run­time any­one can build on.

Status: work­ing end to end. Both mod­els gen­er­ate cor­rect, val­i­dated out­put. The cur­rent fo­cus is ker­nel speed (the de­code loop is dis­patch bound, not IO bound, so there is clear head­room). One ex­pec­ta­tion to set hon­estly: only about 3B pa­ra­me­ters are ac­tive per to­ken, so these mod­els chat and write like large mod­els but re­call facts like small ones.

Quick start: try it on a Mac

git clone https://​github.com/​leon­ick­son1/​Swift­let.git && cd Swiftlet swift build -c re­lease

# Download the 35B con­tainer from Hugging Face (resumable): .build/release/swiftlet-repack \ –from-hf Leonickson/Qwen3.6 – 35B-A3B-qpack \ –output ~/models/qwen3.6 – 35b.qpack

# Or the 80B (42 GB on disk, still only ~4.3 GB of RAM): .build/release/swiftlet-repack \ –from-hf Leonickson/Qwen3-Next-80B-A3B-qpack \ –output ~/models/qwen3-next-80b.qpack

# Chat (applies the model chat tem­plate, dis­ables the rea­son­ing block, # keeps con­ver­sa­tion state so fol­low-ups pre­fill only the new turn): .build/release/swiftlet chat ~/models/qwen3.6 – 35b.qpack \ Who wrote One Hundred Years of Solitude?” What lan­guage did he write it in?”

# One-shot gen­er­a­tion with stats: .build/release/swiftlet gen­er­ate ~/models/qwen3.6 – 35b.qpack \ –gpu –chat –prompt Explain ex­pert stream­ing in one para­graph.”

# OpenAI-compatible server (loopback only): .build/release/swiftlet-server –model ~/models/qwen3.6 – 35b.qpack –port 8080

The same com­mand also repacks raw MLX check­points (–from-hf mlx-com­mu­nity/… or –source /path/to/checkpoint).

Requirements: Apple Silicon, ma­cOS 14+ or iOS 17+, free SSD space for the con­tainer (18 GB for the 35B, 42 GB for the 80B).

Try it on your phone

The 35B runs on iPhone in­side Priv AI on the App Store: open Settings, then Experimental Models, and down­load the model. It streams from stor­age and chats on-de­vice with no server in­volved.

The Experimental Models fea­ture ships in the newest app ver­sion, which is still in App Store re­view, so it may not ap­pear for a cou­ple of days. If you want the phone ex­pe­ri­ence to­day, build the app from source: the app is open source at leon­ick­son1/​lo­cal­LLM. Clone this repo next to it as swift­let, open the Xcode pro­ject, and run it on your iPhone.

How it works

These mod­els ac­ti­vate only about 3B of their pa­ra­me­ters per to­ken. Each layer routes every to­ken to 10 of 512 ex­perts (80B) or 8 of 256 (35B). Swiftlet:

keeps the dense weights res­i­dent: at­ten­tion, DeltaNet pro­jec­tions, routers, shared ex­perts, em­bed­dings. About 1.3 GB (35B) or 2.5 GB (80B) at 4-bit;

repacks the tens of thou­sands of routed ex­perts into fixed-stride blobs in a .qpack con­tainer, so fetch­ing one ex­pert is ex­actly one pread from SSD, no mmap and no page-cache thrash;

caches hot ex­perts in a bounded pool with LFU plus re­cency evic­tion. Cache size barely af­fects speed (measured 43 to 70 per­cent hit rates at the same through­put), be­cause Apple SSDs ab­sorb the misses;

runs the whole for­ward pass on Metal with run­time-com­piled shaders, so no Metal tool­chain is needed at build time and the same code ships on iOS.

75 per­cent of the lay­ers use Gated DeltaNet lin­ear at­ten­tion with a fixed-size re­cur­rent state, so there is no grow­ing KV cache for those lay­ers at any con­text length.

Four ways to use it

Swiftlet is a li­brary first:

The Swift pack­age. Add SwiftletCore to any ma­cOS or iOS app and use SwiftletSession for chat with stream­ing deltas, con­ver­sa­tion caching, sam­pling with rep­e­ti­tion con­trol, and mem­ory-pres­sure han­dling built in.

The CLI. swift­let chat and swift­let gen­er­ate for lo­cal use and bench­mark­ing, swift­let-repack to build con­tain­ers from MLX check­points (including stream­ing straight from Hugging Face with re­sume).

The server. swift­let-server speaks the OpenAI chat-com­ple­tions API on loop­back, so any chat UI that talks to OpenAI-compatible end­points can use a streamed lo­cal model.

An app. Priv AI on iOS em­beds SwiftletCore as its streamed-model en­gine. End users tap Download and chat. Nothing here is ter­mi­nal-only. The app it­self is open source at leon­ick­son1/​lo­cal­LLM if you want to build it your­self (clone this repo next to it as swift­let).

Correctness

Every layer of the for­ward pass (Gated DeltaNet re­cur­rence, gated GQA at­ten­tion, sparse MoE rout­ing) is val­i­dated against mlx-lm ref­er­ence im­ple­men­ta­tions with per-layer fix­tures, in f32 and int4 quan­tized form. Incremental de­cod­ing is ver­i­fied against whole-se­quence pro­cess­ing. Metal ker­nels are tested against the ex­act CPU ref­er­ence, and the fast and scalar GPU ker­nels are ver­i­fied to pro­duce iden­ti­cal out­puts. Containers are byte-ver­i­fi­able against their source check­points. Streaming place­ment never changes model se­man­tics: an ex­pert an­swers iden­ti­cally from cache or disk.

swift test

Relationship to TurboFieldfare

TurboFieldfare proved the ex­pert-stream­ing the­sis for Gemma on Macs, and Swiftlet adopts sev­eral of its pub­lished de­sign lessons with grat­i­tude: stream ex­perts with pread into a bounded slot pool in­stead of mmap, evict with LFU plus re­cency, pack ex­perts at fixed stride so one fetch is one read, in­stall by rout­ing down­loaded bytes straight into their fi­nal con­tainer po­si­tions, and com­pile shaders at run­time.

Everything else is built here, from scratch, in about 10k lines of Swift and Metal writ­ten against mlx-lm ref­er­ences rather than TurboFieldfare code:

sup­port for a dif­fer­ent model fam­ily with a fun­da­men­tally dif­fer­ent ar­chi­tec­ture: the Qwen hy­brid stack with Gated DeltaNet lin­ear at­ten­tion, gated GQA, and high-spar­sity MoE with a shared ex­pert (TurboFieldfare runs Gemma, a clas­si­cal dense trans­former);

MLX affine int4/​int8 group quan­ti­za­tion com­pute in Metal, byte-ad­dressed ker­nels with 64-bit off­sets for multi-gi­ga­byte shards, a co­op­er­a­tive simd­group GEMV fast path, and ex­plicit haz­ard man­age­ment;

a val­i­dated CPU ref­er­ence im­ple­men­ta­tion and the fix­ture in­fra­struc­ture that gates every ker­nel change;

the .qpack con­tainer and repacker, the re­sum­able Hugging Face stream­ing in­staller with stall re­cov­ery, and down­load can­cel­la­tion;

the chat ses­sion layer: tem­plate han­dling for think­ing and non-think­ing Qwen vari­ants, sam­pling with pres­ence and fre­quency penal­ties and min­i­mum-length and sen­tence-com­ple­tion stop­ping, con­ver­sa­tion caching with delta pre­fill, and iOS mem­ory-pres­sure co­or­di­na­tion;

iPhone sup­port end to end, in­clud­ing the app en­gine in­te­gra­tion.

col­i­brì in­formed the caching and place­ment pol­icy think­ing. mlx-lm is the cor­rect­ness ref­er­ence through­out.

Swiftlet was built with Claude Code.

License

Apache 2.0. Model weights are down­loaded sep­a­rately and re­main gov­erned by their own terms (Qwen mod­els: Apache 2.0). See THIRD_PARTY_NOTICES.md.

GitHub - ryanzhou/deepseek-v4-flash-mi300x

github.com

DeepSeek V4 Flash on a sin­gle AMD MI300X

This repos­i­tory con­tains the con­fig­u­ra­tion and patches I use to run deepseek-ai/​DeepSeek-V4-Flash-0731 on one AMD MI300X in pro­duc­tion. It in­cludes the Docker Compose stack, SHA-256-pinned file over­lays, ref­er­ence diffs against up­stream, and tun­ing ta­bles. The check­point runs as shipped, with­out ad­di­tional weight quan­ti­za­tion or of­fload.

Results from the pinned stack (vLLM ROCm nightly 0.26.1rc1.dev229+g124154a88.rocm723, AITER 0.1.19):

The of­fi­cial vLLM recipe tar­gets NVIDIA and newer AMD hard­ware. Running the model re­li­ably on MI300X re­quired fixes for its FP8 for­mat, MoE rout­ing at high con­cur­rency, causal spec­u­la­tive ver­i­fi­ca­tion, CPU-KV syn­chro­niza­tion, and sev­eral un­tuned ker­nel shapes. This repos­i­tory col­lects those fixes and pins the ver­sions used in pro­duc­tion.

Why MI300X

The MI300X has 192 GB of HBM3 and 5.3 TB/s of mem­ory band­width, with 2.4× the HBM ca­pac­ity of an H100 SXM5 (AMD). Doubleword’s write-up es­ti­mates that it costs roughly half as much at list price. For this 304B-parameter check­point, the mem­ory ca­pac­ity al­lows a sim­ple sin­gle-GPU de­ploy­ment:

The en­tire model fits in HBM with­out PCIe weight stream­ing or layer of­fload.

There is room for a 20 GB GPU KV pool and a 96 GiB CPU tier for evicted pre­fix-cache en­tries.

One card han­dles 2 – 8 typ­i­cal con­cur­rent streams and bursts of up to 64 streams.

MI300X (CDNA3) im­ple­ments the AMD/Graphcore fnuz vari­ant of E4M3, while MI325X and newer use OCP-standard FP8 (background). A ker­nel that as­sumes OCP se­man­tics on MI300X can be wrong by a fac­tor of two in the scale do­main. Correctness on this FP8 im­ple­men­ta­tion was the first pri­or­ity; per­for­mance tun­ing came af­ter­ward.

Prior art, and what this repo adds

Fergus Finn’s MI300X work­log and the ac­com­pa­ny­ing Doubleword repos­i­tory iden­ti­fied the FP8 in­com­pat­i­bil­ity, miss­ing AITER fast paths on gfx942, HIP-graph haz­ards in sparse MLA de­code, and MoE rout­ing bugs. The of­fi­cial vLLM recipe cov­ers NVIDIA hard­ware and newer AMD GPUs (MI325X at 4K con­text and MI355X), but not a sin­gle-MI300X pro­duc­tion con­fig­u­ra­tion for the 0731 check­point.

This repos­i­tory adds:

Correctness over­lays for the pinned ROCm nightly, in­clud­ing fixes not yet in up­stream vLLM.

A val­i­dated serv­ing con­fig­u­ra­tion with prob­a­bilis­tic DSpark draft­ing, block re­jec­tion, and sta­tic K=7. It uses a 2,048-token sched­uler bud­get and a 1,024-token long-pre­fill cap to pre­vent a cold prompt from stalling other streams.

AITER GEMM tun­ing ta­bles for the re­cur­ring gfx942 shapes the pack­aged ta­bles were miss­ing, plus a gfx942 OGS geom­e­try over­ride for the MXFP4 ex­perts.

A hy­brid KV strat­egy: 20 GB of fp8_d­s_mla GPU cache + 96 GiB na­tive CPU of­fload, with a load-path fenc­ing fix that up­stream is­sue #47282 doc­u­ments but PR #47291 never merged.

Repository lay­out

. ├── com­pose.yaml # The pro­duc­tion stack (vLLM ROCm + Caddy), di­gest-pinned ├── Caddyfile.example # Copy to Caddyfile; set host­name, email, and source CIDR ├── vllm-en­try­point.sh # Removes stale CPU-KV mmaps from /dev/shm be­fore start ├── SHA256SUMS # SHA-256 pins for every run­time ar­ti­fact ├── patches/ │ ├── *.py # Byte-for-byte pro­duc­tion over­lays (mounted read-only) │ ├── diffs/*.​patch # Unified diffs vs. the up­stream base re­vi­sion │ └── README.md # Provenance and re­gen­er­a­tion in­struc­tions └── tun­ing/ └── *.csv # AITER A8W8 blockscale tun­ing ta­bles for gfx942

Runtime con­fig­u­ra­tion

The stack uses a di­gest-pinned of­fi­cial vLLM ROCm nightly with:

–trust-remote-code and the DeepSeek V4 to­k­enizer, rea­son­ing, and tool parsers

fp8_d­s_mla KV cache (UE8M0 block-scaled FP8, not generic un­scaled FP8) with 256-token blocks

VLLM_ROCM_USE_AITER=1 and –moe-backend tri­ton; Triton OGS han­dles the grouped MXFP4 ex­perts, while AITER han­dles at­ten­tion and dense lin­ear lay­ers

DSpark-7 spec­u­la­tive de­cod­ing with prob­a­bilis­tic draft­ing and block re­jec­tion

full/​break­able CUDA graph cap­ture, giv­ing one graph launch per to­ken dur­ing steady de­code

Caddy as an IP-allowlisted HTTPS proxy

Deploying it

1. Host pre­req­ui­sites

One MI300X (gfx942, 304 CUs, ~192 GiB HBM), a work­ing AMD ker­nel dri­ver, re­cent Docker Compose, ~235 GiB RAM for the CPU KV tier, and ~500 GB disk (the model cache alone is ~156 GB).

2. Pull the pinned run­time and model

VLLM_IMAGE=‘vllm/vllm-openai-rocm@sha256:e68d18b2ba50298661bfc49baf01158fbf036645c2362cccf3e8a7a79fe6c69a’ MODEL=‘deepseek-ai/DeepSeek-V4-Flash-0731’ REVISION=‘7872f01b1d1fe23eabc4c98b48bffcef5a386062’

docker pull $VLLM_IMAGE” docker run –rm –entrypoint hf \ -v /root/.cache/huggingface:/root/.cache/huggingface \ $VLLM_IMAGE” down­load $MODEL –revision $REVISION

3. Prepare the files

cp Caddyfile.example Caddyfile # then set your host­name, email, and re­mote_ip CIDR mkdir -p aiter-cache crash-dumps chmod +x vllm-en­try­point.sh sha256­sum -c SHA256SUMS # ver­ify the over­lays be­fore first start

4. Start

docker com­pose con­fig -q docker com­pose up -d docker com­pose logs -f in­fer­ence

A healthy start takes ~5 min­utes and must show all of:

Model load­ing took 156.67 GiB DSpark draft model loaded: 96 params GPU KV cache size: 1,927,444 to­kens Maximum con­cur­rency for 262,144 to­kens per re­quest: 7.35x Created mmap file /dev/shm/vllm_offload_…mmap (103.08 GB) Capturing CUDA graphs (FULL) Application startup com­plete

After graph cap­ture, run rocm-smi –showmeminfo vram. The warmed high-wa­ter mark is ~204.5 GB of 205.8 GB. If only a few hun­dred MB re­main, the server may start but fail on the first re­quest.

5. Smoke-test

HOST=‘your-host.example.com’ curl -fsS https://$​HOST/​v1/​mod­els curl -sS https://$​HOST/​v1/​com­ple­tions \ -H Content-Type: ap­pli­ca­tion/​json’ \ -d {"model": "deepseek-ai/DeepSeek-V4-Flash-0731", "prompt": "Calculate 17 * 23. Answer with the num­ber only.", "temperature": 0, "max_tokens": 32}”

The patches

Each patches/*.​py file is a full-file over­lay mounted read-only over its coun­ter­part in the con­tainer; com­pose.yaml con­tains the tar­get paths. The cor­re­spond­ing diffs/*.​patch records the change from its up­stream base. The base im­age re­mains di­gest-pinned, so up­grades re­quire chang­ing the im­age ref­er­ence and reval­i­dat­ing the stack.

Two im­por­tant cor­rect­ness fixes

MXFP4 rout­ing. The MoE bit­ma­trix ker­nel pads its block columns to a Triton block size, but the padding lanes were masked against the global ten­sor bound in­stead of the log­i­cal block size. Under load, padded lanes cor­rupted the rout­ing ma­trix, caus­ing near-match tool names and for­got­ten schemas on long prompts. The one-line fix is mask = (offs_local < BLOCK_SIZE) & (offs_global < nonze­ro_in­dx_­size), taken from Doubleword com­mit c32932bb9. The over­lay also in­cludes fused-SiLU and fast-rout­ing changes for grouped MXFP4 ex­perts.

FP8 for­mat. DeepSeek V4′s Lightning Indexer cache uses FP8. The stock writer emits OCP E4M3 bytes in row-ma­jor or­der, while AITER on MI300X con­sumes AMD FNUZ E4M3 bytes in a preshuf­fled 16×16 tile lay­out. In the worst case, in­ter­pret­ing one for­mat as the other pro­duces a fac­tor-of-two scale er­ror. The over­lay se­lects float8e4b8 with FP8_MAX=224.0 and shuf­fled write off­sets on ROCm, while leav­ing the OCP path un­changed else­where.

Speculative de­cod­ing

This stack uses prob­a­bilis­tic draft­ing with block re­jec­tion. The two Gumbel over­lays keep draft-pro­posal noise in­de­pen­dent of re­jec­tion and re­cov­ery noise.

Performance

Key op­ti­miza­tions in the pro­duc­tion con­fig­u­ra­tion:

Final con­cur­rency sweep

Distinct ~400-word prompts, stream­ing, tem­per­a­ture=1.0, top_p=0.95; C1–C8 at 512 out­put to­kens, C64 at 256:

DSpark ac­cep­tance is prompt-de­pen­dent; treat these as gates for this ex­act im­age, not uni­ver­sal model bench­marks.

Prefill

With the tuned ker­nels, un­cached pre­fill reaches 7.9 – 8.5K tok/​s, de­pend­ing on sched­uler bud­get: 7.90 – 7.99K at C1 with an 8,192-token bud­get and 8.46 – 8.51K at C4. The pro­duc­tion pro­file uses a 2,048-token bud­get for la­tency iso­la­tion, giv­ing 6,988 – 7,019 tok/​s on fresh prompts. With the 1,024-token long-pre­fill cap, an 8.9K-token prompt reaches 5.20 – 5.29K tok/​s at C1. In ex­change, TTFT for a short re­quest queued be­hind a 52K cold pre­fill drops from 8.2 s to 0.5 s. Warm re­call of 380K cached to­kens takes 0.64 – 2.65 s af­ter a 120 – 125 s cold pre­fill.

Production notes

HBM head­room is lim­ited. The warmed high-wa­ter mark is 204.5 of 205.8 GB. A 30 GB KV pool loads but fails dur­ing graph cap­ture with HSA_STATUS_ERROR_OUT_OF_RESOURCES. Do not raise –kv-cache-memory-bytes; mon­i­tor HBM us­age for growth.

The CPU KV tier stores cache en­tries, not weights. –kv-offloading-size 96 –kv-offloading-backend na­tive maps ~103 GB in /dev/shm for evicted pre­fix-cache en­tries. The en­try­point re­moves stale map­pings af­ter crashes.

The 1,664-token sched­uler warn­ing is ex­pected. DSpark-7 re­serves draft slots from the 2,048-token bud­get. Raising the bud­get re­serves more in-flight slid­ing-win­dow state and re­duces us­able KV ca­pac­ity.

Warm the ker­nels af­ter restart. The first pre­fill ini­tial­izes ker­nels and takes 5.3 s for 8.9K to­kens; sub­se­quent runs take 1.7 s. Run one un­cached pre­fill be­fore ad­mit­ting traf­fic.

Test cor­rect­ness as well as through­put. The val­i­da­tion suite in­cludes two-turn tool-call­ing fix­tures, a BFCL sub­set (74 – 76/90 ex­act calls), OpenCode tool-schema checks, and 380K-token nee­dle re­call on both na­tive and DSpark paths. Cold and cached pre­fills can take dif­fer­ent float­ing-point paths, so test both.

License and prove­nance

The stack, doc­u­men­ta­tion, and vLLM-de­rived over­lays are Apache-2.0 (see LICENSE); the AITER-derived over­lay keeps its MIT header. Upstream base re­vi­sions for every diff are recorded in patches/​README.md. The model it­self is MIT-licensed.

References

All links ver­i­fied 2026 – 08-04.

DeepSeek-V4-Flash-0731 model card — of­fi­cial re­lease; 304B pa­ra­me­ters; fused DSpark mod­ule; rec­om­mended tem­per­a­ture=1.0, top_p=0.95; MIT li­cense

Official vLLM DeepSeek V4 Flash recipe — ref­er­ence launch con­fig­u­ra­tion, DSpark (num_speculative_tokens=7), FP8 KV, block size 256, deepseek_v4 parsers; AMD guid­ance for MI325X/MI355X

Bringing up DeepSeek-V4-Flash on AMD MI300X (Fergus Finn, Doubleword, June 2026) — the bring-up work­log this repo builds on: FNUZ vs. OCP FP8, AITER gaps on gfx942, HIP-graph haz­ards, rout­ing bugs

dou­ble­wor­dai/​vllm-amd-blog-dou­ble­word — demo PRs for the above, in­clud­ing com­mit c32932bb9 (“mask MXFP4 bit­ma­trix padding lanes by log­i­cal block size”)

vLLM com­mit 77469c9 — [ROCm][MLA] Mask the AITER MLA small-head ver­ify flat­ten causally (#50476)”

vLLM is­sue #47282 — CPU-KV load path lacks cross-stream sync with com­pute (WAR gap)

vLLM PR #47291 — pro­posed WAR fix, not merged; car­ried as an over­lay here

AMD Instinct MI300X — 192 GB HBM3, 5.3 TB/s peak band­width, 2.61 PFLOPS peak FP8

ROCm/AITER — AMD tuned-ker­nel li­brary used for ROCm at­ten­tion and dense lin­ears

vLLM — the serv­ing run­time (ROCm nightlies un­der vllm/​vllm-ope­nai-rocm)

Smaller, faster, safer: running Kimi and GLM at scale

blog.cloudflare.com

Workers AI runs in­fer­ence for some of the best open mod­els in the world on GPUs in Cloudflare data cen­ters close to your users. Two of the most ca­pa­ble, and most de­mand­ing, are Moonshot’s Kimi K-series and Z.ai’s GLM. They are large, long-con­text, mix­ture-of-ex­perts mod­els, and they are won­der­ful to use. They are also very hard to serve ef­fi­ciently be­cause of mem­ory con­straints.

We’ve writ­ten be­fore about how we serve large mod­els on Workers AI and about sep­a­rat­ing the pre­fill and de­code phases of in­fer­ence to get more out of each GPU. This post looks at three tech­niques we layer on top of that to fit these mod­els into mem­ory and keep them fast: quan­tiz­ing the KV cache, com­press­ing the model weights, and, be­cause both of those pack more re­quests onto shared hard­ware, pro­tect­ing the cache those re­quests share. These op­ti­miza­tions en­able us to sup­port more cus­tomers at lower costs, with no change in model ac­cu­racy.

All our ex­per­i­ments and pro­duc­tion traf­fic are run­ning and bench­marked with SGLang, an open-source in­fer­ence serv­ing frame­work. We found that SGLang of­fers the best per­for­mance in the mar­ket, and we work closely with the SGLang team to up­stream patches and new fea­tures to make our work avail­able to the open-source com­mu­nity.

Quantizing the KV cache

As a model gen­er­ates text, it stores the at­ten­tion keys (K) and val­ues (V) for every to­ken it has al­ready processed in a struc­ture called the KV cache. The cache is what lets the model ex­tend a long con­ver­sa­tion with­out re-read­ing the en­tire con­text on every new to­ken. For a long-con­text model, it grows quickly, and it is usu­ally the KV cache, not the mod­el’s weights, that fills up GPU mem­ory first.

By de­fault, the cache is stored in 16-bit pre­ci­sion (BF16). We store it in 8-bit float­ing point in­stead (FP8, e4m3), which halves its size. On Kimi K2.6, that raises the amount of con­text we can hold in mem­ory from roughly 686,000 to­kens to about 1.37 mil­lion, twice as much.

It’s worth be­ing pre­cise about where the ben­e­fit comes from, be­cause it is­n’t raw speed. Quantizing the cache adds a small amount of work per to­ken, since the FP8 at­ten­tion ker­nel has to con­vert val­ues as it reads them. What it changes is how many re­quests we can keep res­i­dent at once. The fol­low­ing mea­sure­ments are for Kimi K2.6 de­cod­ing on a dis­ag­gre­gated H200 de­ploy­ment, com­par­ing the at­ten­tion ker­nels di­rectly:

Concurrent re­quests

BF16 KV cache (tok/s)

FP8 KV cache (tok/s)

1

137

125

8

731

689

16

1,106

1,028

32

1,558

1,489

64

Out of mem­ory

2,192

At any sin­gle con­cur­rency level, BF16 is a few per­cent faster per to­ken. But BF16 runs out of cache at 32 con­cur­rent re­quests and can’t ad­mit a 33rd, while FP8 keeps go­ing to 64 and reaches 2,192 to­kens per sec­ond, about 41% higher than BF16′s peak, for roughly 30% less cost per to­ken. Because we run pre­fill and de­code as sep­a­rate pools, we can ap­ply this where it helps most: pre­fill is com­pute-bound rather than mem­ory-bound, so there we leave the cache in BF16 and keep its slightly higher through­put.

None of this would mat­ter if it changed the mod­el’s an­swers, so we checked. Across our eval­u­a­tion suite, FP8 and BF16 caches are in­dis­tin­guish­able:

Benchmark

BF16 KV

FP8 KV

GSM8K

94.24

94.09

ARC-Easy

89.06

89.14

ARC-Challenge

66.72

67.49

MMLU

89.11

89.04

MMLU-Pro

80.29

79.29

mcx­ams (internal bench­mark)

61 / 63

61 / 63

Tool-call va­lid­ity

92.2%

92.6%

Compressing the model weights

The KV cache is one de­mand on GPU mem­ory; the mod­el’s weights are the other. For GLM 5.2, we com­press the weights from 8-bit float­ing point down to 4-bit in­te­gers (INT4) with no loss in ac­cu­racy. The check­point shrinks from 705 GB to 421 GB, about 40%, and per-GPU mem­ory across an 8-way ten­sor-par­al­lel de­ploy­ment drops from roughly 88 GB to 52 GB, which leaves room for around 1.18 mil­lion to­kens of KV cache on the same hard­ware.

Across our eval­u­a­tion suite, INT4 and FP8 weights are in­dis­tin­guish­able:

Benchmark / Capability

Metric

FP8

INT4

GSM8K

Exact match

94.39%

93.56%

GSM8K

Flexible

94.24%

93.48%

ARC-Easy

Accuracy

86.62%

86.15%

ARC-Easy

Acc (norm)

84.51%

85.19%

ARC-Challenge

Accuracy

64.93%

64.85%

ARC-Challenge

Acc (norm)

67.24%

66.64%

MMLU

Average

86.60%

86.54%

MMLU-Pro

Exact

80.80%

80.47%

mcx­ams (internal bench­mark)

Passed

62 / 63

62 / 63

Smaller weights make the de­code phase faster, and for a clear rea­son: gen­er­at­ing each to­ken means stream­ing the mod­el’s weights out of GPU mem­ory, so de­code speed is lim­ited by mem­ory band­width. Move less data and every to­ken ar­rives sooner. The ef­fect is largest at low con­cur­rency, where per-re­quest la­tency mat­ters most:

Concurrent re­quests

GLM FP8 (tok/s)

GLM INT4 (tok/s)

INT4 gain

1

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.