10 interesting stories served every morning and every evening.

Initiative detail | European Citizens' Initiative

citizens-initiative.europa.eu

European Citizens’ Initiative

Substack writers, you need a website!

elizabethtai.com

But I al­ready have a web­site on Substack,” you ar­gue.

No, no, Substack is just a dis­tri­b­u­tion tool to am­plify your web­site. It should not be your dig­i­tal home.

In the last few years, I’ve no­ticed a pat­tern of writ­ers leav­ing their web­sites to make Substack their dig­i­tal home.

Now, it’s kinda okay if they have bought a do­main and linked it to Substack. (Meaning, it’s bet­ter than noth­ing.)

Rachel from Conscious Living is a good ex­am­ple. This way, Substack more or less func­tions like a con­tent man­age­ment sys­tem (CMS) for you.

However, com­pared to other CMS it’s very lim­ited, such as the abil­ity to man­age your SEO and cus­tomize your pages to add more fea­tures, but I di­gress. If you just want a fuss-free plat­form, this is one way to get it and Substack’s con­di­tions for do­mains are very rea­son­able and cost-ef­fi­cient. As I will ex­plain later, this could change on a dime with­out warn­ing.

However, there are some writ­ers who are say­ing: Hey read­ers, I’m now writ­ing on Substack, so head on over there (and ig­nore my web­site)!”

Some writ­ers do have a web­site, but link to their Substacks, call­ing them their blogs”. If your Substack has a do­main name they own, it’s okay, but if it’s xx.sub­stack.com, Substack is say­ing All your con­tent are be­long to us”.

In con­clu­sion: Writers, don’t do this. It’s short-sighted and un­wise and can de­rail your long-term vis­i­bil­ity on the Internet.

The siren call of con­ve­nience

Every few years, the in­ter­net con­vinces writ­ers that a new dig­i­tal par­adise has ar­rived. First, it was so­cial me­dia like Facebook. Then blog­ging net­works like Tumblr. Then it was Medium. More re­cently, it’s been Substack.

Platforms promise us an ea­ger au­di­ence, built-in mon­e­ti­za­tion, a smooth user in­ter­face, and a sup­port­ive com­mu­nity. As a writer who just wants to fo­cus on writ­ing, it’s in­cred­i­bly tempt­ing to hand over the keys to our cre­ative king­doms and let these por­tals han­dle every­thing. (Believe me, I gave in at one point. For years, I just stopped blog­ging al­to­gether and even gave up a do­main that had high traf­fic! But I got back in 2012 and never left.)

However, this is the truth that has not changed since the dawn of the Internet: When you build your au­di­ence en­tirely on some­one else’s plat­form, you aren’t a home­owner. You are a ten­ant. Or worse, a dig­i­tal share­crop­per.

And cor­po­rate land­lords al­ways change the rules even­tu­ally. It’s not per­sonal, it’s just busi­ness.

The il­lu­sion of the safe space

It’s easy to feel se­cure when a plat­form is in its golden era. But we’ve watched the down­falls of Twitter, the pol­icy shifts of Reddit, and the chang­ing tides of al­go­rith­mic net­works. Relying blindly on a cen­tral­ized por­tal not owned by you means your life’s work can al­ter overnight based en­tirely on a cor­po­rate board­room de­ci­sion.

When I looked at how frag­ile our dig­i­tal ecosys­tems re­ally are, I re­al­ized I needed a space that would­n’t go poof” be­cause a com­pany needed to please its in­vestors or share­hold­ers. This re­al­iza­tion com­pletely changed my ap­proach, push­ing me to pro­tect my con­tent by learn­ing to blog the IndieWeb way.

Your writ­ing needs a per­ma­nent home­base—a do­main that you own and con­trol. Full stop.

Moving from renting” to syn­di­cat­ing

The biggest push­back I hear from writ­ers is: But my web­site does­n’t have an au­di­ence! Substack does.”

But you don’t have to com­pletely aban­don so­cial me­dia or plat­forms like Substack to pro­tect your au­ton­omy (personally, I pre­fer the word sov­er­eignty but it does sound a tad dra­matic).

You just need to change the or­der of op­er­a­tions. Instead of pub­lish­ing di­rectly to a por­tal, you can shift your mind­set to POSSE: Publish (on your) Own Site, Syndicate Elsewhere. (I ex­plain the POSSE/PESOS method in an older post.)

By treat­ing your web­site as the de­fin­i­tive source of truth and us­ing plat­forms sim­ply as dis­tri­b­u­tion pipes, you get the best of both worlds. I dug deep into this shift when I com­mit­ted to be­ing an im­per­fect gar­dener of my dig­i­tal gar­den, ex­plor­ing how a less mar­ket-y way of pre­sent­ing my con­tent on­line let me share my wild gar­den of thoughts with­out danc­ing to the al­go­rithm.

A re­al­ity check on plat­form hype

If you are still hold­ing out hope that Substack is different” from the so­cial plat­forms that came be­fore it, let’s look at the num­bers and be­hav­iors be­hind the mar­ket­ing copy.

After spend­ing a sig­nif­i­cant amount of time ob­serv­ing the plat­form ecosys­tem first­hand, I wrote a bru­tally hon­est take­away in What I learned from one year of Substack. The net­work ef­fects are real, but so is the pres­sure to con­form to what the plat­for­m’s ecosys­tem fa­vors.

This post, by the way, des­per­ately needs to be up­dated be­cause things have got­ten much, much worse since I wrote it.

When you hand your con­tent over to a plat­form, you have to con­form to their rules and their lo­cal­ized bi­ases. For those of us writ­ing from out­side the dom­i­nant US-centric echo cham­bers, plat­form al­go­rithms heav­ily pri­or­i­tize spe­cific west­ern nar­ra­tives, mak­ing it in­cred­i­bly tough for lo­cal­ized or mi­nor­ity voices to be seen un­less they con­form.

I wrote about this ex­act frus­tra­tion re­cently in Linkblog: Dwelling on the Internet, high­light­ing how al­go­rith­mic com­pla­cency forces us into ho­mog­e­nized bub­bles.

The flip side — the writ­ers who re­fused to leave their web­sites

Each time there’s a new drama on some plat­form, and writ­ers are shak­ing their sabers and de­clar­ing that they will leave for yet an­other so­cial me­dia plat­form they don’t con­trol, I think about writ­ers like John Scalzi.

As of date, John scalzi has been blog­ging on https://​what­ever.scalzi.com/ for 28 years!

This sci-fi nov­el­ist has main­tained a sin­gle in­de­pen­dent web­site con­tin­u­ously for nearly three decades; this makes him one of the longest-run­ning, most con­sis­tent orig­i­nal blog­gers on the in­ter­net. Imagine the amount of dig­i­tal foot­print on that web­site! Unbroken by time or plat­forms.

(Specifically, he uses word­press.com like I do, as we both don’t want to bother with the pain of set­ting up your own self-hosted word­press web­site and just want the folks at Automattic to do it.)

He blogs in the clas­sic Indieweb way, though I doubt he is even aware he’s do­ing it. He treats his so­cial me­dia chan­nels such as X or Bluesky as a way to am­plify his web­site. All roads lead back to https://​what­ever.scalzi.com/, and this is some­thing I wish every sin­gle writer would do.

He wrote re­cently in Various & Sundry, 6/3/26:

this site acts as my own in­sti­tu­tional mem­ory, if I post some­thing about it here it con­sti­tutes an of­fi­cial record. I mean, all the posts I ever placed on the for­mer Twitter are now en­tirely lost to time, since I have gone in and purged my en­tire time­line there. This site, how­ever, en­dures. — John Scalzi

this site acts as my own in­sti­tu­tional mem­ory, if I post some­thing about it here it con­sti­tutes an of­fi­cial record. I mean, all the posts I ever placed on the for­mer Twitter are now en­tirely lost to time, since I have gone in and purged my en­tire time­line there. This site, how­ever, en­dures. — John Scalzi

Breaking free from plat­form blues

Trying to adapt your pres­ence across var­i­ous plat­forms in an ever-shift­ing dig­i­tal land­scape is ex­haust­ing. One minute a plat­form is a writer’s dar­ling; the next, it’s be­ing boy­cotted. Railing against a plat­for­m’s fo­cus shift or the pres­ence of (long sigh) Nazis is a use­less en­deavor.

As I noted in Linkblog March 12, 2026: Platform blues, chas­ing plat­form pu­rity is an il­lu­sion. Tech will change, cor­po­rate al­go­rithms will con­tinue to pri­or­i­tize profit over hu­man con­nec­tion, and plat­forms will con­tinue to cy­cle through hype and de­cline.

The an­ti­dote to this ex­haus­tion is­n’t mov­ing to the next shiny new app. It’s an­chor­ing your work on an in­de­pen­dent web­site with open dis­tri­b­u­tion chan­nels like RSS. It also means ruth­lessly us­ing plat­forms as dis­tri­b­u­tion chan­nels. When one col­lapses or you pre­fer to just move, it’s easy to just change strate­gies be­cause your dig­i­tal home re­mains un­changed.

Use plat­forms to find your read­ers, but bring them back to your house. It’s time to stop dig­i­tal share­crop­ping on rented land.

Featured photo is by vivek vk on Unsplash

GitHub - openai/codex-security: SDKs and CLI for Codex Security

github.com

@openai/codex-security is a CLI and TypeScript SDK for find­ing, val­i­dat­ing, and fix­ing se­cu­rity vul­ner­a­bil­i­ties in your code. Scan repos­i­to­ries, re­view changes, track find­ings over time, and run se­cu­rity checks in CI.

Documentation

Quick start

Requires Node.js 22 or later, Python 3.10 or later, and ac­cess to Codex Security.

npm in­stall @openai/codex-security npx codex-se­cu­rity lo­gin npx codex-se­cu­rity scan .

For CI, set OPENAI_API_KEY in­stead of sign­ing in.

If both a ChatGPT sign-in and an API key are avail­able, in­ter­ac­tive scans ask which cre­den­tial to use. CI and other non­in­ter­ac­tive scans keep the ex­ist­ing API-key prece­dence. Select a cre­den­tial ex­plic­itly when needed:

npx codex-se­cu­rity scan . –auth chat­gpt npx codex-se­cu­rity scan . –auth api-key

To make your ChatGPT sign-in the au­to­matic de­fault, un­set any con­fig­ured API keys:

un­set OPENAI_API_KEY CODEX_API_KEY

Scan his­tory is stored in the Codex Security work­bench state di­rec­tory. If that di­rec­tory can­not be writ­ten, set CODEX_SECURITY_STATE_DIR to a writable di­rec­tory out­side the repos­i­tory.

TypeScript SDK

im­port { CodexSecurity } from @openai/codex-security”;

const se­cu­rity = new CodexSecurity(); const re­sult = await se­cu­rity.run(”.“);

con­sole.log(re­sult.re­port­Path); await se­cu­rity.close();

For in­stal­la­tion, au­then­ti­ca­tion, scan op­tions, and CI setup, see the of­fi­cial doc­u­men­ta­tion.

Kimi K3 Architecture Notes

sebastianraschka.com

The Kimi K3 ar­chi­tec­ture fig­ure for yes­ter­day’s big open-weight model re­lease, along with some ob­ser­va­tions and thoughts.

Yes, it looks rel­a­tively com­pli­cated, but it’s es­sen­tially a scaled-up pro­duc­tion ver­sion of their Kimi Linear model they re­leased last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)

Yes, it looks rel­a­tively com­pli­cated, but it’s es­sen­tially a scaled-up pro­duc­tion ver­sion of their Kimi Linear model they re­leased last year (scaled up from 48B -> 2.8T; K3 is by far the biggest open-weight model right now)

The one new com­po­nent com­pared to Kimi Linear is the LatentMoE. I omit­ted it in the fig­ure be­low since it’s al­ready very crowded, but that’s es­sen­tially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are cu­ri­ous). The idea here is to com­press (down-project) large lin­ear lay­ers sim­i­lar to multi-head la­tent at­ten­tion.

The one new com­po­nent com­pared to Kimi Linear is the LatentMoE. I omit­ted it in the fig­ure be­low since it’s al­ready very crowded, but that’s es­sen­tially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are cu­ri­ous). The idea here is to com­press (down-project) large lin­ear lay­ers sim­i­lar to multi-head la­tent at­ten­tion.

Kimi K3s over­all trend (similar to Nemotron 3, DeepSeek V4, and oth­ers) is also to­wards bet­ter in­fer­ence ef­fi­ciency. That is, there are many com­po­nents that re­place ex­ist­ing com­po­nents with ef­fi­ciency-tweaked ver­sions. I.e., MoE -> LatentMoE, reg­u­lar at­ten­tion -> multi-head la­tent at­ten­tion and Kimi Delta Attention. (I also have short tu­to­ri­als and write-ups in my gallery if you are cu­ri­ous about ad­di­tional de­tails).

Kimi K3s over­all trend (similar to Nemotron 3, DeepSeek V4, and oth­ers) is also to­wards bet­ter in­fer­ence ef­fi­ciency. That is, there are many com­po­nents that re­place ex­ist­ing com­po­nents with ef­fi­ciency-tweaked ver­sions. I.e., MoE -> LatentMoE, reg­u­lar at­ten­tion -> multi-head la­tent at­ten­tion and Kimi Delta Attention. (I also have short tu­to­ri­als and write-ups in my gallery if you are cu­ri­ous about ad­di­tional de­tails).

The one com­po­nent change that is not an ef­fi­ciency tweak is at­ten­tion resid­u­als. Like DeepSeek V4 im­proved the resid­ual path with mHC (manifold-constrained Hyper-Connections), at­ten­tion resid­u­als are a way to im­prove the resid­ual path, but it works a bit dif­fer­ently. I.e., mHC made the resid­ual path wider. Attention resid­u­als (also al­ready part of Kimi Linear) con­nect the resid­u­als across lay­ers; the con­nec­tion it­self uses an at­ten­tion score for an im­por­tant/​con­tri­bu­tion weight. According to the re­port, it im­proves the val­i­da­tion loss and down­stream per­for­mance (a bit) con­sis­tently and adds about 4% in train­ing cost and 2% in in­fer­ence cost.

The one com­po­nent change that is not an ef­fi­ciency tweak is at­ten­tion resid­u­als. Like DeepSeek V4 im­proved the resid­ual path with mHC (manifold-constrained Hyper-Connections), at­ten­tion resid­u­als are a way to im­prove the resid­ual path, but it works a bit dif­fer­ently. I.e., mHC made the resid­ual path wider. Attention resid­u­als (also al­ready part of Kimi Linear) con­nect the resid­u­als across lay­ers; the con­nec­tion it­self uses an at­ten­tion score for an im­por­tant/​con­tri­bu­tion weight. According to the re­port, it im­proves the val­i­da­tion loss and down­stream per­for­mance (a bit) con­sis­tently and adds about 4% in train­ing cost and 2% in in­fer­ence cost.

Interestingly, Kimi K3 got rid of all RoPE lay­ers and uses NoPE (No Positional Embeddings) every­where in­stead. (Again, this is in­her­ited from Kimi Linear). In other ar­chi­tec­tures, the re­cent trend was to­wards RoPE in lo­cal at­ten­tion lay­ers (like slid­ing win­dow at­ten­tion) and NoPE in the global lay­ers. There were a few ar­chi­tec­tures that only used NoPE every­where, but this is the first fron­tier-level one as far as I know.

Interestingly, Kimi K3 got rid of all RoPE lay­ers and uses NoPE (No Positional Embeddings) every­where in­stead. (Again, this is in­her­ited from Kimi Linear). In other ar­chi­tec­tures, the re­cent trend was to­wards RoPE in lo­cal at­ten­tion lay­ers (like slid­ing win­dow at­ten­tion) and NoPE in the global lay­ers. There were a few ar­chi­tec­tures that only used NoPE every­where, but this is the first fron­tier-level one as far as I know.

Kimi K3 now also has na­tive mul­ti­modal sup­port, which is great!

Kimi K3 now also has na­tive mul­ti­modal sup­port, which is great!

There are sev­eral other in­ter­est­ing train­ing tid­bits in the tech­ni­cal re­port, but that’s it from the ar­chi­tec­ture front so far. A re­ally great re­lease over­all.

Source: web­site ver­sion of my Substack note.

Read Next

A Few Notable Open-Weight Models This Week

Short note on the ar­chi­tec­tures of six new open-weight mod­els, in­clud­ing Nanbeige 4.2, Laguna S 2.1, Motif-3-Beta, Solar Open 2, Antares 1B, and BTL-3.

Correction for Listing 6.5 in Build a Reasoning Model From Scratch

Short cor­rec­tion note for the ran­dom seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch.

Inkling: A New Open-Weight 975B MoE with a Few Surprises

Short note on Thinking Machines Lab’s 975B Inkling model, in­clud­ing bench­marks, sparse MoE de­sign, short con­vo­lu­tions, RMSNorm, and po­si­tion bias.

GitHub - twalichiewicz/HNewhere: A lightweight userscript that adds Hacker News discussions to any article.

github.com

A light­weight user­script that adds Hacker News dis­cus­sions to any ar­ti­cle.

HNewhere de­tects Hacker News sto­ries, loads com­ments into a side­bar, and lets you browse dis­cus­sions with­out leav­ing the page.

Install

Install HNewhere

Requires a user­script man­ager such as Tampermonkey, Violentmonkey, or Userscripts.

Features

Opens HN dis­cus­sions be­side ar­ti­cles

Automatically de­tects match­ing Hacker News sto­ries

Tracks links opened from Hacker News

Resizable side­bar

Collapsible com­ments

Preserves side­bar width

Reply links open di­rectly to HN

Install

Install a user­script man­ager:

Userscripts (Safari) Tampermonkey Violentmonkey

Install a user­script man­ager:

Userscripts (Safari)

Tampermonkey

Violentmonkey

Install or paste HNewhere.user.js

Install or paste HNewhere.user.js

Visit an ar­ti­cle with a Hacker News dis­cus­sion.

Visit an ar­ti­cle with a Hacker News dis­cus­sion.

Requirements

Browser with user­script sup­port

Access to:

Hacker News API HN Algolia search API

Hacker News API

HN Algolia search API

License

MIT

You Could Have Come Up With Kimi Delta Attention | Doubleword

blog.doubleword.ai

A note on no­ta­tion: this ar­ti­cle de­faults to bra-ket no­ta­tion be­cause (in my quan­tum-in­spired opin­ion) it makes the shapes in this de­riva­tion very clear. The Math no­ta­tion switch above rewrites every equa­tion us­ing con­ven­tional bold vec­tors and ex­plicit trans­poses in­stead. In bra-ket mode, ∣q⟩\lvert q\ran­gle is a col­umn vec­tor, ⟨k∣\langle k\rvert is a row vec­tor, ⟨k∣q⟩\langle k\rvert q\ran­gle is a num­ber, and ∣v⟩⟨k∣\lvert v\ran­gle\lan­gle k\rvert is a ma­trix. Vectors face right by de­fault, while keys face left when writ­ten into the lin­ear-at­ten­tion state. We work with one causal at­ten­tion head and real-val­ued vec­tors, as­sume DeltaNet’s keys are nor­mal­ized, and let the state map from key space to value space.

Modern lin­ear at­ten­tion vari­ants are com­plex, and a upon first glance it is not so easy to see what they are de­signed to achieve. For ref­er­ence here is the state up­date equa­tion for Kimi Delta Attention (KDA):

S~t=St−1Diag⁡(αt)\widetilde S_t = S_{t-1}\operatorname{Diag}(\alpha_t) ∣v^t⟩=S~t∣kt⟩\lvert\widehat v_t\ran­gle = \widetilde S_t\lvert k_t\ran­gle ∣et⟩=βt(∣vt⟩−∣v^t⟩)\lvert e_t\ran­gle = \beta_t \left( \lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle \right) St=S~t+∣et⟩⟨kt∣S_t = \widetilde S_t+\lvert e_t\ran­gle\lan­gle k_t\rvert ∣ot⟩=St(dk−1/2∣qt⟩)\lvert o_t\ran­gle = S_t\left(d_k^{-1/2}\lvert q_t\ran­gle\right)

The rea­son they are so dif­fi­cult to un­der­stand is that this is the lat­est in a fam­ily of lin­ear at­ten­tion vari­ants that have been de­vel­oped over the last few years and the com­plex­ity of them has in­evitably bal­looned such that from the out­side the lat­est vari­ants ap­pear in­ac­ces­si­ble.

In this post we are go­ing to walk through the DeltaNet fam­ily of lin­ear at­ten­tion vari­ants, two of which are used by the lat­est Qwen and Kimi model fam­i­lies, and show how you might have ar­rived at the same equa­tions by as­sert­ing sim­ple things about your hid­den state.

That is the route we will take:

soft­max at­ten­tion → lin­ear at­ten­tion → DeltaNet → Gated DeltaNet → KDA

Only af­ter de­riv­ing KDA will we turn to the re­cur­rent and chunk­wise Triton pro­grams that ex­e­cute it.

1. Begin with qua­dratic at­ten­tion

For a query at to­ken tt, or­di­nary causal soft­max at­ten­tion is

ati=exp⁡ ⁣(s⟨ki∣qt⟩)∑j≤texp⁡ ⁣(s⟨kj∣qt⟩),s=dk−1/​2,∣ot⟩=∑i≤tati∣vi⟩.\be­gin{aligned} a_{ti} &= \frac{ \exp\!\left(s\langle k_i\rvert q_t\ran­gle\right) }{ \sum_{j\leq t} \exp\!\left(s\langle k_j\rvert q_t\ran­gle\right) }, \qquad s=d_k^{-1/​2},\\ \lvert o_t\ran­gle &= \sum_{i\leq t}a_{ti}\lvert v_i\ran­gle. \end{aligned}

Every at­ten­tion weight is a scalar. It mea­sures the sim­i­lar­ity be­tween one key and one query, then soft­max turns all of the scores for that query into a dis­tri­b­u­tion. The out­put is a weighted sum of value vec­tors.

Over a se­quence of length TT, there are T2T^2 key-query pairs. During au­tore­gres­sive in­fer­ence we can cache the keys and val­ues in­stead of re­com­put­ing them, but the cache still grows with the se­quence and every new query still has to in­spect the en­tire his­tory.

The ob­sta­cle to re­ar­rang­ing this com­pu­ta­tion is the soft­max. Its de­nom­i­na­tor de­pends jointly on the cur­rent query and every ear­lier key. So, for the mo­ment, re­move it.

1.1 Remove the soft­max

For clar­ity, ab­sorb the con­stant scale ss into the query. The de­lib­er­ately bare ver­sion of at­ten­tion is then

∣ot⟩=∑i≤t⟨ki∣qt⟩∣vi⟩.\lvert o_t\ran­gle = \sum_{i\leq t} \langle k_i\rvert q_t\ran­gle \lvert v_i\ran­gle.

The scalar in­ner prod­uct can move to the right:

∣ot⟩=∑i≤t∣vi⟩⟨ki∣qt⟩=(∑i≤t∣vi⟩⟨ki∣)∣qt⟩.\begin{aligned} \lvert o_t\ran­gle &= \sum_{i\leq t} \lvert v_i\ran­gle \langle k_i\rvert q_t\ran­gle\\ &= \left( \sum_{i\leq t} \lvert v_i\ran­gle\lan­gle k_i\rvert \right) \lvert q_t\ran­gle. \end{aligned}

Everything that de­pends on the past can now be col­lected into one ma­trix of a fixed size V×KV \times K:

St=∑i≤t∣vi⟩⟨ki∣\boxed{ S_t = \sum_{i\leq t} \lvert v_i\ran­gle\lan­gle k_i\rvert }

and at­ten­tion be­comes a re­cur­rent write fol­lowed by a read:

St=St−1+∣vt⟩⟨kt∣,∣ot⟩=St∣qt⟩.\boxed{ \begin{aligned} S_t &= S_{t-1} + \lvert v_t\ran­gle\lan­gle k_t\rvert,\\ \lvert o_t\ran­gle &= S_t\lvert q_t\ran­gle. \end{aligned} }

The iden­tity

(∣v⟩⟨k∣)∣q⟩=⟨k∣q⟩∣v⟩\left(\lvert v\ran­gle\lan­gle k\rvert\right)\lvert q\ran­gle = \langle k\rvert q\ran­gle\lvert v\ran­gle

is the whole trick. The outer prod­uct is a ma­trix; the in­ner prod­uct is a num­ber. We no longer store every past key and value. We store their summed outer prod­ucts in the fixed-size state StS_t.

This is lin­ear in se­quence length rather than qua­dratic: scan the to­kens once, up­dat­ing the same dv×dkd_v\times d_k state at every step. We have paid for that ef­fi­ciency by dis­card­ing soft­max’s nor­mal­iza­tion and se­lec­tiv­ity. More so­phis­ti­cated lin­ear-at­ten­tion meth­ods use fea­ture maps and nor­mal­iz­ers, but this un­adorned form ex­poses the mem­ory prob­lem that mo­ti­vates DeltaNet.

1.2 Addition is not as­sign­ment

Suppose we write a pair ∣vt⟩⟨kt∣\lvert v_t\ran­gle\lan­gle k_t\rvert and im­me­di­ately query the new state with that same key:

St∣kt⟩=(St−1+∣vt⟩⟨kt∣)∣kt⟩=St−1∣kt⟩+∣vt⟩⟨kt∣kt⟩⏟1=St−1∣kt⟩+∣vt⟩.\begin{aligned} S_t\lvert k_t\ran­gle &= \left( S_{t-1} + \lvert v_t\ran­gle\lan­gle k_t\rvert \right) \lvert k_t\ran­gle\\ &= S_{t-1}\lvert k_t\ran­gle + \lvert v_t\ran­gle \underbrace{\langle k_t\rvert k_t\ran­gle}_{1}\\ &= S_{t-1}\lvert k_t\ran­gle+\lvert v_t\ran­gle. \end{aligned}

The write does not make the mem­ory re­turn ∣vt⟩\lvert v_t\ran­gle. It adds ∣vt⟩\lvert v_t\ran­gle to what­ever the mem­ory al­ready re­turned.

If the old state al­ready pro­duced the cor­rect value, the ad­di­tive write makes the new state pro­duce twice that value. More gen­er­ally, keys are not mu­tu­ally or­thog­o­nal, so every write can in­ter­fere with pre­vi­ous writes. Linear at­ten­tion has given us a com­pact as­so­cia­tive mem­ory, but its up­date be­haves like += when what we want is closer to =.

2. DeltaNet: write the er­ror, not the value

DeltaNet re­places the un­con­di­tional lin­ear-at­ten­tion write with a delta-rule cor­rec­tion. There are two use­ful ways to de­rive it.

2.1 Derivation one: de­mand that the write can be read back

Before writ­ing to­ken tt, ask the mem­ory what it cur­rently as­so­ci­ates with the new key:

∣v^t⟩=St−1∣kt⟩.\lvert\widehat v_t\ran­gle = S_{t-1}\lvert k_t\ran­gle.

If we want the mem­ory to re­turn ∣vt⟩\lvert v_t\ran­gle, we should not add the whole value. We should add only the dif­fer­ence:

∣vt⟩−∣v^t⟩.\lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle.

Introduce a learned write strength βt∈[0,1]\be­ta_t\in[0,1] and de­fine

∣et⟩=βt(∣vt⟩−St−1∣kt⟩).\lvert e_t\ran­gle = \beta_t \left( \lvert v_t\ran­gle - S_{t-1}\lvert k_t\ran­gle \right).

Then write this er­ror at the cur­rent key:

St=St−1+∣et⟩⟨kt∣.\boxed{ S_t = S_{t-1} + \lvert e_t\ran­gle\lan­gle k_t\rvert. }

Now im­me­di­ately read the same key:

St∣kt⟩=St−1∣kt⟩+∣et⟩⟨kt∣kt⟩=(1−βt)St−1∣kt⟩+βt∣vt⟩.\begin{aligned} S_t\lvert k_t\ran­gle &= S_{t-1}\lvert k_t\ran­gle + \lvert e_t\ran­gle \langle k_t\rvert k_t\ran­gle\\ &= (1-\beta_t)S_{t-1}\lvert k_t\ran­gle + \beta_t\lvert v_t\ran­gle. \end{aligned}

When βt=1\be­ta_t=1, the re­sult is ex­actly ∣vt⟩\lvert v_t\ran­gle. Smaller βt\be­ta_t moves the old pre­dic­tion part­way to­wards the tar­get.

The cor­rec­tion is also lo­cal in key space. For any query ∣x⟩\lvert x\ran­gle or­thog­o­nal to the cur­rent key,

⟨kt∣x⟩=0⟹(St−St−1)∣x⟩=∣et⟩⟨kt∣x⟩⏟0=0.\langle k_t\rvert x\ran­gle=0 \quad\Longrightarrow\quad (S_t-S_{t-1})\lvert x\ran­gle = \lvert e_t\ran­gle \underbrace{\langle k_t\rvert x\ran­gle}_{0} =0.

So the rank-one write changes the re­sponse in the se­lected key di­rec­tion while leav­ing every or­thog­o­nal di­rec­tion alone.

2.2 Derivation two: take one step on re­con­struc­tion loss

The same up­date falls out of an on­line learn­ing ob­jec­tive. Treat the cur­rent key-value pair as one train­ing ex­am­ple for the lin­ear map SS:

Lt(S)=12∥S∣kt⟩−∣vt⟩∥22.\mathcal L_t(S) = \frac12 \left\| S\lvert k_t\ran­gle-\lvert v_t\ran­gle \right\|_2^2.

Its gra­di­ent with re­spect to the state is

∇SLt(S)=(S∣kt⟩−∣vt⟩)⟨kt∣.\nabla_S\mathcal L_t(S) = \left( S\lvert k_t\ran­gle-\lvert v_t\ran­gle \right) \langle k_t\rvert.

This is vis­i­bly an outer prod­uct: a value-space pre­dic­tion er­ror times the key bra at which that er­ror was ob­served. Take one gra­di­ent-de­scent step of size βt\be­ta_t from St−1S_{t-1}:

St=St−1−βt∇SLt(St−1)=St−1−βt(St−1∣kt⟩−∣vt⟩)⟨kt∣=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣.\begin{aligned} S_t &= S_{t-1} - \beta_t\nabla_S\mathcal L_t(S_{t-1})\\ &= S_{t-1} - \beta_t \left( S_{t-1}\lvert k_t\ran­gle-\lvert v_t\ran­gle \right) \langle k_t\rvert\\ &= S_{t-1} + \beta_t \left( \lvert v_t\ran­gle-S_{t-1}\lvert k_t\ran­gle \right) \langle k_t\rvert. \end{aligned}

This is ex­actly the up­date we got by re­quir­ing im­me­di­ate re­con­struc­tion. The two in­ter­pre­ta­tions are the same:

as a mem­ory op­er­a­tion, βt\be­ta_t con­trols how strongly to re­place the old as­so­ci­a­tion;

as on­line learn­ing, βt\be­ta_t is the step size;

as lin­ear al­ge­bra, the change is a rank-one outer prod­uct.

2.3 The DeltaNet state tran­si­tion

Expanding the er­ror ex­poses DeltaNet as a struc­tured state tran­si­tion plus a new in­put:

St=St−1+βt(∣vt⟩−St−1∣kt⟩)⟨kt∣=St−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1} + \beta_t \left( \lvert v_t\ran­gle-S_{t-1}\lvert k_t\ran­gle \right) \langle k_t\rvert\\ &= S_{t-1} \left( I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert \right) + \beta_t\lvert v_t\ran­gle\lan­gle k_t\rvert. \end{aligned}

For a unit key, I−βt∣kt⟩⟨kt∣I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert has eigen­value 1−βt1-\beta_t in the cur­rent key di­rec­tion and eigen­value 11 in every or­thog­o­nal di­rec­tion. It re­moves the old as­so­ci­a­tion along the cur­rent key be­fore adding the new one.

DeltaNet fixes the write. It does not yet fix the life­time of the state.

3. Gated DeltaNet: some­times old in­for­ma­tion should dis­ap­pear

The lin­ear state com­presses the whole his­tory into one ma­trix. A read

St∣q⟩=∑i≤t⟨ki∣q⟩∣vi⟩S_t\lvert q\ran­gle = \sum_{i\leq t} \langle k_i\rvert q\ran­gle\lvert v_i\ran­gle

can­not choose to skip an in­di­vid­ual old to­ken af­ter that to­ken has been folded into StS_t. Every stored di­rec­tion that over­laps the query con­tributes. The delta rule can cor­rect the state around the cur­rent key, but stale in­for­ma­tion in other di­rec­tions re­mains avail­able and can dis­tort fu­ture reads.

We there­fore need a way to for­get the old state be­fore us­ing it. Let αt∈[0,1]\al­pha_t\in[0,1] be a learned scalar re­ten­tion gate:

S~t=αtSt−1.\widetilde S_t = \alpha_t S_{t-1}.

Run the same delta rule against this gated state:

S~t=αtSt−1,forget,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write.\boxed{ \begin{aligned} \widetilde S_t &= \alpha_tS_{t-1}, &&\text{forget},\\ \lvert\widehat v_t\ran­gle &= \widetilde S_t\lvert k_t\ran­gle, &&\text{predict},\\ \lvert e_t\ran­gle &= \beta_t \left( \lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\ran­gle\lan­gle k_t\rvert, &&\text{write}. \end{aligned} }

This is Gated DeltaNet. The or­der mat­ters: for­get first, pre­dict from the re­tained state, then cor­rect that pre­dic­tion. If we pre­dicted be­fore for­get­ting, the er­ror would de­scribe a dif­fer­ent mem­ory from the one we up­date.

Expanding the re­cur­rence gives

St=αtSt−1(I−βt∣kt⟩⟨kt∣)+βt∣vt⟩⟨kt∣.S_t = \alpha_tS_{t-1} \left( I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert \right) + \beta_t\lvert v_t\ran­gle\lan­gle k_t\rvert.

The delta rule gives tar­geted re­place­ment; the scalar gate gives global era­sure. They solve dif­fer­ent prob­lems and are com­ple­men­tary.

But αt\al­pha_t still makes one de­ci­sion for the en­tire ma­trix. The model must re­tain or for­get every key chan­nel at the same rate.

4. Kimi Delta Attention: for­get each chan­nel in­de­pen­dently

Kimi Delta Attention re­places Gated DeltaNet’s scalar re­ten­tion with a vec­tor αt∈[0,1]dk\al­pha_t\in[0,1]^{d_k}. Put the vec­tor on the di­ag­o­nal:

Dt=Diag⁡(αt)∈Rdk×dk.D_t = \operatorname{Diag}(\alpha_t) \in\mathbb R^{d_k\times d_k}.

Our state maps keys to val­ues, so the key chan­nels are the columns of SS. Right-multiplication ap­plies a dif­fer­ent re­ten­tion fac­tor to every one:

S~t=St−1Dt.\widetilde S_t = S_{t-1}D_t.

Everything else is the delta rule we have al­ready de­rived:

S~t=St−1Dt,forget each key channel,∣v^t⟩=S~t∣kt⟩,predict,∣et⟩=βt(∣vt⟩−∣v^t⟩),correct,St=S~t+∣et⟩⟨kt∣,write,∣ot⟩=St(s∣qt⟩),s=dk−1/2,read.\boxed{ \begin{aligned} \widetilde S_t &= S_{t-1}D_t, &&\text{forget each key chan­nel},\\ \lvert\widehat v_t\ran­gle &= \widetilde S_t\lvert k_t\ran­gle, &&\text{predict},\\ \lvert e_t\ran­gle &= \beta_t \left( \lvert v_t\ran­gle-\lvert\wide­hat v_t\ran­gle \right), &&\text{correct},\\ S_t &= \widetilde S_t+\lvert e_t\ran­gle\lan­gle k_t\rvert, &&\text{write},\\ \lvert o_t\ran­gle &= S_t(s\lvert q_t\ran­gle), \qquad s=d_k^{-1/​2}, &&\text{read}. \end{aligned} }

That is KDA. Compared with Gated DeltaNet, the con­cep­tual change is only the pro­mo­tion

αt⟶Dt=Diag⁡(αt).\al­pha_t \quad\longrightarrow\quad D_t=\operatorname{Diag}(\alpha_t).

The ef­fect is sub­stan­tial: one chan­nel can be cleared while an­other is re­tained.

4.1 Why the tran­si­tion is di­ag­o­nal-plus-low-rank

Expand the KDA cor­rec­tion:

St=St−1Dt+βt(∣vt⟩−St−1Dt∣kt⟩)⟨kt∣=St−1Dt(I−βt∣kt⟩⟨kt∣)⏟At+βt∣vt⟩⟨kt∣.\begin{aligned} S_t &= S_{t-1}D_t + \beta_t \left( \lvert v_t\ran­gle - S_{t-1}D_t\lvert k_t\ran­gle \right) \langle k_t\rvert\\ &= S_{t-1} \underbrace{ D_t \left( I-\beta_t\lvert k_t\ran­gle\lan­gle k_t\rvert \right) }_{A_t} + \beta_t\lvert v_t\ran­gle\lan­gle k_t\rvert. \end{aligned}

The key-space tran­si­tion is

At=Dt−βtDt∣kt⟩⟨kt∣=Dt−∣bt⟩⟨at∣,\begin{aligned} A_t &= D_t-\beta_tD_t\lvert k_t\ran­gle\lan­gle k_t\rvert\\ &= D_t-\lvert b_t\ran­gle\lan­gle a_t\rvert, \end{aligned}

where

∣bt⟩=Dt∣kt⟩,⟨at∣=βt⟨kt∣.\lvert b_t\ran­gle=D_t\lvert k_t\ran­gle, \qquad \langle a_t\rvert=\be­ta_t\lan­gle k_t\rvert.

So AtA_t is a di­ag­o­nal ma­trix mi­nus a rank-one ma­trix: a di­ag­o­nal-plus-low-rank, or DPLR, tran­si­tion. DPLR de­scribes the dk×dkd_k\times d_k tran­si­tion act­ing on key space. The mem­ory state it­self is still the dv×dkd_v\times d_k ma­trix StS_t.

The full jour­ney can now be sum­ma­rized com­pactly:

The im­ple­men­ta­tion usu­ally stores gt=log⁡αt­g_t=\log\al­pha_t with gt≤0g_t\leq0, then ob­tains the re­ten­tion fac­tors as exp⁡(gt)\exp(g_t). In the trans­posed dk×dvd_k\times d_v lay­out used by the ref­er­ence code, the re­cur­rence is only five lines:

state = state * g_t.exp().un­squeeze(-1) pre­dic­tion = ein­sum(“bhkv,bhk->bhv”, state, k_t) resid­ual = be­ta_t.un­squeeze(-1) * (v_t - pre­dic­tion) state = state + ein­sum(“bhk,bhv->bhkv”, k_t, resid­ual) out­put = ein­sum(“bhk,bhkv->bhv”, q_t * scale, state)

See the of­fi­cial naive_re­cur­ren­t_kda ref­er­ence.

Half-Life ported to Mac OS 9 | Mac Classic

mac-classic.com

Half-Life has fi­nally landed for PowerPC based Macintosh com­put­ers 28 years af­ter it’s orig­i­nal re­lease! Half-Life is a story dri­ven first-per­son shooter, that fol­lows sci­en­tist Gordon Freeman, who is a the­o­ret­i­cal physi­cist try­ing to sur­vive and es­cape the Black Mesa Research Facility af­ter a failed ex­per­i­ment opens a por­tal to an alien di­men­sion.

The game was orig­i­nally planned to be re­leased for Mac OS 9 by Valve in 1999, but was can­celled shortly be­fore launch. Valve did­n’t bring Half-Life to Mac OS X un­til 2013 which was well into the in­tel based CPU era, and now we fi­nally have a re­lease for PowerPC based ma­chines.

This port has been ac­com­plished by GitHub user doc­tashay us­ing a fork of Xash3D FWGS, (a re-im­ple­men­ta­tion of the GoldSrc en­gine). It’s playable from start to fin­ish, in­cludes mul­ti­player sup­port, a demo of Uplink, along with down­loads for Blue Shift and Opposing Force.

This re­lease sup­ports G3 and G4 PowerPC based com­put­ers run­ning Mac OS 9.0 or later. Performance heav­ily de­pends on the GPU pre­sent in your ma­chine, iMacs, iBooks etc. may strug­gle with per­for­mance you have less than 8Mb VRAM.

This re­lease also in­cludes:

Half-Life

Half-Life

Half-Life: Blue Shift

Half-Life: Blue Shift

Half-Life: Opposing Force

Half-Life: Opposing Force

This is a huge achieve­ment for the Macintosh gam­ing com­mu­nity, and doc­tashay de­serves con­sid­er­able credit for the work in­volved in bring­ing Half-Life to the PowerPC plat­form.

User Interfaces of the Demo Scene

www.datagubbe.se

Exploring the tools of the trade

July 2026

Ahh, the demo scene - a dig­i­tal art sub­cul­ture, a mot­ley gang of cre­ative nerds, and a favourite pas­time of mine. So much amaz­ing art, mu­sic and code has been pro­duced by sceners, and any­one who comes into con­tact with the scene might start to won­der ex­actly how - apart from hours of ded­i­cated grind, of course.

The scene has a long and sto­ried tra­di­tion of build­ing its own tools. Sometimes from scratch, some­times by steal­ing ideas or even code from ex­ist­ing of­fer­ings. In a com­bi­na­tion of teenage in­ex­pe­ri­ence, old habits, a pen­chant for ex­per­i­ment­ing and a de­sire to make things look cool, this has re­sulted in some rather pe­cu­liar user in­ter­faces.

A few of them are pre­sented be­low. Most of them are for the Amiga, but sev­eral other plat­forms are rep­re­sented as well.

Elite Sinus Producer

Not want­ing to be ac­cused of click­bait­ing, let’s start off with one of the main at­trac­tions: Elite Sinus Producer (sinus here is to be un­der­stood as sine), made by Ipec Elite for the Amiga. Demos are usu­ally de­scribed as real time”, which is true in the sense that they’re (almost) never just an­i­ma­tion play­ers, and that demo ef­fects are pro­duced, frame by frame, from code. However, to achieve seem­ingly im­pos­si­ble tech­ni­cal feats, ex­ten­sive cheating” is em­ployed. The most com­mon cheat is prob­a­bly the so called pre­calc, mean­ing that in­stead of do­ing com­plex maths on a 7 MHz (or even slower) CPU, lookup ta­bles are uti­lized. A plethora of tools for cre­at­ing such lookup ta­bles ex­ist. This is one of them.

When first start­ing Elite Sinus Producer, the user is met with this menu. Upon press­ing an F-key to make a se­lec­tion, the cor­re­spond­ing menu op­tion is high­lighted and a sam­ple of a cuckoo clock is played. Loudly.

Here I’ve se­lected the Flower” op­tion. Pretty, no? This can then be saved to disk as (presumably) as­sem­bly source code. Any sprite would look nice when mov­ing along this path!

There’s also a handy help screen, which is so out­landish I had to grab a short film clip of it: Part of the back­ground con­sists of mov­ing blue raster bars, and the other part flashes be­tween red and cyan. This ob­vi­ously helps im­mensely when read­ing the text, dis­played in a font de­signed with noth­ing but leg­i­bil­ity in mind.

Text Based Interfaces

Plenty of scene re­lated tools are ei­ther com­mand line util­i­ties or pre­dom­i­nantly text based. First out, we have the as­sem­blers. There was a wide va­ri­ety of as­sem­blers for the Amiga, but the scene al­ways favoured Seka, Asm-One and other de­riv­a­tives of the same con­cept. There were so many dif­fer­ent ver­sions and hacks (Trash’m-one, for ex­am­ple) that the sprawl­ing fam­ily tree ri­vals that of Unix sys­tems.

Here’s Seka 2.0, and as we can see, it’s based on a com­mer­cial as­sem­bler. This type of as­sem­bler al­ways asks the user for the size of the work­ing mem­ory to be al­lo­cated. They then en­ter a com­mand line mode, which can be used to ex­am­ine RAM mem­ory and CPU reg­is­ters, and load source files into the ed­i­tor proper.

Here’s AsmOne in one of its many in­car­na­tions. It’s quite sim­i­lar to Seka (in fact, it’s Seka-Updated”), but has more built-in com­mands and pre­sum­ably other im­prove­ments as well (I’ve never been much of a coder). Interface-wise, it opens its own screen, as op­posed to run­ning in a win­dow on the de­fault Workbench screen.

What if you found a piece of cool mu­sic or graph­ics in, say, a game? What if you wanted to steal some sam­ples, or a sprite, or per­haps just save an en­tire tune for easy lis­ten­ing? Then you’d need a rip­per - a tool for hunt­ing through your com­put­er’s mem­ory, look­ing for rem­nants of such data af­ter quit­ting the game. There were tons of var­i­ous such rip­pers for the Amiga. Here’s Multi-Ripper, which has a fairly rep­re­sen­ta­tive set of fea­tures.

Here’s an­other type of rip­per, specif­i­cally writ­ten to look for Seka as­sem­bly sources in mem­ory af­ter the com­puter had crashed. The Amiga, like most other home com­put­ers, had no mem­ory pro­tec­tion, and demo cod­ing is a no­to­ri­ously crash-prone ac­tiv­ity. Saving of­ten was com­mon prac­tice, but even sea­soned coders some­times messed up and for­got. With a bit of luck, the code could be ex­tracted from mem­ory af­ter a warm re­boot.

Here’s an­other type of sine pre­cal­cu­la­tor, called The Sinus Creator. I have no idea if the num­bers I’ve en­tered in the screen­shot make sense, but the two-win­dow text in­ter­face is in­ter­est­ing. Of course, the re­sult can be saved as a Seka source file.

Music Trackers

Demo mu­sic has his­tor­i­cally been made in track­ers. Far from tra­di­tional no­ta­tion, a tracker lets the mu­si­cian en­ter tones along with var­i­ous mod­i­fiers and ef­fects in some­thing that’s more rem­i­nis­cent of a pro­gram­ming ed­i­tor rather than reg­u­lar com­pos­ing. They also let the user man­age in­stru­ments, whether sam­pled or syn­the­sized.

Sample-based track­ers on the Amiga have an even more sprawl­ing fam­ily tree than that of Amiga as­sem­blers, but they all orig­i­nate from Karsten Obarski’s com­mer­cial Ultimate Soundtracker from 1987. Being com­mer­cial and thus cost­ing money, it was soon picked apart by sceners, which re­sulted in NoiseTracker, which was then re­vamped into ProTracker, which in turn ex­ists in so many var­i­ous ver­sions and re-hashed hacks that it’s nigh im­pos­si­ble to keep track (heh) of. The sprawl is even sprawlier than that of Amiga as­sem­blers!

SoundMonitor 1.0 for the Commodore 64 (by Chris Huelsbeck) is­n’t strictly speak­ing a scene re­lease, and was­n’t called tracker”. However, the in­ter­face (with one track” per avail­able sound chan­nel) is what in­spired the pre­vi­ously men­tioned Ultimate Soundtracker, and is thus in­cluded here for pos­ter­ity and cor­rect­ness.

NoiseTracker by the Swedish duo Mahoney and Kaktus was­n’t the first, but it was im­mensely pop­u­lar and came to de­fine the tracker ex­pe­ri­ence for years to come. Its legacy lives on just not in Protracker on the Amiga, but on sev­eral other plat­forms as well.

Here’s Protracker’s file picker. Just like in NoiseTracker above, it’s ac­cessed by click­ing the Disk Op.” but­ton in the rather dense in­ter­face. It’s hard to de­scribe what a strange ex­pe­ri­ence this UI de­liv­ers, be­cause it’s al­most - but not quite - in­tu­itive to some­one used to more main­stream Amiga pro­grams. It’s easy to misclick, mis­un­der­stand or just plain miss things. To il­lus­trate its idio­syn­cratic de­sign, take note of the ver­ti­cal EXIT but­ton, which quits back to the main menu. It’s con­ve­niently placed be­tween the up and down ar­rows used for scrolling the pur­ple file and di­rec­tory list­ing.

Here’s Digicomposer 1.0 for the Atari ST/e, which in turn builds on Noisetracker for the Atari. The in­ter­face is clearly more than just in­spired by its Amiga coun­ter­part. It’s but one of a whole menagerie of track­ers for the Atari ST, some digi” (sample-based) like this one, some for YM-chip based mu­sic.

Fasttracker II for MS-DOS is iconic in its own right. With sup­port for Gravis UltraSound and other ad­vanced PC sound hard­ware, it has fea­tures for 16-bit sam­ples, bizarre amounts of sound chan­nels, and even lets the user play a sim­ple ver­sion of Snake if they so de­sire.

Abyss’ Highest Experience is a chip­tune tracker for the Amiga, in­tended to sound like the Commodore 64′s SID chip. The user in­ter­face is an in­ter­est­ing crossover be­tween Soundtracker and a more mod­ern, Workbench 2.0-like look.

I don’t know much about Megatizer for the Atari ST, but it sure does look cool!

JamCrackerPro for the Amiga es­chewed the tra­di­tional tracker UI and opted for a sys­tem-friendly, multi-win­dow in­ter­face.

Disk Copiers

In or­der to dis­trib­ute demos (and pi­rated soft­ware), disk copy­ing was a com­mon ac­tiv­ity on the scene. Commodore pro­vided a disk copier in AmigaOS, but it was­n’t al­ways up to the task of copy­ing demos and games that by­passed the file sys­tem by writ­ing di­rectly to the tracks of a floppy. Thus, spe­cial soft­ware was needed!

Like SoundMonitor, X-Copy is­n’t strictly a scene re­lease. Although orig­i­nally writ­ten by sceners, it was re­leased com­mer­cially and then heav­ily pi­rated on the scene. The promi­nently fea­tured grids con­tain one square for each track on the disk, dis­play­ing the sta­tus for copy­ing that par­tic­u­lar part of the floppy. When a copy was fin­ished (which could take some time), the pro­gram help­fully played a lit­tle boing” sound.

X-Copy was re­leased in many ver­sions dur­ing the Amiga hey­days. Here’s X-Copy 3.0, ap­par­ently in an il­licit va­ri­ety dis­trib­uted by the crack­ing group Paradox.

Personally, I pre­ferred D-Copy, mostly be­cause I thought the user in­ter­face looked cool (and I still do). It is, as far as I know, a bona fide non-com­mer­cial scene prod­uct.

Other Tools

A short ar­ti­cle like this can only ever scratch the sur­face of the nu­mer­ous demo scene tools ever cre­ated. Here are but a few more var­i­ous tools and in­ter­faces, se­lected for be­ing in­ter­est­ing and/​or rep­re­sen­ta­tive of their kind.

This is Titanics Cruncher for the Amiga. A cruncher uses asym­met­ri­cal com­pres­sion of ex­e­cutable files. This saves space on disk, with the trade­off be­ing longer load­ing times, due to the de­crunch­ing (decompression) of the ex­e­cutable upon run­ning it. Several other crunch­ers ex­ist, and on a va­ri­ety of plat­forms.

Before the Internet, there was the Bulletin Board System, or BBS. Sceners were usu­ally in­ter­ested in Elite BBS:es, mean­ing ones that of­fered pi­rated soft­ware for down­load. There were few, if any, Elite BBS:es with­out cool ANSI (or PETSCII, on the C64) graph­ics, for ex­am­ple in an­i­mated menu screens. Thus, many var­i­ous ANSI ed­i­tors popped up for dif­fer­ent plat­forms. This is Digital Intelligence’s Ansi-Editor v2.4 for the Amiga, and it has a very pe­cu­liar user in­ter­face. The tool­bar at the bot­tom dis­plays the cur­rently ac­tive colours, but you can’t se­lect them from there. Click all you want, to no ef­fect - you have to use the pull-down menu to ac­tu­ally pick one.

Before Twitter, Facebook, IRC and even wide­spread ac­cess to modems and BBS:es, scroll texts were the com­mu­ni­ca­tion medium of choice for the scene. Cool scrollers re­quired cool fonts, and it could be help­ful with a ded­i­cated font (or charset) ed­i­tor. Here’s Charedit by Escape, for MS-DOS.

Home com­put­ers of­ten came with cus­tom or oth­er­wise es­o­teric hard­ware. The Atari Falcon, for ex­am­ple, sported a Motorola 56001 dig­i­tal sig­nal proces­sor. DSPdit by tSCc (short for The Sirius Cybernetics Corporation) is an ed­i­tor and as­sem­bler made specif­i­cally for DSP56k pro­gram­ming. It uses the stan­dard GEM toolkit, which gives it a very clean and pro­fes­sional look.

The first com­puter virus for the Amiga was a boot­block virus, in­fect­ing the boot sec­tor of floppy disks. As such, it could po­ten­tially ruin games and other flop­pies with cus­tom boot­blocks. It was con­structed by the Swiss Cracking Association, SCA. The virus spread like wild­fire and per­haps SCA was plagued by a guilty con­science, be­cause later they also pro­duced the very first virus killer for the Amiga - de­signed to counter their own virus. The mega-mighty in­ter­face is easy enough to grasp!

As men­tioned above, plenty of Amiga demos were so called track­mos, mean­ing they did­n’t use the file sys­tem, but rather stored data straight on the tracks of a floppy disk. This soon re­sulted in var­i­ous tool­ing be­com­ing avail­able for cre­at­ing such track­mos. Mostly it was just code shared be­tween sceners, but there were also com­plete soft­ware suites, like TrackmoDOS by Poison of NOVA. Among other things, it came with this or­tho­dox file man­ager for writ­ing and delet­ing files to a trackmo floppy. Being a scener tool, the GUI of course sports some rather spiffy colour gra­di­ents.

This is RAW, which is­n’t re­ally a tool, but rather a disk mag­a­zine, or diskmag. A diskmag is just that - a pe­ri­od­i­cal pub­lished on one or more floppy disks. Most of the ar­ti­cles were usu­ally scene re­lated, al­though some mags branched out and fea­tured every­thing from short sto­ries and po­ems to es­says about his­tory and pol­i­tics. RAW was one of the most pop­u­lar Amiga diskmags dur­ing the early 1990s, and as we can see, the UI was shiny and tex­tured long be­fore Frutiger Aero was a thing. It even came with a built-in palette ed­i­tor (pictured), al­low­ing the user to cus­tomize the text colours.

This is FuckPaint, a pixel painter for the Atari Falcon. It is in­cluded here solely on the merit of its name, and its equally classy Analizer” tool.

Although not at all a scene pro­duc­tion, Deluxe Paint must be men­tioned here. To the best of my knowl­edge, the scene never pro­duced a pixel painter for the Amiga (except later ports of Grafx2), pre­sum­ably be­cause no­body saw the need for one. I don’t think a sin­gle Amiga user ex­isted that did­n’t have a copy of this soft­ware, ei­ther bought sep­a­rately, bun­dled with the ma­chine or just pi­rated. It was also com­pletely dom­i­nant on the PC. There were other Amiga pixel painters, such as Brilliance and Personal Paint, but their com­bined mar­ket share was a mere frac­tion of this gi­ant, which also dom­i­nated graph­ics cre­ation for games well into the the mid-1990s.

That’s enough demo scene in­ter­faces for one help­ing. For those still want­ing more, I rec­om­mend this gallery of utild­isk menus.

Take care and happy hack­ing!

More Tailscale tricks for your jailbroken Kindle

tailscale.com

If you man­aged to put Tailscale on a jail­bro­ken Kindle be­fore it up­dated too far ahead, you got some­thing pretty great, even if it was­n’t the full Tailscale ex­pe­ri­ence. But good things come to those who wait (or dig around on GitHub).

Open-source de­vel­op­ers have im­proved the Tailscale ex­pe­ri­ence on one of the weak­est com­put­ers you own. If your Kindle is jail­bro­ken, an up­dated ver­sion of Mitanshu Sukhwani’s Tailscale im­ple­men­ta­tion of­fers a few new things:

Tailscale SSH en­abled by de­fault, so you don’t have to en­able USBnetworking SSH and its very ob­vi­ous de­fault user/​pass­word

A proxy mode that lets apps like KOReader to reach other nodes on your tail­net, like a Calibre/OPDS or Wallabag server

A full TUN mode that, on some Kindles, can make Tailscale net­work­ing work at the de­vice level

Let’s dig into each one and how to set them up. As be­fore, this is com­mu­nity code work­ing on a very un­of­fi­cial de­vice state; bring your pa­tience along.

Tailscale on a Kindle, now with prox­ies

The last time we wrote about Tailscale on a Kindle, the client was ba­sic, but it worked. The Kindle showed up on your tail­net, com­plete with a green dot in the web ad­min con­sole. You could reach the Kindle by its Tailscale IP ad­dress. You could even SSH into the Kindle over Tailscale, which was handy for fur­ther tin­ker­ing.

But reachable via Tailscale” is not the same as routing all in­com­ing and out­go­ing traf­fic across your tail­net,” it turns out. Tailscale on a jail­bro­ken Kindle is typ­i­cally forced to run in user­space mode, which means it can­not use the de­vice’s own net­work rout­ing layer, known as TUN mode. You could start Tailscale, and then start an app like KOReader, but when you tried to con­nect to an­other Tailscale de­vice, like your Calibre server at 100.x.y.z, it would go like this:

KOReader (or any app) asks the Kindle’s root OS how to reach 100.x.y.z

The Kindle, lack­ing Tailscale rout­ing, can­not reach that Tailscale IP ad­dress

KOReader drops the con­nec­tion

An up­date to the Kindle KUAL app by grey­wolf1499 pro­vides dif­fer­ent modes that work around this. Now, when you try to reach an­other Tailscale IP ad­dress on your Kindle, it can go like this:

You start Tailscale in proxy mode

You set up KOReader or an­other ap­p’s proxy set­tings to con­nect to 127.0.0.1:1055

KOReader tells the proxy it wants to reach 100.x.y.z

Tailscale’s dae­mon tailscaled, lis­ten­ing on port 1055, routes the con­nec­tion through Tailscale

E-books, ar­ti­cles, and other data flows be­tween your Tailscale-running Kindle and other Tailscale de­vices

This Tailscale proxy of­fers two modes, SOCKS5 and HTTP CONNECT, for apps that may pre­fer one or the other. This opens up a good bit more util­ity for your more-con­nected Kindle.

What you can do with a prox­ied Kindle

A few wild ideas, de­pend­ing on how dug in you want to get:

Calibre or Wallabag servers, as men­tioned

Audiobookshelf con­nec­tion through KOReader

Use Readest to track read­ing progress across de­vices

Linking KOReader’s RSS reader to a a self-hosted feed server

Accessing min­i­mal­ist dash­boards and web pages in the (pretty bad) Kindle browser

Using a Bluetooth key­board and the kterm app to SSH into tail­net de­vices

Is that last one all that prac­ti­cal? Not re­ally. But is there a pleas­ant warmth, know­ing that you’ve added the least likely thin client to your what-if kit? For some types, types I know quite well: yes.

The Tailscale plu­gin for KOReader (with Kobo and PocketBook sup­port)

If you don’t re­ally need any Tailscale pow­ers out­side the highly ca­pa­ble KOReader app, check out this Tailscale KOReader plu­gin. It does­n’t make your Kindle ac­ces­si­ble over your tail­net, like the KUAL-based app. But it does au­to­mat­i­cally cre­ate the proxy in­ter­faces that are needed for reach­ing your con­tent servers from your KOReader-running Kindle—or your Kobo de­vice, or your PocketBook.

I haven’t been able to re­ally try this ex­ten­sion out my­self; my 11th-generation stan­dard Kindle does­n’t play well with it at the mo­ment. It’s been Tested on Kindle PW5/PW6, Kobo, and PocketBook”—it’s nice to see Tailscale come to some other KOReader-friendly de­vices, too.

Installation is not too hard, at least if you made it this far into jail­break­ing al­ready. You copy the plu­gin into KOReader’s plu­g­ins di­rec­tory, trig­ger an Install/Update Tailscale” from KOReader’s menu, copy a Tailscale key into a di­rec­tory, then tog­gle Tailscale on in the KOReader menu. From there, you con­fig­ure KOReader with one of its proxy ad­dresses (127.0.0.1:1055 for SOCKS5, :1056 for HTTP CONNECT), then give other plu­g­ins the Tailscale IP ad­dresses you need to reach.

Victoria Riley Barnett’s repos­i­tory notes that the plu­gin works great with a SyncThing plu­gin for KOReader. KOReader is like its own sep­a­rate OS for jail­bro­ken Kindles at this point,

So now you’ve got a lot more op­tions and weird pro­jects avail­able to you, through this al­ready quite-strange lit­tle slab. If you’ve worked up a weirdly use­ful Tailscale setup on your Kindle, Kobo, or other e-pa­per de­vice, we’d love to hear about it. Let us know on Red­dit, Dis­cord, Bluesky, X, Mastodon, or LinkedIn.

Inside Zig's Incremental Compilation

mlugg.co.uk

As a mem­ber of the Zig core team, one of the most im­pact­ful pro­jects I’ve been in­volved with is the im­ple­men­ta­tion of in­cre­men­tal com­pi­la­tion into the Zig com­piler. This fea­ture al­lows the com­piler to de­tect which in­di­vid­ual func­tions and de­c­la­ra­tions have changed since a pro­ject was last built, re­com­pile only that code, and di­rectly patch the re­sult­ing bytes into the out­put bi­nary, mak­ing the re­build ex­tremely fast.

The Zig pro­ject has been work­ing to­wards this fea­ture for a long time, and over the last few re­lease cy­cles, it has fi­nally gone from a proof-of-con­cept qual­ity fea­ture to one which is vi­able for real-world pro­jects and which most of the Zig core team makes daily use of.

Today, us­ing Zig’s in­cre­men­tal com­pi­la­tion, you can make changes to real, com­plex ap­pli­ca­tions in a mat­ter of mil­lisec­onds.

But don’t just take my word for it! Here’s a sim­ple video (no au­dio) demon­strat­ing me us­ing Zig to quickly make and test some changes to Fizzy, a pixel ed­i­tor ap­pli­ca­tion. The ini­tial build takes around 5 sec­onds, and then every time I make a change, a re­build com­pletes in 50 – 70ms.

For this demo, I had to up­grade Fizzy to Zig’s mas­ter branch. This is be­cause while Zig 0.16.0 does have sup­port for in­cre­men­tal com­pi­la­tion, it is miss­ing some im­por­tant linker fea­tures which have since been im­ple­mented. This means that if you pre­fer to stick to tagged re­leases of Zig, you likely won’t be able to try this out un­til 0.17.0 drops; sorry!

If you’re al­ready con­vinced and just want to know how to use this, great! Head on down to the last sec­tion of this post to find out. But per­haps you’re un­der­stand­ably skep­ti­cal that this is ap­plic­a­ble to most pro­jects, or, like me, you just en­joy learn­ing how stuff like this works. For all of you folks, let’s dig into the de­tails!

Processing Source Files

The Zig com­pil­er’s pipeline can be split up into a few dif­fer­ent parts, which we’ll look at in or­der. The first part works at the gran­u­lar­ity of en­tire source files, and ba­si­cally con­sists of run­ning the fol­low­ing process in a loop:

Read in a source file from disk

Parse that file into an AST

Convert that AST into a for­mat named ZIR us­ing a pass named AstGen”

If you’re cu­ri­ous, ZIR (Zig Intermediate Representation) is an un­typed SSA-form IR—but don’t worry if you have no idea what that means, be­cause it won’t re­ally mat­ter here. All we care about is that we’re con­vert­ing an en­tire source file into a dif­fer­ent for­mat.

While AstGen runs, it learns about all Zig im­ports (@import(“foo.zig”)) in the source file, so we can re­peat this en­tire process on all of the im­ported files. So by run­ning this process in a loop, we will ul­ti­mately dis­cover every Zig source file in the com­pi­la­tion, and will con­vert them all to ZIR.

This part of the pipeline ac­tu­ally has sev­eral use­ful prop­er­ties:

The pro­cess­ing run on each file is a pure func­tion of that file’s con­tents, in­volv­ing no shared or ex­ter­nal state

Parse and AstGen are both quite fast on their own: on my lap­top, run­ning them both over the en­tire src/ di­rec­tory of the Zig com­piler (with no par­al­lelism at all) takes around 920ms

Thanks to Zig’s us­age of data-ori­ented de­sign pat­terns, ZIR can be triv­ially writ­ten to and read from disk with one writev/​readv sys­tem call—there is no serialization” step.

These prop­er­ties have two nice con­se­quences.

Firstly, as­sum­ing one task” per source file, this en­tire process is em­bar­rass­ingly par­al­lel. That means we can triv­ially run it on a thread pool by queu­ing up a task every time we dis­cover a new source file from an im­port—the only shared state (which we’ll just pro­tect with a mu­tex) is a hash set keep­ing track of which file paths we have al­ready seen.

Secondly, and ar­guably even more im­por­tantly, these prop­er­ties make it very straight­for­ward to im­ple­ment in­cre­men­tal com­pi­la­tion for this part of the pipeline. All we need to do is cache each source file’s gen­er­ated ZIR on disk, and only re­build it when we de­tect that the file changed.

Both of these op­ti­miza­tions have been en­abled by de­fault in Zig for years—they are bat­tle-tested and make this part of the pipeline near-in­stan­ta­neous in most cases. If you’re us­ing Zig, you can see how fast this is us­ing the progress out­put on stderr—when it says AST Lowering”, this part of the pipeline is run­ning. I’d guess that a lot of Zig users only even no­tice that hap­pen­ing the very first time they run the com­piler (because on its first run the com­piler needs to do this work for the en­tire Zig stan­dard li­brary and com­pil­er_rt).

Okay, so, we made this part fast! That’s great, but the bad news is that this was the easy part—lots of com­pil­ers can al­ready do this kind of caching. From here, things will get trick­ier.

Semantic Analysis

The next part of the pipeline is ar­guably the most im­por­tant: se­man­tic analy­sis. This in­cludes both type check­ing and comp­time eval­u­a­tion.

The job of se­man­tic analy­sis is es­sen­tially to interpret” the ZIR we pro­duced ear­lier, emit­ting com­pile er­rors (such as type er­rors) along the way; and, for run­time func­tions, build­ing an­other in­ter­me­di­ate rep­re­sen­ta­tion (Analyzed Intermediate Representation, or AIR for short) which can be sent on to later parts of the pipeline.

Before we move for­ward, a quick ter­mi­nol­ogy clar­i­fi­ca­tion. A container-level de­c­la­ra­tion” is the Zig equiv­a­lent of what other lan­guages call a top-level de­c­la­ra­tion”. That term is in­ac­cu­rate in Zig, be­cause con­tainer-level de­c­la­ra­tions do not have to be at the top level syn­tac­ti­cally, but the con­cept is the same. If I say container-level de­c­la­ra­tion”, I ba­si­cally mean a func­tion, global con­stant, or global vari­able”.

Semantic analy­sis is the most dif­fi­cult part of the com­piler to han­dle in­cre­men­tally. Perhaps un­sur­pris­ingly then, this is where lan­guage de­sign starts to mat­ter a lot: while I am pretty con­fi­dent that most mod­ern lan­guages could sup­port in­cre­men­tal com­pi­la­tion sim­i­lar to how we do, cer­tain de­sign de­ci­sions can make that much more dif­fi­cult. Zig has had its de­sign tweaked over the years (sometimes con­tro­ver­sially) specif­i­cally so that it is eas­ier to sup­port fast in­cre­men­tal com­pi­la­tion.

The name of the game here is to split up your com­pi­la­tion into a bunch of pieces which you can mostly an­a­lyze in­de­pen­dently of one an­other, and, cru­cially, where the de­pen­den­cies that do ex­ist be­tween those pieces can be eas­ily mod­eled in a de­pen­dency graph.

In the Zig com­piler, we call these pieces analysis units”, or I might some­times just say unit” for short. I’m go­ing to ever so slightly sim­plify things here and tell you that the Zig com­piler has four dif­fer­ent kinds of analy­sis unit:

The lay­out (size, align­ment, etc) of a struct or union type.

The type of a con­tainer-level de­c­la­ra­tion.

The value of a con­tainer-level const de­c­la­ra­tion.

The body of a run­time func­tion.

During se­man­tic analy­sis of a par­tic­u­lar unit, we pop­u­late a set of other units which this unit de­pends on. Let’s look at a ba­sic ex­am­ple:

var glob­al_0: u32 = 123; const glob­al_1: u32 = 456;

pub fn foo(cond: bool) u32 { if (cond) { re­turn glob­al_0; } else { re­turn glob­al_1; } }

Here’s what hap­pens when we an­a­lyze the body of the func­tion foo:

Because the ar­gu­ment cond is not comp­time-known, we se­man­ti­cally an­a­lyze both branches of the if

Take a pointer to glob­al_0, in prepa­ra­tion to load from itAdd de­pen­dency: type of glob­al_0

Add de­pen­dency: type of glob­al_0

Load glob­al_0 at run­time, be­cause it is var so does not have a comp­time-known value

Take a pointer to glob­al_1, in prepa­ra­tion to load from itAdd de­pen­dency: type of glob­al_1

Add de­pen­dency: type of glob­al_1

Load glob­al_1 at com­pile time, be­cause it has a comp­time-known val­ueAdd de­pen­dency: value of glob­al_1

Add de­pen­dency: value of glob­al_1

So we end up with this func­tion body de­pend­ing on the types of glob­al_0 and glob­al_1, and the value of glob­al_1 (since that’s comp­time-known). This tells the com­piler that if the type of glob­al_0 or glob­al_1 changes, or the comp­time-known value of glob­al_1 changes, the func­tion should be re-an­a­lyzed.

Dependencies on the body of a run­time func­tion are im­pos­si­ble (at least in the sim­pli­fied view I’m pre­sent­ing here). This means that func­tion body analy­sis units can only have outgoing” edges in the de­pen­dency graph (i.e. they may de­pend on other units, but other units do not de­pend on them).

Dependencies on the value of a const de­c­la­ra­tion only arise due to Zig’s abil­ity to use those at comp­time. If not for that lan­guage fea­ture, de­pen­den­cies on the value of a de­c­la­ra­tion would be im­pos­si­ble, just as it is im­pos­si­ble to de­pend on the body of a run­time func­tion.

Dependencies on a type’s lay­out arise, in short, from hav­ing val­ues of that type, or from need­ing to know some­thing about the type’s lay­out. I’m not go­ing to dis­cuss this any fur­ther here, be­cause it’s a bit com­pli­cated and quite spe­cific to Zig’s type sys­tem, but it’s not fun­da­men­tally dif­fer­ent.

Okay, so, we’ve told the com­piler about when re-analy­sis of one thing needs to also trig­ger re-analy­sis of an­other thing. However, there’s one more puz­zle piece here—source code de­pen­den­cies. By it­self, this de­pen­dency graph is use­less: what do we ac­tu­ally do when the user asks for a re­com­pile (what we call an incremental up­date”)? We don’t know the first thing to re-an­a­lyze!

To solve this prob­lem, we track de­pen­den­cies of analy­sis units, not only on other units, but also on pieces of source code. In the cases we’ve looked at so far, these are all re­ally sim­ple: in the snip­pet above, the unit type of glob­al_0” de­pends on the source code of glob­al_0, the units type of glob­al_1” and value of glob­al_1” both de­pend on the source code of glob­al_1, and the unit body of foo” de­pends on the source code of foo. Whenever any byte of source code in the given re­gion is mod­i­fied, the de­pen­dent analy­sis unit will be marked as outdated” and re-an­a­lyzed.

Note that in re­al­ity, things can get more com­pli­cated than each unit de­pend­ing on one piece of source code. For ex­am­ple, an in­line func­tion call in Zig per­forms se­man­tic in­lin­ing, which means that it es­sen­tially trig­gers se­man­tic analy­sis of a dif­fer­ent piece of code but in the caller’s analy­sis unit. Therefore, in­line func­tion calls in­tro­duce de­pen­den­cies from the caller’s analy­sis unit on the source code of the callee.

Of course, we still need to be able to fig­ure out which re­gions of source code have changed since an in­cre­men­tal up­date. For this, ZIR con­tains hashes for spe­cific interesting” re­gions of source code (e.g. the en­tire source code for each con­tainer-level de­c­la­ra­tion), and those hashes are what you are ac­tu­ally de­pend­ing on. If the source code changes, the hash changes, and that’s easy for the com­piler fron­tend to de­tect.

Okay, that was a lot of ex­plain­ing—now let’s look at some pretty pic­tures! Here’s some Zig source code:

const luck­y_num­ber = 42;

const S = struct { x: u32 };

fn get­Some­thing() S { re­turn .{ .x = luck­y_num­ber }; }

fn testLuck(x: u32) void { if (x == luck­y_num­ber) { // do some­thing } }

ex­port fn en­try() void { const re­sult = get­Some­thing(); testLuck(re­sult.x); }

…and here’s its de­pen­dency graph (with some re­dun­dant edges re­moved for leg­i­bil­ity). The nodes on the right rep­re­sent the source code which has been hashed, while the re­main­ing nodes are all analy­sis units.

Now, let’s say we change the first line of the file to read const luck­y_num­ber = 43;. First, the com­piler low­ers the new ZIR for this file. It maps de­c­la­ra­tions from the old ZIR to the new ZIR based on the de­c­la­ra­tion names, and com­pares the source hashes as­so­ci­ated with each de­c­la­ra­tion. In this case, it suc­cess­fully maps every de­c­la­ra­tion, and it sees that one source hash changed—the one as­so­ci­ated with const luck­y_num­ber. Next, it looks at the de­pen­dency graph to find every­thing which de­pends on that source hash. In this case, it only finds one di­rect de­pen­dency:

So, the com­piler re-an­a­lyzes the value of the luck­y_num­ber de­c­la­ra­tion. If our change to the line had been a no-op (e.g. we just added some white­space), then it would de­ter­mine that the value did not change, and stop here. But in this case, the value did change! Therefore, the com­piler con­tin­ues this process, by next con­sid­er­ing any analy­sis units which de­pend on the value of luck­y_num­ber, of which there are two:

The com­piler an­a­lyzes those two units—the bod­ies of testLuck and get­Some­thing. There are no de­pen­den­cies on these units (since they’re func­tion bod­ies), so the se­man­tic analy­sis loop stops here. However, se­man­tic analy­sis of those func­tions does gen­er­ate new AIR, which brings us neatly to our next topic: code gen­er­a­tion.

Code Generation

Code gen­er­a­tion, some­times called codegen” for short, is the stage in the com­piler pipeline where AIR from se­man­tic analy­sis is con­verted to some­thing re­sem­bling ma­chine in­struc­tions. Codegen does­n’t quite emit ma­chine in­struc­tions yet—in­stead it’s some­thing called MIR (Machine Intermediate Representation)—but there is al­most a 1 – 1 map­ping be­tween MIR in­struc­tions and ma­chine in­struc­tions. There are sep­a­rate code­gen im­ple­men­ta­tions for each tar­get ar­chi­tec­ture (x86_64, aarch64, etc).

A nice thing about code­gen is that just like the whole-file pro­cess­ing ear­lier, it is an em­bar­rass­ingly par­al­lel task (at least in builds where you aren’t do­ing in­ter-func­tion op­ti­miza­tions like in­lin­ing). There is no state shared be­tween code gen­er­a­tion of dif­fer­ent func­tions, so we can have a queue of pend­ing func­tions whose AIR needs con­vert­ing to MIR, and process that queue across ar­bi­trar­ily many threads. There’s just one small gotcha, which is that we need to be care­ful to cap the size of that queue, be­cause if code­gen is ever run­ning be­hind se­man­tic analy­sis for any rea­son, the size of the queued-up AIR can add up fast!

In terms of in­cre­men­tal com­pi­la­tion, this phase of the pipeline is ac­tu­ally as sim­ple as it gets, be­cause AIR and MIR both ex­ist at the gran­u­lar­ity of in­di­vid­ual func­tions, which is the same gran­u­lar­ity in­cre­men­tal com­pi­la­tion works at. This means that there is no need for the com­piler to cache AIR or MIR at all! The AIR is thrown away as soon as code gen­er­a­tion is done, and the MIR will be thrown away right af­ter it’s con­sumed by our next stop: the linker.

Linking

Incremental link­ing is kind of a dif­fi­cult prob­lem, and I sus­pect is a big rea­son that no other ma­jor tool­chain sup­ports this kind of in­cre­men­tal com­pi­la­tion yet. General-purpose in­cre­men­tal link­ers aren’t re­ally a thing at the mo­ment, and though Wild was orig­i­nally con­cep­tu­al­ized as one, that pro­ject seems to have shifted its fo­cus firmly to­wards cold-link per­for­mance over the past cou­ple of years, with no ex­plicit time­frame for in­cre­men­tal link­ing.

David Lattimore, the cre­ator of Wild, has a blog post dis­cussing some of the dif­fi­cul­ties of in­cre­men­tal link­ing. One of those is diff­ing in­put ob­jects to fig­ure out what ac­tu­ally changed on an up­date. However, when you con­trol the en­tire com­pi­la­tion pipeline, a sim­pler de­sign pre­sents it­self which neatly side­steps that en­tire prob­lem: tightly in­te­grat­ing the linker with the com­piler.

To be­gin with, let’s just look at how the linker might work with­out in­cre­men­tal com­pi­la­tion. Because link­ing in­volves a lot of shared state, our linker is en­tirely sin­gle-threaded (maybe we’ll look into multi-threaded link­ing in the fu­ture, but for now we’re keep­ing things sim­ple). When the linker re­ceives MIR from code­gen, it first needs to con­vert that MIR into the ac­tual ma­chine code. This logic is spe­cific to the code­gen back­end, but we can’t run it un­til now be­cause it re­quires co­op­er­a­tion with the linker. That’s be­cause while emit­ting ma­chine code, the code­gen back­end gen­er­ates re­lo­ca­tions—ba­si­cally, in­struc­tions for the linker to over­write cer­tain parts of the code with spe­cific ad­dresses or val­ues (for in­stance the ad­dress of an­other sym­bol). The linker needs to save all of these re­lo­ca­tions in­ter­nally, so we need to be on the linker thread for this.

After gen­er­at­ing the ma­chine code and as­so­ci­ated re­lo­ca­tions, we re­serve space for that ma­chine code in the out­put sec­tion (usually .text). We save the ma­chine code in a buffer, save the re­lo­ca­tions to ap­ply later, and do some mis­cel­la­neous book­keep­ing work, such as adding a sym­bol table en­try.

For a non-in­cre­men­tal linker, this would be the end of the story. At the end of com­pi­la­tion, we would as­sign ad­dresses to every sec­tion, write out every­thing we re­served space for, and ap­ply all of the re­lo­ca­tions. Incremental link­ing is a bit trick­ier—writ­ing the ma­chine code to the file, as­sign­ing ad­dresses, and ap­ply­ing re­lo­ca­tions, all ide­ally needs to hap­pen be­fore we know the full con­tents of the bi­nary, and we need to be able to up­date those things later.

A lot of the com­plex­ity here is ac­tu­ally just in mov­ing things around. For ex­am­ple, if we want to add a func­tion to the .text sec­tion, but there is­n’t enough space, we need to ex­pand that sec­tion. But the sec­tion might be sur­rounded by other sec­tions, which we can’t just over­write, so we’ll need to move some­thing—ei­ther the .text sec­tion it­self, or one of the sur­round­ing sec­tions. In do­ing so, we’re go­ing to change not only file off­sets but also vir­tual ad­dresses of every­thing we move—this means we’ll need to up­date sym­bol table ad­dresses, re-ap­ply re­lo­ca­tions, etc. That’s a lot to keep track of! (There’s also a sim­i­lar prob­lem for seg­ments, one level up.)

To solve this prob­lem, Jacob Young in­tro­duced a nifty ab­strac­tion into the Zig com­piler called link.Mapped­File. It mem­ory-maps the out­put file, but more im­por­tantly tracks a tree of nodes” in that file. The root node cov­ers the en­tire file, and child nodes re­fer to spe­cific re­gions within their par­ent node. The API user can add nodes, or grow a node to a given size—in both cases, if there is not space in the par­ent to triv­ially per­form the op­er­a­tion, MappedFile deals with mov­ing other nodes around to make space. Whenever it re­sizes or moves a node, the im­ple­men­ta­tion sets a dirty” flag on that node, so that at some point the linker im­ple­men­ta­tion can de­tect this and ap­ply any nec­es­sary fix­ups, e.g. re-ap­ply­ing re­lo­ca­tions whose tar­get moved.

Right now, MappedFile has fairly prim­i­tive logic for node al­lo­ca­tion, so some­times makes sub­op­ti­mal de­ci­sions—but be­cause we’ve ab­stracted it be­hind a neat lit­tle API, we can im­prove it in­de­pen­dently go­ing for­ward.

To get to in­cre­men­tal link­ing, then, we need only slightly change the process I de­scribed ear­lier. After we fin­ish emit­ting ma­chine code, we cre­ate a node in the mapped file, large enough to hold the code—or if this func­tion al­ready ex­isted, we just re­size the ex­ist­ing node—and we copy the ma­chine code into it. Allocating this node in the file might (in rare cases) need to move some other stuff in the file around, in which case the ap­pro­pri­ate dirty” flags are set on those nodes. We al­ways set the dirty” flag for the func­tion’s node it­self, so that its re­lo­ca­tions will be ap­plied at some point.

Because of the pend­ing re­lo­ca­tions, we prob­a­bly don’t have a valid bi­nary right now—but that’s okay! When the linker thread is next idle (i.e. its work queue is empty), or at the end of com­pi­la­tion if the linker thread re­mains busy un­til then, we’ll check all of those dirty” flags and clean up af­ter our­selves. This could in­volve work such as as­sign­ing new vir­tual ad­dresses, up­dat­ing the sec­tion head­ers and pro­gram head­ers, up­dat­ing ad­dresses in the sym­bol table, and re-ap­ply­ing re­lo­ca­tions.

It might sound like that fixup” work is ex­pen­sive. Sometimes, it can be—if you’re cre­at­ing a dy­namic ex­e­cutable and the PLT has to move, that can take a mo­ment, be­cause there are usu­ally a lot of re­lo­ca­tions tar­get­ing the PLT. However, most of the time, we don’t need to move any­thing! By us­ing ex­po­nen­tial growth fac­tors on nodes (similar to how dy­namic data struc­tures like ArrayList work), we amor­tize this cost and make it ex­tremely rare in re­al­ity (at the cost of a slightly in­creased bi­nary size, which is­n’t usu­ally a ma­jor con­cern dur­ing de­vel­op­ment). This de­sign means that you might very oc­ca­sion­ally see one up­date run slightly slower than usual (maybe a few hun­dred mil­lisec­onds?), but I’ve not per­son­ally hit this a sin­gle time, de­spite us­ing in­cre­men­tal com­pi­la­tion with this linker near-daily for the past cou­ple of months.

Flush

Okay, we’ve made it to the end, and kept every­thing in­cre­men­tal along the way. Files were low­ered to ZIR with a sim­ple per-file cache; se­man­tic analy­sis of de­c­la­ra­tions kept track of a de­pen­dency graph to fig­ure out what might have changed; code gen­er­a­tion re-ran only for up­dated func­tions; and our linker wrote new code into the file with­out chang­ing any other bytes. We just have a few more loose ends to tie up.

Firstly, be­cause of how Zig’s lazy analy­sis” fea­ture in­ter­acts with in­cre­men­tal com­pi­la­tion, we need to do a graph tra­ver­sal to fig­ure out which func­tions/​de­c­la­ra­tions/​etc are ac­tu­ally ref­er­enced. It’s pos­si­ble that some­thing was ref­er­enced on a pre­vi­ous in­cre­men­tal up­date (so we com­piled it), but has since be­come un­ref­er­enced, which means we need to ig­nore any com­pile er­rors it emit­ted, not per­form sym­bol ex­ports from it, etc. There are prob­a­bly some op­ti­miza­tions you can do here, but at least right now, we just tra­verse the full ref­er­ence graph on every up­date. We can get away with this even on big pro­jects, be­cause com­put­ers are re­ally fast!

Once we’ve fig­ured out what’s ref­er­enced, we can tell the linker every global sym­bol which is ex­ported from Zig code, so that it can add any nec­es­sary en­tries to the sym­bol table. We will also re­port com­pile er­rors if there are any, and some other mis­cel­la­neous tasks like that. Finally, we call the link­er’s flush func­tion, whose job is just to do any re­main­ing link­ing work be­fore the file is closed. If the linker has any MappedFile node still marked as dirty”, we’ll need to han­dle that, but oth­er­wise we want to do as lit­tle work as pos­si­ble—re­mem­ber, any­thing we do here is go­ing to hap­pen on every up­date, so we want to keep it pretty much O(1). Therefore, all that the ELF linker re­ally does here is write out the .dynamic sec­tion, and write the en­try field in the ELF header.

We then close the file, and the com­pi­la­tion is com­plete!

Tracing an Update

Explanations are cool and all, but we can ac­tu­ally see this hap­pen­ing. Tracy is a real-time pro­filer—it’s de­signed for games, but you can in­te­grate it into any­thing. The Zig com­piler has op­tional Tracy in­te­gra­tion, en­abled us­ing a build flag. (I guess in­cre­men­tal up­dates kinda re­sem­ble frames in a video game if you squint?)

This can oc­ca­sion­ally be use­ful for var­i­ous com­piler per­for­mance analy­sis, but I ac­tu­ally find it re­ally cool to use for in­cre­men­tal com­pi­la­tion, be­cause we can see the dif­fer­ent parts of the com­piler pipeline clear as day. Let’s take a look at the Tracy out­put for a change to Fizzy, much like the changes in the video from ear­lier. This par­tic­u­lar up­date took 37ms (a lit­tle faster than the ones we saw in the video), but at first I’m go­ing to zoom in on the first 6ms or so of this 37ms up­date—I’ll ex­plain why later.

At the start we can see a flurry of ac­tiv­ity across all threads—that’s the thread pool do­ing all of the per-file work. Although we only changed one file, the com­piler does­n’t as­sume that, and in­stead checks every source file. There’s one small op­ti­miza­tion here, which is that be­cause we al­ready know which source files were in the com­pi­la­tion on the last up­date, we can guess that those files will all still be reach­able and so check for changes to all of them. That just means we don’t need to wait for the first file to be processed so that we can dis­cover its im­ports.

Next we see a good chunk of time (around 1ms) in com­puteAlive­Files. This func­tion is tra­vers­ing the graph of file im­ports to as­sign every file to a Zig module” (because it’s pos­si­ble for a file to move from one mod­ule to an­other be­tween up­dates). We also use this im­port tra­ver­sal to check whether all source files are, in fact, still in the com­pi­la­tion. If any are not—be­cause all im­ports of them were re­moved—then we’ll ba­si­cally just ig­nore those files for the rest of this up­date.

Then we have an­other mil­lisec­ond in up­dateZir­Refs. This func­tion is re­spon­si­ble for cor­re­lat­ing the old and new ZIR of any changed files, and up­dat­ing all in­ter­nal ref­er­ences to ZIR in­struc­tions to re­fer to the in­struc­tion’s in­dex in the new ZIR rather than its in­dex in the old ZIR. This is a fairly sim­ple task, but the cur­rent im­ple­men­ta­tion in­volves it­er­at­ing every ZIR in­struc­tion we hold a ref­er­ence to at all and com­pletely re­build­ing a hash map’s meta­data. This can prob­a­bly be op­ti­mized.

Now we’re done with the sin­gle-threaded per-file stuff, and we can fi­nally get onto the meat and pota­toes of the pipeline: se­man­tic analy­sis, code­gen, and link­ing. The sema_loop” zone con­tains all of the time spent in se­man­tic analy­sis—around 1.2ms. We then see that func­tion’s AIR get picked up by code­gen on a dif­fer­ent thread (the green zones named run­Code­genIn­ner), which runs for around 240us. The re­sult­ing AIR is picked up by the linker thread and emit­ted to the bi­nary—this link­ing work (the pur­ple emit­Func­tion zone and the lit­tle green zones next to it) takes around 170us. Overall, this en­tire part of the pipeline—which by far dom­i­nates cold builds—comes in at around 1.6ms for this up­date. Not bad!

After that’s all done, we’re onto flush. The lit­tle pur­ple zone at the bot­tom-right is a small bit of link­ing work the fron­tend re­quests dur­ing flush: re­gen­er­at­ing a lookup table which we use to im­ple­ment Zig’s @errorName builtin. That takes around 50us, and brings us to the end of the 6ms re­gion I’ve zoomed in on. That means it’s fi­nally time to zoom out and see what the re­main­ing 31ms are…

Basically all of the re­main­ing time is spent in one func­tion, re­solveRef­er­encesIn­ner. Remember a bit ear­lier I men­tioned do­ing a graph tra­ver­sal dur­ing flush, to de­ter­mine which Zig de­c­la­ra­tions are ref­er­enced? Well, that’s this func­tion’s job! It’s not in­ef­fi­cient by any means, but that graph is kinda big, so it’s per­haps un­sur­pris­ing that it starts to mat­ter when we’re try­ing to go fast.

On the one hand, this seems pretty silly, so much so that I se­ri­ously con­sid­ered try­ing to im­prove it be­fore putting out this blog post (after all, a 7ms time is more im­pres­sive than a 37ms time). This is am­pli­fied when you con­sider that the ref­er­ence graph did­n’t ac­tu­ally change here. So the vast ma­jor­ity of the du­ra­tion of this in­cre­men­tal up­date is be­ing spent fig­ur­ing out that a graph did­n’t change!

But ac­tu­ally, I think this is re­ally cool, be­cause it shows how much ef­fi­ciency is still left to squeeze out. This 30ms zone is re­al­is­ti­cally not a big is­sue, but we can get rid of it nonethe­less—firstly by avoid­ing re­com­put­ing this data when the ref­er­ence graph is un­changed, but also by only re­com­put­ing what we need to when the ref­er­ences do change (that prob­lem is called dynamic sin­gle-source short­est path” and is a fairly well-stud­ied prob­lem in graph the­ory).

Basically, we’re far from done on the per­for­mance front! If you want to keep up with what we’re do­ing in the fu­ture, you might con­sider adding the Zig de­vlog to your RSS reader, or check­ing out the re­lease notes when new ver­sions of Zig are re­leased.

Okay, I’ve been ram­bling about com­pil­ers for long enough; let me ac­tu­ally show you how to use this thing. I’ll as­sume you have a Zig pro­ject with a build script, and that it com­piles on a re­cent mas­ter branch build of Zig (or, if you’re read­ing this af­ter Zig 0.17.0 re­leases, that’ll also work.)

At the time of writ­ing, this will only re­ally work if you tar­get x86_64-linux, be­cause our other code gen­er­a­tion and linker back­ends are not ma­ture enough yet. The ma­jor­ity of the Zig core team runs Linux on x86_64, so by fo­cus­ing on it first, we’ve sped up our work­flows, mean­ing it’ll be faster for us to add sup­port for other tar­gets—which is now top pri­or­ity!

The bad news is that right now, this is­n’t zero-ef­fort. Eventually it will be—we’ll cache all of the com­piler state to disk and au­to­mat­i­cally re­load the last saved state when you run zig build, so in­cre­men­tal com­pi­la­tion will just hap­pen au­to­mat­i­cally—but we’re not quite there yet. However, the good news is that us­ing it to­day re­quires very lit­tle work!

The short ver­sion is that you just need to run this com­mand:

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.