10 interesting stories served every morning and every evening.

AI;DR (AI; Didn’t Read)

www.rickmanelius.com

I’m SUPER jeal­ous that I did­n’t think of this first…

Alas! Hat tip to se­clilc for tweet­ing this gem out two days ago.

lil c@se­clilc

AI;DR

(AI; did­n’t read)

4:13 PM · Aug 15, 2026 · 346K Views

83 Replies · 2.09K Reposts · 16.6K Likes

I’ve been think­ing about it ever since. Why? Because there is grow­ing grum­bling among every­one about AI writ­ing. And it’s not just oth­ers; it’s me! I am get­ting to the point where I phys­i­cally flinch (sometimes drop­ping my shoul­ders and hunch­ing, or hav­ing a slight eye twitch) when some­one I re­spect sends me un­fil­tered and unedited AI out­put.

Look, I get it. It’s Q3 2026, and we should ex­pect that every­one is uti­liz­ing AI at SOME point in their process (sourcing ideas, cre­at­ing out­lines, re­fin­ing prose, etc.).

However, I have a new pol­icy.

If you’re not both­ered enough to re­view and edit it…

…then I’m not go­ing to bother read­ing it.

Yes, there are cer­tain sit­u­a­tions in which we should ex­pect 100% AI-generated copy. Customer sup­port would be a per­fect ex­am­ple. We’re not look­ing for ar­ti­sanal did you make sure to re­set your phone” style di­a­logue.

But if you’re my col­league and we’re in a Slack dis­cus­sion and you post a wall of Claude out­put, then I’m afraid I re­ceived a dif­fer­ent mes­sage than you in­tended.

The same is true for peo­ple’s newslet­ters and so­cial con­tent. It’s your name on it; are you proud of the prose and weird AI-isms sprin­kled through­out it? If so, great. But I can ask Claude di­rectly if I wanted to.

TL;DR (too long; did­n’t read) was the so­lu­tion for so­cial me­dia.

AI;DR (AI; did­n’t read) is the so­lu­tion for AI slop.

May you em­brace this pol­icy your­self and seek out those will­ing to care enough to pri­or­i­tize a hu­man touch when they talk to you.

No posts

403 Forbidden

responsiblestatecraft.org

Error 403 Forbidden

Forbidden

Error 54113

Details: cache-iad-ki­ad7000146-IAD 1787076104 555180333

Varnish cache server

How Bluesky draws its logo on screenshots

timmarinin.net

Sometimes I take a screen­shot of a post I like, ei­ther to send it to friends/​meme chan­nel or to save a durable” copy. Like this one (I’ve cropped out the rest of the in­ter­face):

I no­ticed the Bluesky logo in the right cor­ner and thought that it was weird that the logo does­n’t bother me when I use the app. Then I looked at the post in the app again—logo was­n’t there, re­placed by the Follow” but­ton.

I re­mem­bered that a few apps hide their logo where the iPhone notch is, so that it does­n’t stick out, un­less you take a screen­shot. But here the logo is placed in the open, so how do they do it?

I tried to take an­other screen­shot, this time mid-switch­ing to the other app:

Did they some­how set up a lis­tener for two but­tons I’m press­ing to take a screen­shot and do a switcheroo at the last mo­ment? I’m not an iOS de­vel­oper, so I’m not sure what’s pos­si­ble and what is not over there.

At this point I was mildly in­trigued. Thankfully, I re­mem­bered that Bluesky app is open source (or at least the code is avail­able to look at).

The an­swer was in the file lit­er­ally called GrowthHack.tsx , in­tro­duced in January 2026 by mozzius. But it merely used a de­pen­dency, so to un­der­stand I looked into pack­age expo-pri­vacy-sen­si­tive, also by them.

The pack­age cre­ates UITextField with is­Se­cure­Tex­tEn­try prop­erty set to true and ren­ders the ac­tual con­tent (the but­ton) into that field’s  .layer. When I take the screen­shot, iOS hides this UITextField by blank­ing the layer, al­low­ing the Bluesky logo to flut­ter its wings through (it was here the whooole time). For other plat­forms it sim­ply ren­ders con­tent as-is, with­out mask­ing.

Why does­n’t it work when I switch be­tween the apps? I sup­pose that iOS takes a snap­shot it­self at the start of the ges­ture (without trig­ger­ing blank­ing), and when I do a screen­shot, there is no live UITextField in­stance to re­act to that, only the in­ert snap­shot. But once again, I’m not an iOS de­vel­oper.

Nifty trick or an abuse of API meant for pri­vacy? The peo­ple in the thread adding the be­hav­ior mostly did­n’t like it, be­fore the thread got locked. I think it’s cute.

I googled a bit, and the trick is well-known. Telegram im­ple­mented sim­i­lar thing for its secret” chats, as did Signal, so I don’t ex­pect it to be patched by Apple any time soon.

GPT-5.6 Sol - API Pricing & Benchmarks

openrouter.ai

Not avail­able in this work­space

The Amazon tax

seths.blog

It’s not tech­ni­cally a tax. Taxes pro­duce valu­able pub­lic ben­e­fits, like med­ical re­search and parks. This is sim­ply le­gal theft.

Amazon makes nearly a bil­lion dol­lars in profit from search ads. Every week. Each week, they sell mer­chants and pub­lish­ers enough search-dis­tort­ing ads to cap­ture a bil­lion dol­lars in rev­enue. Amazon makes enough in search ad rev­enue to give every sin­gle one of their em­ploy­ees a $35,000 cash bonus and still have change left over.

My pub­lisher is ter­rific, and they’re work­ing hard to in­tro­duce peo­ple to my new book. Last week, they be­gan buy­ing search ads on Amazon.

At first glance, this is com­pelling. Someone who is­n’t sure what they’re look­ing for, who is look­ing for a book or a kitchen ap­pli­ance, might find one if the right ad showed up at the right time.

But of course, that’s not what yields, or what most of the ads you see on Amazon do.

If you’re search­ing for an air fryer, Amazon al­ready knows quite a bit. They know the best-re­viewed, least-re­turned, best-priced model. The only pur­pose of the ads is to get you to pick an air fryer that is­n’t that one (or for the best air fryer, to keep you on track to buy the one you wanted in the first place). The ads make the search worse. [Cory wrote about this three years ago, and the scale has al­ready dou­bled.]

When there are plenty of ads, the maker of the best air fryer now has to bid on ads as well, if only to pro­tect the sales they were en­ti­tled to in the first place. Businesses con­tinue to buy the ads—not be­cause they’re dumb, but be­cause the sys­tem has cre­ated a sit­u­a­tion with few op­tions. Folklore im­plies that buy­ing the ads some­how shifts how search re­sponds in the long run, even af­ter the ads stop run­ning, but there’s lit­tle data to con­firm this.

Traditional ads in­crease de­mand. We see some­thing that’s clearly an ad, it might spark de­sire, and sales go up. But zero-sum search ads aren’t like that–the to­tal sales in the cat­e­gory stay the same, and mer­chants are merely com­pet­ing for a share of a sta­tic pie. This study ar­gues that an ecom­merce site with search ads ac­tu­ally sells fewer items than the same site with­out ads.

The high­est-yield­ing ad my pub­lisher has tested so far is the search Seth Godin The Knot“. It costs about a dol­lar per click. My pub­lisher is pay­ing Amazon a dol­lar to show you an ad for the book you went to buy in the first place.

Who ends up pay­ing the more than $50 bil­lion a year spent on these ads? It’s not the sell­ers. Sellers can’t make heart­felt do­na­tions for long. It’s you. By mak­ing the mar­ket­ing of prod­ucts sig­nif­i­cantly less ef­fi­cient, Amazon’s theft makes prod­ucts more ex­pen­sive or sucks the en­ergy out of the de­vel­op­ment of new prod­ucts.

It leads to two per­verse side ef­fects. First, pro­duc­ers re­al­ize that if brand rep­u­ta­tion mat­ters less than a bud­get for clicks, they will shift to shoddy and cheap ver­sions of their prod­ucts so they have a big­ger bud­get for clicks. And sec­ond, Amazon (and Google be­fore it) have an in­cen­tive to make their or­ganic search re­sults worse–giv­ing pro­duc­ers more in­cen­tive to buy more ads.

For decades, Amazon cre­ated value for con­sumers by low­er­ing the price of just about every­thing. And they opened the doors to mer­chants who did­n’t have suf­fi­cient dis­tri­b­u­tion. They claimed to be cus­tomer-cen­tric, and they were.

I don’t think they can claim this any longer. The ad sys­tem they built is­n’t il­le­gal, but it’s pretty clear who it’s for.

Amazon is steal­ing from the cus­tomers they said they were here to serve.

Google buys crashed airline Spirit’s data at auction, because AI

www.theregister.com

AI and ml

$10 mil­lion buys over 100 mil­lion emails, 30 mil­lion recorded phone calls, reams of stuff from Teams, Oracle, and SAP

Google has ac­quired the data of failed US air­line Spirit.

Spirit hit fi­nan­cial tur­bu­lence when COVID-19 blew in dur­ing 2020 and started mak­ing losses. Its bal­ance sheet never climbed back to a safe al­ti­tude and in May 2026 the air­line grounded it­self per­ma­nently.

The low-cost car­rier en­tered liq­ui­da­tion to wind it­self up and is now auc­tion­ing as­sets to raise cash and set­tle at least some of its debts.

REG AD

A court doc­u­ment [PDF] filed last week re­veals that one of the as­sets up for sale is a huge trove of dei­den­ti­fied data, which Google bid for and won for just $10 mil­lion.

REG AD

For that sum, Google bought it­self 100 mil­lion emails and 500 mil­lion items from Microsoft Teams, 17 mil­lion OneDrive files and 20.5 mil­lion items from SharePoint. The search gi­ant also now owns over 30 mil­lion recorded cus­tomer ser­vice calls, and more than 15 mil­lion cus­tomer ser­vice chat records.

600,000 ServiceNow tick­ets are an­other el­e­ment of the col­lec­tion, along with 13.7 mil­lion ac­tive emails ad­dresses from Oracle’s Responsys mar­ket­ing ap­pli­ca­tion, and de­tails of 11 mil­lion sales of in-flight Wi-Fi ser­vices.

There’s also op­er­a­tional data in the trove, de­scrib­ing over 763,000 flights, five mil­lion crew pair­ings, more than 1.2 mil­lion fuel slips, and records de­scrib­ing pur­chases of 787,452 parts.

Google has re­port­edly said it bought the data to im­prove its AI ser­vices. The un­der­bid­der was Mercor, a com­pany that pro­vides data to train AI mod­els. So clearly Spirit’s data is of value to AI com­pa­nies.

Indeed, The Register re­cently re­ported how AI ex­perts in­creas­ingly be­lieve large lan­guage mod­els are blunt in­stru­ments, and that smaller mod­els trained on spe­cific fields of knowl­edge are more use­ful in some ap­pli­ca­tions.

Google might have it­self the ba­sis for an avi­a­tion ops model, or just with all sorts of quo­tid­ian fi­nan­cial records that could be use­ful for an­other AI.

If you’ve flown Spirit and worry that Google will soon know about a testy con­ver­sa­tion you had with the air­line’s call cen­ter, you’re be­ing told not to worry. The court fil­ing says the data was dei­den­ti­fied be­fore be­ing put on sale and Google has promised to scrub any PII it finds in the trove.

The Register awaits ev­i­dence of the in­evitable SNAFUs that mean some per­sonal info ap­pears as the re­sult of a fu­ture prompt, or search.

REG AD

Fasten your seat belts! ®

Quake Shareware, a CD-ROM just a little too full

fabiensanglard.net

Aug 17, 2026

In the mid-90s the coolest thing to buy for a PC, be­sides the in­cred­i­bly ex­pen­sive Intel Pentium, was a CD-ROM drive. With their ca­pac­ity of 640 MiB (three times the stor­age of PC HDD at the time), CDs al­lowed en­thu­si­asts to step into a world of mul­ti­me­dia, made of high-res­o­lu­tion 640x480 256 col­ors palette-in­dexed pho­tos[1], VOC sound­tracks, and play with Video For Windows but­ter-smooth 12 fps 240x179 videos[2][3] last­ing up to sev­eral sec­onds.

For video game de­vel­op­ers, the CD-ROM was an odd beast. The ca­pac­ity far ex­ceeded the quan­tity of as­sets they were able to pro­duce. A few ti­tles, like 7th Guest (1993), Wing Commander III: Heart of the Tiger (1994), or Phantasmagoria (1995) in­tro­duced Full Motion Video (in a world where only part of the screen could be an­i­mated). Some added high qual­ity mu­sic. My most mem­o­rable take on the mat­ter was from id Software’s lead de­vel­oper, John Carmack.

People ex­pect CD games to have tons of dig­i­tized speech and video […] The joke here is that if we ever do a CD ver­sion of DOOM, you are go­ing to get the game and The Making of DOOM a one hour fea­ture film. John Carmack (Jan 94) for ATARI EXPLORER ONLINE

By June 1996, af­ter three years of hard work, id Software had com­pleted their next ti­tle, Quake. As for their pre­vi­ous ti­tle, they were go­ing to re­lease both a share­ware ver­sion and a full ver­sion of their game. Since it used a mere 22 MiB of stor­age, peo­ple at id Software had the idea of lever­ag­ing the re­main­ing ca­pac­ity of a CD-ROM. Why not in­clude en­crypted ver­sions of the full id cat­a­logue of games? Not only this would cut out the mid­dle­men, it would give in­stant ac­cess to gamers with a sim­ple phone call and a credit card.

The con­cept was im­ple­mented. The CD was an­nounced[4] on July 3, 1996 and re­leased on August 30th[5]. The hacker group GNOMON re­leased Quakecrk.zip only 39 days later[6]. The archive con­tained QCRACK.EXE, a tool al­low­ing to de­crypt every sin­gle game on the CD-ROM.

Quake’s share­ware re­tail ex­per­i­ment had proved dis­as­trous. In the­ory id was go­ing to cut out re­tail­ers by al­low­ing gamers to buy the share­ware and then call an 800 num­ber to place an or­der and re­ceive a pass­word that would un­lock the rest of the game.

But gamers wasted no time hack­ing the share­ware to un­lock the full ver­sion of the game for free. Worse, all the mun­dane as­pects of dis­tri­b­u­tion and or­der ful­fill­ment were spin­ning out of con­trol. In a des­per­ate mea­sure, id tried to put the brakes on the re­tail share­ware, but it was too late. They were stuck with al­most 150,000 CDs sit­ting in a ware­house. David Kushner (Masters of Doom)

But gamers wasted no time hack­ing the share­ware to un­lock the full ver­sion of the game for free. Worse, all the mun­dane as­pects of dis­tri­b­u­tion and or­der ful­fill­ment were spin­ning out of con­trol. In a des­per­ate mea­sure, id tried to put the brakes on the re­tail share­ware, but it was too late. They were stuck with al­most 150,000 CDs sit­ting in a ware­house. David Kushner (Masters of Doom)

So what hap­pened? Let’s dive in!

Thanks to usenet archives of rec.games.com­puter.quake.misc, we have de­tailed dis­cus­sions[7] of how it worked. A gamer could go to any of the hun­dreds of CompUSA/Computer City stores and buy the CD for $9.95.

The pack­ag­ing was ac­tu­ally pretty high qual­ity for a share­ware prod­uct. On my copy, a sticker on the front clar­i­fies this is the Shareware ver­sion” with in­struc­tions to call 1 – 800-669 – 9342 (or 1 – 800-ID-GAMES) to un­lock the full game.

The phone num­ber is still ac­tive to­day. However you don’t reach the un­lock cen­ter since CompUSA went out of busi­ness. The Superstore chain closed be­tween 2007 and 2008. Instead of an un­lock op­er­a­tor we get an au­to­mated mes­sage to sell us el­derly stuff.

Upon con­tact­ing the op­er­a­tor, users were to also com­mu­ni­cate a SOURCE CODE. It played no part in gen­er­at­ing the Unlock code. It may have been a way for the dis­trib­u­tors to claim a trans­ac­tion fee. Browsing eBay, I found many with names in­dica­tive of past/​pre­sent re­tail­ers. 12-BSTBY BestBuy, 24-CCITY Computer City, 22-CUSA CompUSA, 88, 11 – 1111, 38-EB Electronic Boutique, 44-FTRSP Future Shop (Canada), 34-EGGH Egghead Software, and 56-MCTR Media Play/ Musicland.

The first con­tact was sleek. The GUI was well done. Users could jump di­rectly into Quake Shareware but they could also click on QUAKE UNLOCK.

The CD-ROM also fea­tures an ID STUFF sec­tion, al­low­ing to browse the cat­a­logue of id games and un­lock any of them. Several ver­sions of DOOM are there, along with HEXEN, and HERETIC.

Once the un­lock process was started, the GUI gen­er­ated a CODE NUMBER (which I call CHALLENGE) that was to be com­mu­ni­cated to a Service Agent over the phone. Upon pay­ing the fees, an UNLOCK CODE NUMBER (which I call SERIAL) would be re­ceived. Checksums on both num­bers mit­i­gated is­sues re­lated to this prim­i­tive land­line mode of com­mu­ni­ca­tion.

To avoid re­play at­tacks, the CHALLENGE changes every time the pro­gram is run and ro­tates every 5 min while the GUI is ac­tive.

The screen even had the sig­na­ture provoca­tive tone of the early days of id Software (“those who are too cheap”). Note that there was also a warn­ing that users should back up the game once un­locked. Since the CHALLENGE in­cluded some ran­dom­ness, there was no way to reuse a SERIAL.

At first sight, the process looked solid. Users had to call an un­lock ser­vice to ob­tain a pass­word. And only with that pass­word could the fi­nal un­lock be com­pleted. So what went wrong?

The tool pow­er­ing the lock/​un­lock was pro­vided by TestDrive Corp’s (archived www.test­drive.com). The idea was to give play­ers a way to try-before-you-buy” and al­low them to im­me­di­ately pur­chase the full ver­sion of a pro­gram.

Their en­crypter was ca­pa­ble of denaturing” an .EXE ex­e­cutable. It re­placed the first 32 KiB with a cus­tom header (attempting to run a de­na­tured ex­e­cutable dis­played This ap­pli­ca­tion has been dis­abled”), re­named the file to .MJ3, en­crypted the orig­i­nal header as a .ST3, and is­sued a seed.

Normal EXE

Denatured MJ3

Chunk ST3

Secret seed

Let’s peek in­side the Quake share­ware CD with a filemap I gen­er­ated (Split View rec­om­mended). In that tree, we can see one MJ3 file for each game avail­able. We can also see all the ST3 files in­side the PAGEMKR archive. Everything is there to renature” an MJ3 back into an EXE game in­staller, ex­cept for the se­cret seed” that is as­surely de­rived from the SERIAL.

Described as is, there is no flaw in this process. The se­cret seed comes from the un­lock server, tied to a CHALLENGE/SERIAL that could not be reused. But the hacker team GNOMON found a way.

Released on 10/08/96 (gnomon.nfo), only 39 days af­ter Quake re­tail share­ware CD hit the stores, QCRACK.EXE was a tool able to gen­er­ate a SERIAL au­to­mat­i­cally given the CHALLENGE.

What mem­bers of GNOMON group fig­ured out was that the SERIAL re­ceived over the phone con­tained no se­cret at all. It was just a proof of pay­ment.

The QUAKE un­lock pro­gram FLOW.EXE that ships on the CD is ca­pa­ble of gen­er­at­ing the SERIAL from the CHALLENGE on its own. All it does is check that its own lo­cally-gen­er­ated SERIAL and the SERIAL en­tered by the user match! The en­tire pro­tec­tion mech­a­nism re­lies on se­cu­rity by ob­scu­rity.

The pipeline from CHALLENGE to SERIAL is con­vo­luted but was re­versed in 2016 by rmolina[8].

The 11-digit CHALLENGE is split into a 4-digit GAME-ID, and a 7-digit num­ber re­sult­ing in an OFFSET, and a DEPTH.

The GAME-ID in­dexes an en­crypted data­base SKU.17, which gives a co­de­name (e.g.: doom2).

The CODENAME al­lows to re­trieve a 512-byte DOC file (e.g.: DOOM2.DOC) in­side the FLOWLIB.LIB archive.

Mixed with the CODENAME and the string Testdrive Corp.”, the DOC trans­forms a 508-byte hard-coded table into a table of 254 16-bit val­ues unique to the ti­tle.

DEPTH and OFFSET walk that table back­wards, XOR-ing DEPTH to gen­er­ate a sin­gle 16-bit value MEM.

The SERIAL is then un­lock = ((reverse7(GAME_ID) + MEM + 0x18) & 0x7F) + 0x83 * ((MEM ^ 0x1EA3) + 0x1700A1) printed with a lead­ing B.

The more I re­searched the mat­ter, the more it looked like who­ever was in charge had no time to pol­ish the re­sult.

Never at­tribute to mal­ice what you can at­tribute to stu­pid­ity. And never at­tribute to stu­pid­ity what you can at­tribute to time pres­sure. Fab’s Razor

Digging in­side Quake share­ware CD re­veals many more is­sues.

Some parts of the un­lock sys­tem look like they were never tested. Final DOOM can­not be un­locked by call­ing the un­lock cen­ter be­cause of a bug. The GAME-ID for Final Doom” is 12. There is a typo in SKU.17 which makes GAME-ID 12 cor­re­spond to CODENAME Final” (with a cap­i­tal F). This makes the SERIAL gen­er­a­tion re­trieve the wrong DOC. The cor­rect value was final”. This cre­ated an only-too-fa­mil­iar sit­u­a­tion where il­le­git­i­mate users en­joyed a bet­ter ex­pe­ri­ence than pay­ing cus­tomers.

The step-by-step sum­mary men­tions an en­crypted SKU.17 file. There is a plain-text ver­sion, com­pletely un­en­crypted, of the very same file named SKU.TXT in­side the FLOWDIR archive. There are many more TXT files match­ing their .17 en­crypted ver­sions (PRODUCT.TXT, EXE.TXT).

Several files are tem­po­raries (DM.TMP), ed­i­tor ar­ti­facts (FLOWWORK.BAK), or not used at all (ENCRYPT.EXE).

The li­brary for­mat is not en­crypted or scram­bled. Figuring out the .DIR for­mat of­fered low re­sis­tance and easy ac­cess to all DOC files nec­es­sary to gen­er­ate SERIALs.

References

^[1]MediaPack 10-CD Roms ^[2]MediaPack Tropical Rainforest ^[3]MediaPack Wild Places ^[4]Quake’ Is Here ^[5]When will it be avail­able in stores (rec.games.computer.quake.misc) ^[6]gnomon.nfo ^[7]Question about Quake Shareware CD (rec.games.computer.quake.misc) ^[8]Quake, TestDrive y Qcrack

VRAM Management Part 2: Beyond the Limits of Physical VRAM

pixelcluster.dev

Earlier this year, I blogged about work I did to im­prove VRAM man­age­ment for games. Now, af­ter many months of float­ing around in mail­ing lists, the ker­nel patches are fi­nally merged up­stream and queued for Linux 7.3! Hooray!

To cel­e­brate, let’s look a bit deeper at one sen­tence I wrote in my pre­vi­ous post:

[Games] should per­form much more sta­ble - as long as the game it­self does­n’t use more VRAM than you ac­tu­ally have.

[Games] should per­form much more sta­ble - as long as the game it­self does­n’t use more VRAM than you ac­tu­ally have.

So, one may ask: What if they do, in fact, use more VRAM than you ac­tu­ally have?

Typical ex­pec­ta­tions for this seem to be that once this hap­pens you’re pretty much screwed. Games will start crash­ing left and right, per­for­mance plum­mets to un­playable lev­els, a good gam­ing ex­pe­ri­ence be­comes im­pos­si­ble.

But is that re­ally just an un­avoid­able fact of life? What re­ally makes run­ning out of VRAM suck so hard? And, most im­por­tantly: How can we make it suck as lit­tle as pos­si­ble?

Setting ex­pec­ta­tions

In the­ory, run­ning out of VRAM should ex­clu­sively be a per­for­mance is­sue, not a sta­bil­ity one. Support for over­com­mit­ting VRAM has ex­isted for as long as GPU dri­vers have: If the dri­ver over­com­mits VRAM, you are gen­er­ally al­lowed to re­quest as much VRAM as you’d like, and you’ll get as much as the ker­nel dri­ver de­cides it can fit into the phys­i­cal mem­ory that ex­ists on GPU.

On the per­for­mance side, the big-pic­ture rea­son for bad per­for­mance when you run out of VRAM is fairly sim­ple. As soon as the game re­quests more VRAM than is phys­i­cally pre­sent, some of the game’s mem­ory will have to be moved/​evicted to CPU RAM in­stead. For the GPU, ac­cess­ing CPU RAM is much slower than VRAM: Not only is CPU RAM slower than a ded­i­cated GPUs VRAM in gen­eral, all mem­ory ac­cesses also have to go over the PCI bus. The PCI bus adds la­tency and is typ­i­cally also the lim­it­ing fac­tor in band­width when fetch­ing from CPU mem­ory.

Due to PCI speed lim­i­ta­tions, there are some truly un­avoid­able per­for­mance con­straints when over­com­mit­ting VRAM. Assuming the GPU is hooked up via a PCIe 4.0x16 con­nec­tion, you get a lit­tle less than 32GiB/s of band­width. Each mil­lisec­ond, that PCIe bus can trans­fer ~32.2MiB of data. For a min­i­mum fram­er­ate of 30 frames per sec­ond (33.3ms per frame), the ab­solute max­i­mum amount of data the GPU is able to ac­cess is ~1,075.5MiB, a tiny bit over 1GiB of data. In other words, if so much mem­ory gets evicted that the GPU needs to fetch more than 1GiB from evicted mem­ory in one sin­gle frame, it is sim­ply im­pos­si­ble to still hit 30 FPS.

Not all mem­ory is equal

At the same time, just read­ing a lit­tle bit of CPU mem­ory on the GPU is not im­me­di­ately a death sen­tence for per­for­mance. In fact, GPU dri­vers some­times de­cide to let things like com­mand buffer data and re­lated al­lo­ca­tions live in CPU RAM even when there’s plenty of VRAM avail­able! Whenever the GPU ex­e­cutes these com­mands, it has to ac­cess CPU mem­ory, and yet in these cases every­thing runs com­pletely fine. So what makes these ac­cesses dif­fer­ent - why are they fine and yet run­ning out of VRAM seems cat­a­strophic?1

One thing that in­flu­ences the cal­cu­lus sig­nif­i­cantly is caching. Since the ac­cess la­tency in case of a cache hit is the same re­gard­less of whether the cached mem­ory lives on CPU or GPU, the high ini­tial cost of fetch­ing over the PCI bus can be amor­tized by cache hits (to some ex­tent). We can es­ti­mate la­tency dif­fer­ences be­tween fetch­ing CPU RAM and VRAM by writ­ing mi­crobench­marks that mea­sure ac­cess la­tency for dif­fer­ent buffer sizes (using an ad­ver­sar­ial ac­cess pat­tern to min­i­mize cache hi­trates as far as pos­si­ble). The re­sult you get may look some­thing like this (captured on RDNA3):

As ex­pected, if the buffer fits into L2 (or any higher-level cache), ac­cess la­ten­cies are ex­actly the same for mem­ory backed by CPU RAM and mem­ory backed by VRAM, be­cause the data gets fetched di­rectly from cache in ei­ther case. At a size of 6MB (the L2 cache size on RDNA3), CPU mem­ory la­ten­cies go up to about 2400 cy­cles per ac­cess, while de­vice mem­ory la­ten­cies stay within the same rough ball­park. Note that VRAM ac­cesses also go through the Infinity Cache, but CPU mem­ory ac­cesses do not (they hit PCIe di­rectly on an L2 miss). I sus­pect this is be­cause the Infinity Cache sits di­rectly on top of VRAM, so any ac­cess that does­n’t hit VRAM also does­n’t reach the Infinity Cache.

Obviously, mem­ory does­n’t start off with be­ing cached any­where, so the first ac­cess will still have con­sid­er­ably higher la­tency. Also, los­ing the Infinity Cache def­i­nitely hurts as well: PCIe fetches seem to have some­where around 7.3x as much la­tency than an Infinity Cache hit, and around 4.6x as much la­tency as a fetch from VRAM. This in­creased la­tency needs re­ally high cache hi­trates to fully amor­tize the cost of go­ing over PCIe. That means there is only a small set of use cases where us­ing CPU mem­ory has such mi­nus­cule slow­downs that you’d ac­tively de­cide to use it in fa­vor of VRAM when you have the choice. When you’re evict­ing mem­ory from VRAM, there will al­most un­avoid­ably be at least some de­gree of slower per­for­mance.

Still, even though slow­down is un­avoid­able, there is go­ing to be mem­ory where evic­tion mat­ters more and mem­ory where evic­tion has a lesser ef­fect on over­all perf. Memory that is ac­cessed in very cache-friendly ways is not af­fected by the slow­down of CPU RAM as much. If the ac­cess pat­terns aren’t cache-friendly but the mem­ory is­n’t ac­cessed very of­ten, things may also still be fine since the GPU only rarely needs to ac­tu­ally fetch data from CPU RAM. There might be many mem­ory al­lo­ca­tions where the GPU will only ac­cess a small part of the to­tal al­lo­ca­tion size, and never even read the rest. If these al­lo­ca­tions were to be evicted, you might evict mul­ti­ple GiBs of data, but still re­main well be­low the 1GiB hard limit of data that is ac­tu­ally ac­cessed per frame.

All of these vari­ables make it sur­pris­ingly hard to pre­dict how per­for­mance ac­tu­ally pans out in prac­tice when mem­ory is be­ing evicted. But in short: Depending on how much the evicted mem­ory gets ac­cessed and how well these ac­cesses cache, you might just be able to run out of VRAM with­out (completely) ru­in­ing per­for­mance!

Confronting re­al­ity

We’ve the­o­rycrafted our­selves all the way to­wards hav­ing per­for­mant VRAM over­com­mit­ment now. Great! Let’s just boot up SteamOS, start some game and crank up the setti- radv/​amdgpu: Not enough mem­ory for com­mand sub­mis­sion.

oh.

As it turns out, run­ning out of VRAM in prac­tice does carry plenty of sta­bil­ity is­sues with it.

This er­ror is­n’t quite like a reg­u­lar couldn’t al­lo­cate, out of mem­ory” er­ror, though. Note that the mes­sage specif­i­cally com­plains about com­mand sub­mis­sion: RADV prints this mes­sage when the ker­nel re­turns -ENOMEM when try­ing to sub­mit com­mands2, but merely sub­mit­ting com­mands does not al­lo­cate any new re­sources! All the com­mand buffers were al­lo­cated in ad­vance, and clearly their al­lo­ca­tion suc­ceeded. Even though all mem­ory was suc­cess­fully al­lo­cated, us­ing it in a GPU sub­mis­sion sud­denly re­sults in out of mem­ory” er­rors be­ing thrown.

It’s time for an­other ker­nel ad­ven­ture! Surely get­ting the ker­nel to ac­cept the sub­mis­sion can’t be that hard - af­ter all, the ker­nel al­ready ac­cepted all the al­lo­ca­tions3!

The hor­rors of ker­nel lock­ing

One thing the amdgpu dri­ver has to do on every sub­mis­sion, be­fore it can di­rect the GPU to start ex­e­cut­ing com­mands, is to make sure that all mem­ory that may po­ten­tially be ref­er­enced by the GPU com­mands is ac­ces­si­ble. With more mod­ern bind­less graph­ics APIs, you have to as­sume all al­lo­cated mem­ory may at some point get ref­er­enced. Therefore, amdgpu will try to make sure all al­lo­cated mem­ory is also ac­ces­si­ble.

Each mem­ory al­lo­ca­tion car­ries in­for­ma­tion about which type of mem­ory (for our pur­poses here, sys­tem RAM or GPU VRAM) it can be prop­erly ac­cessed from. Most al­lo­ca­tions can be ac­cessed from ei­ther CPU RAM or VRAM, and amdgpu will be happy with the mem­ory al­lo­ca­tion be­ing in ei­ther of these mem­ory types. Some al­lo­ca­tions, how­ever, have to be placed in VRAM and VRAM only. If these mem­ory al­lo­ca­tions have been evicted to sys­tem RAM be­cause some other ap­pli­ca­tion al­lo­cated VRAM in the mean­time, amdgpu will have to move them back into VRAM. Because there is no free VRAM avail­able at all, mov­ing the al­lo­ca­tion back re­quires evict­ing some­thing else. For some rea­son, that failed and the ker­nel re­ported an out-of-mem­ory con­di­tion.

In or­der to ex­plain why evict­ing some­thing ran­domly fails, we’ll have to take a small de­tour to look at how the ker­nel han­dles (CPU-side) lock­ing for GPU al­lo­ca­tions. In or­der to evict a mem­ory al­lo­ca­tion, you have to ac­quire a lock as­so­ci­ated with that al­lo­ca­tion. However, dur­ing a sub­mis­sion, you also have to lock every al­lo­ca­tion that’s ref­er­enced in a sub­mis­sion, to pre­vent some other ap­pli­ca­tion from mov­ing the al­lo­ca­tion some­where else while you’re busy prepar­ing GPU work. But if an­other GPU sub­mis­sion is do­ing the same thing con­cur­rently, you can end up in a sit­u­a­tion like this:

If one sub­mit wants to evict an al­lo­ca­tion that an­other sub­mit has al­ready locked, but that other sub­mit also needs to lock an al­lo­ca­tion from the first one to make progress, we have a text­book ABBA dead­lock con­di­tion.

But fear not, the ker­nel knows how to de­tect and re­solve dead­locks! The de­tails about how dead­lock de­tec­tion works are ex­plained in this ker­nel doc­u­men­ta­tion page, but in very broad strokes, the ker­nel as­so­ci­ates lock­ing op­er­a­tions with a transaction” (which ba­si­cally just keeps track of which locks were ac­quired). If two trans­ac­tions would dead­lock, one of the trans­ac­tions is marked as wounded”, and the next time it tries to ac­quire a lock, the -EDEADLCK er­ror is re­turned. This er­ror re­quests the trans­ac­tion to be aborted: All locks ac­quired dur­ing the trans­ac­tion should be re­leased, and the trans­ac­tion is restarted from scratch. In the con­text of com­mand sub­mis­sion, this just means the dri­ver will restart the process of go­ing over all mem­ory al­lo­ca­tions and mak­ing sure they’re ac­ces­si­ble.

So where’s the catch? There is­n’t one. This ap­proach is rock solid and works re­ally well.

At least as long as it’s ac­tu­ally im­ple­mented every­where.

In the graph­ics sub­sys­tem, the gritty in­ter­nals of the wound-abort-retry loop are ab­stracted us­ing a small helper li­brary called dr­m_exec. Instead of hav­ing to man­u­ally track which al­lo­ca­tions are locked, and re­lease the locks once you run into -EDEADLCK, you sim­ply use the dr­m_ex­ec_lock­_obj helper. If you study the lock­ing code in TTM, the shared Linux GPU mem­ory man­age­ment layer, you will no­tice a pro­found lack of us­age of dr­m_exec.

Instead, there even is a com­ment not­ing that -EDEADLCK will cause evic­tion to fail. There we go, we found our is­sue! As soon as this dead­lock con­di­tion is en­coun­tered be­cause of in­tense mem­ory pres­sure dur­ing com­mand sub­mis­sion, the ker­nel bails out and re­jects the sub­mis­sion in­stead of retry­ing.

There al­ready are some patch­sets to hook up the dr­m_exec helper in TTM, sent all the way back in 2024, but those never made it in for a few rea­sons, among which were some re­main­ing bugs that had­n’t been fig­ured out. My work had been cut out for me here: Rebase the patch­set on top of my ker­nel ver­sion and fig­ure out what those re­main­ing bugs are.

Rebasing the patch­set was­n’t too much of a has­sle, and fig­ur­ing out the bugs only took one sin­gle week of in­tense suf­fer­ing with games ran­domly hang­ing 3 min­utes into heavy VRAM con­tention. Not the worst!

I tried re­send­ing the patch­set with fixes for all bugs I found in the hopes it would get in this time, but there’s go­ing to be more work need­ing to be done with it be­fore it can be merged.

Now that run­ning out of VRAM at least won’t crash your apps at ran­dom, we can at least prop­erly crank up the set­tings and look at perf. The ini­tial re­sult gave me an ab­solutely glo­ri­ous per­for­mance graph like this:

Hold On Where Did All The Perf Go

Figuring out why per­for­mance is so garbage re­quires fig­ur­ing out what the sys­tem is ac­tu­ally do­ing that’s this slow. For broad what’s the ker­nel dri­ver do­ing??” ques­tions like that, I like us­ing gpu­vis. gpu­vis uses ker­nel tra­ce­points to build a time­line of things that hap­pened (including GPU work sub­mis­sion started/​stopped”, from which the time taken for each sub­mis­sion can be in­ferred).

Booting up gpu­vis with a trace taken while the sys­tem is run­ning out of VRAM, the time­line shows a sit­u­a­tion like this:

Turns out, most of that time is­n’t ac­tu­ally spent on han­dling the sub­mis­sion (that’s the gfx_0.0.0 ac­tiv­ity), but in­stead mov­ing around mem­ory in prepa­ra­tion for that sub­mis­sion (sdma0 ac­tiv­ity)!

The rea­son why there are so many buffer moves all the time be­comes more ob­vi­ous if you use gpu­vis’s event list, to­gether with a fil­ter to show only cap­tured move events for a par­tic­u­lar buffer ob­ject (I chose one at ran­dom here, most buffer ob­jects have a sim­i­lar pat­tern):

The list shows quite clearly that con­tend­ing processes (in this case, gamescope and the game it­self) will con­stantly take turns evict­ing and mov­ing back the same piece of mem­ory, over and over. That’s re­ally bad! And it’s very rem­i­nis­cent of some­thing I wrote in my first blog­post:

Generally, two com­pet­ing ap­pli­ca­tions can be ex­pected to roughly take turns ex­e­cut­ing GPU work - first one ap­pli­ca­tion sub­mits work, then the other, then the first again, and so on. With that ap­proach, mem­ory would keep be­ing moved back and forth af­ter every sin­gle sub­mis­sion. One ap­pli­ca­tion gets kicked out and im­me­di­ately moved back in, kick­ing the other out (which moves mem­ory back in the next step). All this mov­ing ended up with worse per­for­mance than if the mem­ory had never been moved in the first place.

Generally, two com­pet­ing ap­pli­ca­tions can be ex­pected to roughly take turns ex­e­cut­ing GPU work - first one ap­pli­ca­tion sub­mits work, then the other, then the first again, and so on. With that ap­proach, mem­ory would keep be­ing moved back and forth af­ter every sin­gle sub­mis­sion. One ap­pli­ca­tion gets kicked out and im­me­di­ately moved back in, kick­ing the other out (which moves mem­ory back in the next step). All this mov­ing ended up with worse per­for­mance than if the mem­ory had never been moved in the first place.

This de­scribed an old is­sue where overly ag­gres­sive VRAM al­lo­ca­tion would lead to ping-pong-like moves hap­pen­ing con­stantly. But that is­sue had since been fixed by sim­ply not try­ing to claim VRAM when there is­n’t any free VRAM left, and the ker­nel only started be­ing some­what ag­gres­sive when I im­ple­mented VRAM pro­tec­tion with dmem cgroups. Obviously, this must have rein­tro­duced the ping-pong­ing some­how.

Conceptually, the de­sign of the dmem cgroup VRAM pro­tec­tion should never re­sult in ping-pong moves, be­cause the ker­nel is only sup­posed to evict mem­ory that does not have any cgroup VRAM pro­tec­tion as­so­ci­ated with it. Without any VRAM pro­tec­tion, you should typ­i­cally not be al­lowed to evict pro­tected VRAM.

The sin­gle ex­cep­tion to this rule is mem­ory that ab­solutely has to live in VRAM for things to work prop­erly. These kinds of mem­ory al­lo­ca­tions are al­ways al­lowed to be moved to VRAM to en­sure sys­tem sta­bil­ity. Typically, al­most noth­ing com­ing from an ap­pli­ca­tion is re­ally re­quired to live in VRAM for cor­rect op­er­a­tion, but there is one buffer ob­ject com­ing from an ap­pli­ca­tion that does: The buffer con­tain­ing im­age data to be scanned out to the dis­play4.

Display hard­ware is funky

Not only does the dis­play hard­ware like scanned-out im­ages to be in VRAM, it also com­pletely skips past the GPUs vir­tual mem­ory ar­chi­tec­ture and works with phys­i­cal ad­dresses ex­clu­sively. In con­se­quence, scanned-out im­ages also have to be con­tigu­ous in phys­i­cal mem­ory.

With vir­tual mem­ory and the power of page ta­bles, typ­i­cal ap­pli­ca­tion buffers are only con­tigu­ous in vir­tual mem­ory, and may be scat­tered around all over phys­i­cal mem­o­ry5. The first page of a buffer at vir­tual ad­dress 0x5000 may be mapped in the page ta­bles to point to phys­i­cal ad­dress 0x1234000, but the sec­ond page at vir­tual ad­dress 0x6000 might point to phys­i­cal ad­dress 0x4321000, some­where com­pletely dif­fer­ent!

Here is a di­a­gram vi­su­al­iz­ing the map­ping of vir­tual al­lo­ca­tions to phys­i­cal ones in case where there is a lot of frag­men­ta­tion (which typ­i­cally is the case when you’re very low on VRAM):

The ar­rows show page table map­pings to phys­i­cal mem­ory seg­ments for the dif­fer­ent seg­ments of the first al­lo­ca­tion. They’re left out for all other al­lo­ca­tions for read­abil­ity.

If you’re al­lo­cat­ing dis­play scanout data, this frag­men­ta­tion is not an op­tion as the phys­i­cal mem­ory has to be con­tigu­ous. This has very, very un­for­tu­nate in­ter­ac­tions with evic­tion of other data specif­i­cally. Let’s as­sume the scanout data has al­ready been evicted, but now it’s time for that data to be scanned out, so it has to be moved back into VRAM.

Simply evict­ing one buffer won’t be suf­fi­cient, even if that buffer is the same size as the dis­play scanout data, be­cause evict­ing it does not re­sult in enough con­tigu­ous phys­i­cal space to place the scanout data in! To make mat­ters worse, the evic­tion al­go­rithm does not take into ac­count phys­i­cal mem­ory con­straints at all. It is a very sim­plis­tic loop along the lines of

while (true) { evict(getLeas­tRe­cent­lyUsed­Buffer()) if (tryAllocate(newBuffer) == SUCCESS) break; }

Using this al­go­rithm (assuming the al­lo­ca­tions are arranged in LRU or­der), even if you evict the first 3 al­lo­ca­tions (green, blue, and red), there won’t be a large enough space to hold the scanout buffer! Even the largest pos­si­ble free space is ever so slightly too small, as is vis­i­ble in this up­dated di­a­gram:

To find a large enough phys­i­cally con­tigu­ous mem­ory re­gion in our ex­am­ple, every sin­gle al­lo­ca­tion in VRAM would end up be­ing evicted! In real-world sce­nar­ios, I ob­served up to 4GiB of VRAM be­ing nuked just to make space for scanout im­ages (which are ~32MiB of pixel data per im­age for a R11G11B10 pixel for­mat). That’s go­ing to hurt real hard! Simply the act of mov­ing all that data out from VRAM would al­ready cost at least ~130ms, ac­cord­ing to the PCIe trans­fer rate es­ti­mated ear­lier.

Throwing heuris­tics at the prob­lem

While scanout is def­i­nitely the most egre­gious fail­ure case here, this is­sue is more gen­eral: There are al­ways go­ing to be cer­tain mem­ory al­lo­ca­tions that will be moved to VRAM over and over, po­ten­tially kick­ing out some mem­ory that an ap­pli­ca­tion might pre­fer to stay in VRAM. Resisting this and try­ing to move the evicted mem­ory back in will most likely back­fire.

Even though dmem cgroup pro­tec­tion is not a com­plete so­lu­tion to this prob­lem, it does re­duce the prob­lem scope by a lot. With cgroup pro­tec­tion, you can be sure that any ran­dom app won’t try to kick out im­por­tant game re­sources willy-nilly. Any mem­ory that does get moved back into VRAM by force prob­a­bly has a good rea­son to be in VRAM. Therefore, even with dmem cgroup pro­tec­tion, we should be care­ful and not try to re­claim evicted mem­ory back by force.

With some it­er­a­tive test­ing, I think I’ve ar­rived at a set of heuris­tics that work rea­son­ably well for most cases a game would en­counter in the wild (not be­ing too ag­gres­sive when stuff gets evicted by im­por­tant sys­tem al­lo­ca­tions is one thing, but it also needs to be rea­son­ably quick at re­claim­ing evicted mem­ory if e.g. the game is paused and the Steam menu runs in­stead, evict­ing lots of game mem­ory, and then the game is re­sumed).

The heuris­tics work some­thing like this:

When the ker­nel de­tects an ap­pli­ca­tion’s mem­ory is be­ing evicted, it en­ters a hard throt­tle” phase for a few mil­lisec­onds. During this phase, it does not try mov­ing any mem­ory for that app back into VRAM what­so­ever (as long as all mem­ory can be prop­erly ac­cessed, of course).

After this pe­riod, it switches a soft throt­tle” phase, dur­ing which it may re­claim free space by mov­ing things back into VRAM, but does not try evict­ing any mem­ory that other apps have al­lo­cated. This pe­riod may last up to a few sec­onds, to make ex­tra sure every­thing reached a sta­ble state.

If the soft throt­tle” phase has com­pleted with­out any fur­ther mem­ory be­ing evicted again, the sys­tem is as­sumed to have reached a fairly sta­ble state and re­stric­tions on evict­ing other ap­pli­ca­tions’ mem­ory are re­moved.

IME, this achieves an ac­cept­able bal­ance be­tween not shoot­ing one­self in the foot with over­ag­gres­sive evic­tion of other apps, while still re­cov­er­ing rea­son­ably fast when lots of your mem­ory was sud­denly evicted, for ex­am­ple be­cause the game was paused and the user browsed around on Steam in­stead of play­ing.

Getting some­where

With those heuris­tics in place, let’s fi­nally try crank­ing up the set­tings for real this time.

I ended up go­ing with Indiana Jones: The Great Circle, since it con­ve­niently ex­poses a set­ting for stream­ing pool sizes that you can mess with to mod­ify VRAM con­sump­tion pretty much di­rectly.

Lo and be­hold, even if the set­tings are turned up to a some­what ridicu­lous point, where the game re­quests 9GiB of 8GiB VRAM (aka. a whole 1GiB of over­com­mit­ted game re­sources liv­ing in CPU mem­ory), per­for­mance is­n’t cra­ter­ing into obliv­ion any­more! A 19.6ms per frame av­er­age is what I’d still call per­fectly playable.

I can also bump the set­tings to even more ridicu­lous lev­els and dou­ble the amount of over­com­mit­ted mem­ory, with the game re­quest­ing 10GiB of VRAM on this 8GiB sys­tem (and thus 2GiB of re­sources be­ing over­com­mit­ted). Frametime vari­ance goes up quite a lot at this point, with spikes reach­ing above 33.3ms hap­pen­ing fre­quently. The over­all av­er­age is around 29.8ms which is­n’t the worst, but es­pe­cially paired with the vari­ance, this would start be­ing no­tice­able in game­play.

While this is al­ready a huge step for­ward, we aren’t quite there yet. The ex­pe­ri­ence un­der VRAM over­com­mit can some­times still be a bit hit-or-miss, and fram­e­times may no­tice­ably vary de­pend­ing on which ob­jects in the game you’re look­ing at.

Remember that for ac­tu­ally good evic­tion per­for­mance, it mat­ters a lot how the evicted mem­ory is used by the GPU. Right now, this is­n’t taken into ac­count at all! If we were able to base our evic­tion de­ci­sions more on how well the ap­pli­ca­tion’s ac­cesses work with CPU mem­ory, a lot of this vari­ance might sim­ply dis­ap­pear.

Handing over the con­trols

The com­pli­cated thing about the ap­pli­ca­tion’s mem­ory ac­cess pat­terns is that they are only re­ally known to the ap­pli­ca­tion. Therefore, the dri­ver is­n’t re­ally able to take them into ac­count as-is. Ideally there would be some API where the ap­pli­ca­tion can sup­ply hints to the dri­ver about how well a par­tic­u­lar mem­ory al­lo­ca­tion is suited to be­ing evicted.

Something ex­actly like vk­Set­De­vice­Mem­o­ryPri­or­i­tyEXT! The VK_EXT_pageable_device_local_memory ex­ten­sion pro­vides pre­cisely what we need here, by al­low­ing ap­pli­ca­tions to com­mu­ni­cate any pri­or­ity they want for any piece of de­vice mem­ory they want. As long as ap­pli­ca­tions pro­vide rea­son­able hints through this ex­ten­sion, im­ple­ment­ing pri­or­i­ti­za­tion in the ker­nel and then uti­liz­ing app-pro­vided pri­or­i­ties has the po­ten­tial to sta­bi­lize things by a lot!

Hooking up pri­or­i­ties in the ker­nel turns out to be a lot less of an is­sue than you might ex­pect. The ker­nel al­ready main­tains a Least-Recently-Used list of mem­ory al­lo­ca­tions that, on evic­tion, are tra­versed in or­der. For each en­try on that LRU list, evic­tion is at­tempted un­til there is enough free space for what­ever the evic­tion was for.

This LRU list pro­vides a good heuris­tic for which ap­pli­ca­tion’s mem­ory should be evicted first. Applications that haven’t sub­mit­ted any­thing in a long while are un­likely to need the mem­ory soon, and since their mem­ory is Not Recently Used, it will ap­pear early in the LRU list and be evicted first.

When an ap­pli­ca­tion uses a set of buffers, that set of buffers is moved to the very end of the LRU list in one bulk. However, the or­der of al­lo­ca­tions within that bulk is not ex­plic­itly con­trolled at all. That means once the ker­nel closes in on some ap­pli­ca­tion to evict its mem­ory, which spe­cific pieces of mem­ory get evicted is more or less un­de­fined6. A sim­pli­fied vi­su­al­iza­tion could look some­thing like this:

If the ker­nel walks the LRU list like this, it would evict the buffer with a pri­or­ity value of 2 first, even though there are much lower-pri­or­ity buffers else­where in the LRU list. If only the first buffer of pri­or­ity 2 gets evicted, things might be okay, but if the highly im­por­tant buffer with pri­or­ity 4 ends up evicted as well, there are likely go­ing to be prob­lems.

Given that we al­ready know spe­cific pri­or­i­ties for the in­di­vid­ual al­lo­ca­tions, this LRU list is a very sim­ple place to in­te­grate them. It’s as sim­ple as or­der­ing the list en­tries within a sin­gle ap­pli­ca­tion by their pri­or­i­ty7:

Now, when the ker­nel goes over the LRU list to find some­thing to evict, the very first thing it will find and try to evict are the low­est-pri­or­ity buffers. The high­est-pri­or­ity buffers are last in the list, and thus only get evicted when evict­ing all the lower-pri­or­ity buffers was not enough.

Memory pri­or­ity adop­tion in apps

Unfortunately, not all ap­pli­ca­tions ac­tu­ally set pri­or­i­ties via VK_EXT_pageable_device_local_memory. As for na­tive Vulkan ap­pli­ca­tions, I haven’t ob­served any idTech game us­ing the ex­ten­sion di­rectly, at least :/

The D3D side looks a lot bet­ter, be­cause vkd3d-pro­ton al­ready uses VK_EXT_pageable_device_local_memory when avail­able, and trans­lates both the ID3D12Device::MakeResident/ID3D12Device::Evict API calls as well as pri­or­i­ties set via ID3D12Device1::SetResidencyPriority to pri­or­ity val­ues set us­ing the Vulkan vk­Set­De­vice­Mem­o­ryPri­or­ity com­mand. Lots of D3D12 games uti­lize at least one of these APIs, so the hints these games pro­vide will now be uti­lized.

I don’t have su­per solid num­bers for how much mem­ory ex­actly is over­com­mit­ted by most D3D12 apps, as they don’t typ­i­cally ex­pose the to­tal amount of VRAM they re­quest in an easy-to-ac­cess way like idTech’s per­for­mance over­lay does. However, prop­erly hon­or­ing mem­ory pri­or­i­ties gen­er­ally seems to have a good chance to im­prove the ex­pe­ri­ence. Performance gen­er­ally ap­pears more sta­ble over time (because you’re not re­ly­ing on luck with which buffers the ker­nel evicts as much). In some spots I had a good com­par­i­son point at, I sus­pect it in­creased per­for­mance com­pared to the ker­nel evict­ing ran­dom things by up to 30% in the very best case - but again, take this num­ber with a moun­tain of salt as it de­pends al­most en­tirely on luck with re­gards to evic­tion.

Conclusion

When all is said and done, how well does run­ning out of VRAM hold up?

I’d say it’s quite al­right! In many cases, you may be sur­prised how much per­for­mance you can re­tain even when evict­ing a gi­ga­byte or more of mem­ory! Then again, that’s of course a rather op­ti­mistic case, and the wrong thing end­ing up in CPU RAM can very quickly cause very sig­nif­i­cant slow­downs. Eviction is tricky to get just right, and to an ex­tent, per­for­mance will al­ways be dragged down. If a game is strug­gling to hit 30fps even with every­thing in VRAM, need­ing to evict some­thing on top of all that could some­times just un­avoid­ably re­sult in that 30fps tar­get be­ing missed.

Regardless, what I hope this blog­post can demon­strate is that even if you end up with some mem­ory evicted to sys­tem RAM, the slow­down can be man­age­able. There’s mea­sures that dri­vers (particularly, the ker­nel dri­ver) can take to make over­com­mit work as fast as pos­si­ble, and even ap­pli­ca­tions can do their part in co­or­di­nat­ing with the dri­ver stack to mit­i­gate the ef­fects of their mem­ory be­ing evicted. With every­thing in place, VRAM over­com­mit is­n’t re­ally as big of a deal as one may think it is at first sight.

All the work I de­scribed here has al­ready been re­leased in SteamOS for some time now (it’s both in Stable and Preview. As long as your sys­tem is up-to-date, it’s good to go!).

A note on up­stream­ing

Of course, I’m al­ready work­ing on up­stream­ing all this work so it’s avail­able to every­one! However, there’s a lot of mov­ing parts and a lot of deep refac­tors of some pretty core con­cepts at play here, so it will likely need time to cook be­fore every­thing is merged up­stream.

At the same time, I don’t want to put up a blog­post talk­ing about lots of cool code just to fin­ish it with actually you can’t see for your­self, go wait un­til it’s all up­stream lol”, ei­ther.

As a mid­dle ground, I have re­based the ker­nel work onto a re­cent up­stream ver­sion of the ker­nel and pub­lished a git branch here. While it should the­o­ret­i­cally yield sim­i­lar ef­fects, it did not go through as rig­or­ous test­ing the SteamOS ker­nel did. There will likely be bugs and in­sta­bil­i­ties that weren’t there in the SteamOS ver­sion. Use at your own risk, ba­si­cally. I don’t ex­pect to be main­tain­ing this branch in any sig­nif­i­cant ca­pac­ity, as I’d rather fo­cus on get­ting the patches into up­stream prop­erly.

In or­der to pass through ap­pli­ca­tion pri­or­ity hints to the ker­nel, you will also need a cus­tom Mesa branch I pushed here. Similar con­sid­er­a­tions as the ker­nel branch ap­ply here, as well.

Questions of my own

While I would claim to have a fairly good overview of the dri­ver side of mem­ory man­age­ment at this point, I am not very fa­mil­iar with how ap­pli­ca­tions de­cide on sup­ply­ing mem­ory man­age­ment heuris­tics in­ter­nally, at all. I would sus­pect op­ti­miz­ing cases where you’ve al­ready run out of VRAM is­n’t ex­actly the top item on de­vel­oper TODOs (who knows, maybe the mem­ory scarcity is chang­ing that? :P), so maybe there’s some un­ex­plored room for per­for­mance im­prove­ments there?

An Update on Leaving Gmail for Fastmail

moddedbear.com

It’s been a few months since my post on leav­ing Gmail which sparked a lot of dis­cus­sion on Hacker News. Email is­n’t the most ex­cit­ing topic in the world, but I fig­ured I should give a lit­tle up­date on how things have been go­ing af­ter the move to Fastmail since so many peo­ple saw the orig­i­nal post.

The quick ver­sion is that I think I picked right with Fastmail. You can find cheaper email hosts out there. I’m sure they’re great too, but I’m re­ally happy with Fastmail’s value con­sid­er­ing all I’m get­ting.

Inbox or­ga­ni­za­tion

I ended up start­ing fresh in­stead of for­ward­ing my Gmail to my new ad­dress. This was def­i­nitely the right move for me for rea­sons I’ll get into.

My biggest con­cern mov­ing away from Gmail was my in­box or­ga­ni­za­tion. Gmail does a pretty al­right job of au­to­mat­i­cally sort­ing out pro­mo­tional emails from no­ti­fi­ca­tion emails from hu­man-sent emails and so on. It’s the only rea­son my in­box has re­mained some­what man­age­able over the years de­spite no ef­fort on my part.

The so­lu­tion for this that I landed on is sub­do­main ad­dress­ing and so far it’s been keep­ing me even bet­ter or­ga­nized than be­fore. I’ve up­dated my im­por­tant ac­counts with unique or cat­e­gory-based sub­do­main ad­dresses, then any mail sent to those ad­dresses au­to­mat­i­cally gets moved to a match­ing folder with­out me even hav­ing to cre­ate any rules.

If I had for­warded my Gmail to the new Fastmail ad­dress, or­ga­ni­za­tion would have been a whack-a-mole rule cre­ation game.

Going and up­dat­ing my ac­counts with a new ad­dress was eas­ier than I was ex­pect­ing. It does help that I’m keep­ing the Gmail ad­dress ac­tive as a junk in­box that I still oc­ca­sion­ally check though, so I’ve only had to up­date the ac­counts that I ac­tu­ally care about. It was­n’t more than ten or so ac­counts, and I’m get­ting to the rest grad­u­ally as I re­mem­ber about them.

Multiple do­mains

Fastmail lets you con­nect some­thing crazy like 100 do­mains to your ac­count. That was a nice bonus for me be­cause I was able to add a new ad­dress us­ing my blog do­main, some­thing I’d been want­ing to do for a while any­way.

This is prob­a­bly a co­in­ci­dence, but I’ve been get­ting way more email replies to my posts since I up­dated my con­tact page with my new ad­dress on my blog do­main. I used to have a Proton ad­dress I cre­ated specif­i­cally for email replies listed there. It’s been a fun new mo­ti­va­tor for blog­ging, and it’s also helped me re­mem­ber to reach out when I read a cool post.

Since I’m talk­ing about cus­tom do­mains al­ready, here’s some­thing you might want to know if you’re think­ing of mak­ing a move sim­i­lar to mine. You’ll prob­a­bly find your sent emails greylisted” by your re­cip­i­ents’ email providers if your do­main has been freshly reg­is­tered — re­gard­less of your email host. For ex­am­ple, emails sent from my brand new do­main were get­ting de­layed by a few hours when sent to Gmail ad­dresses. Messages from my blog’s do­main, which has been around for a while, were get­ting de­liv­ered im­me­di­ately. It took a day or two for Gmail to fully trust the new do­main and re­move the de­lay.

Masked email

This is some­thing I’m us­ing more than I thought I would. You can gen­er­ate ran­dom­ized ad­dresses that route to your in­box so you don’t have to share your real ad­dress. It’s nice to use for ser­vices that you may want to eas­ily block, since you can flip a tog­gle and start send­ing all mail re­ceived at a masked ad­dress to the trash.

It’s been smooth sail­ing

There’s a bunch of other lit­tle things I’m lik­ing about Fastmail too. Their apps are solid. Their doc­u­men­ta­tion is good. I haven’t re­ally found any­thing not to like yet.

The main thing I want you to take away is that the switch has been much smoother than I ex­pected. If you’re think­ing of mak­ing a switch too — to any­where, not just Fastmail — I don’t think you have a whole lot to worry about. Especially if you move to an ad­dress on your own do­main, then any other moves in the fu­ture will be even eas­ier.

JP

Using the railway network as a flatbed scanner

philo.gay

Using the rail­way net­work as a flatbed scan­ner

August 17th, 2026 — 4,600 words

Over the past few months, I’ve been work­ing on us­ing an in­dus­trial lin­ear scan­ning cam­era to take very wide pho­tos out of trains and fer­ries. Getting it work­ing has been quite the chal­lenge, but I think the re­sults speak for them­selves.

taken on the San Francisco to Oakland ferry in February 2026 (56,894x2,048 pixel grayscale im­age); scroll to zoom in and click and drag to move

More pic­tures are on dis­play in the gallery.

I pre­sented a talk on this pro­ject at EMFcamp 2026, which you can watch be­low or read on for the same story in more de­tail:

What am I even look­ing at?

The process of cap­tur­ing an im­age like the one of the con­tainer port above.

The cam­era is pointed out of a mov­ing ve­hi­cle and is con­stantly cap­tur­ing a sin­gle ver­ti­cal line kinda like these grayscale ones in the di­a­gram, but a lot thin­ner. As the cam­era moves, what ex­actly it sees is chang­ing. If I cap­ture the lines from the cam­era quickly enough and stitch them to­gether, I can pro­duce a com­plete-look­ing im­age. It’s a bit more com­pli­cated than that and get­ting the re­sults look­ing good was rather tricky, but that’s the main idea be­hind it.

Background and Prior Art

Back in the 1990s, dig­i­tal cam­era sen­sor tech­nol­ogy had­n’t caught up to the size and ef­fec­tive res­o­lu­tion of medium and large for­mat film, so dig­i­tal scan­ning backs were de­vel­oped. They cap­ture a high-res­o­lu­tion im­age with­out need­ing a gi­ant grid of pix­els by mov­ing a sin­gle line of pix­els (or three lines for color) across the frame. In the in­ter­ven­ing years, im­age sen­sors have got­ten pretty big (there’s even one that cov­ers 4x5″ large for­mat nowa­days), but this ap­proach is still cheaper to build for large for­mats than a gi­ant sen­sor.

I’d been think­ing about build­ing my own dig­i­tal scan­ning back for my large for­mat cam­era for a while, but I’ve never quite got­ten around to it be­cause build­ing some­thing to mount prop­erly on my cam­era seemed too daunt­ing. (Buying one could have been an op­tion, but ones from the 1990s still go for thou­sands of dol­lars on ebay and re­quire re­con­struct­ing a com­put­ing en­vi­ron­ment of a sim­i­lar vin­tage to use.) Late last year, I was watch­ing a video on Gigawipf’s medium for­mat scan­ning cam­era build and sud­denly thought: what if the en­tire cam­era moved and the sub­ject did­n’t?” and de­cided to give it a shot.

Loading film into my large for­mat cam­era on top of a moun­tain in Vermont be­cause I’m al­ler­gic to do­ing pho­tog­ra­phy in a nor­mal way. (The re­sult­ing pic­tures from that trip are here.)

I found some pre­vi­ous pho­tos in the same vein (the Scannoramic pro­ject, John Hikerbiker’s ex­per­i­ment, Daniel Lawrence Lu’s re­ver­sal of his sta­tion­ary cam­era, and Martin Liebscher’s very in­ter­est­ing film shots), but the re­sults seemed like they could be im­proved upon. Surely tak­ing the speed of mo­tion into ac­count and get­ting cleaner re­sults would­n’t be too hard, right?

Slit Scanning My Sofa

On the night I thought up this big scan­ner” con­cept, I had to give it a shot. It was a bit late to go out and catch a train, so I scanned my sofa in­stead.

I set my phone on my of­fice chair and slowly pushed it along as it cap­tured a video. I then wrote some re­ally slap­dash code (which I am choos­ing not to share here to pro­tect my read­ers) to grab the left­most col­umn (a slit”) of each frame and com­bine them into an im­age.

My com­ments in­cluded lyrics from Future Me Hates Me” by The Beths, which be­came some­thing of a self-ful­fill­ing prophecy when I started writ­ing a post­proces­sor for the next ver­sion of the cam­era loosely based on that code and cursed my de­ci­sions.

It looks vaguely like my sofa, but it’s rather squished and the art on the wall is un­in­tel­li­gi­ble. Surely I can do bet­ter.

I messed around with the post­pro­cess­ing and dou­bled every col­umn, which makes it look less squished, but it’s still a mess be­cause I was­n’t push­ing the chair at a par­tic­u­larly con­sis­tent speed.

I knew from the start that I’d need to mea­sure the speed some­how, but I was naïvely hop­ing that I would­n’t need to mea­sure it that well and could sim­ply fudge it. This im­age, how­ever, shows that even small vari­a­tions of speed mat­ter. This was my first glimpse into how much of a pain deal­ing with speed would turn out to be.

For my next trick, I took a ride on the MBTA or­ange line. I taped my old phone to the seat to use its ac­celerom­e­ter and held my cur­rent phone to the win­dow, mak­ing sure to turn the frame rate up all the way to 60 fps.

The ac­celerom­e­ter data was­n’t very use­ful and was even less so when I took an in­te­gral to get ve­loc­ity. If I re­mem­ber cor­rectly, y was the axis of the train’s move­ment, but the data is so noisy that the train was ap­par­ently mov­ing back­wards at the end.

The re­sult looks in­ter­est­ing, though, but I def­i­nitely need more lines if I want a prop­erly in­tel­li­gi­ble im­age.

While I was get­ting ready for EMFcamp, I no­ticed an­other talk on the sched­ule by Tim Jacobs (better known on­line as mitx­ela) that was also about slit scan cam­eras and started to worry we’d both done the same thing. (He ran up to me af­ter my talk to tell me he’d also wor­ried this.) His talk started in the same way, with tak­ing a slit from a video, but he ended up mak­ing re­ally cool and trippy an­i­ma­tions by go­ing through every pos­si­ble slit po­si­tion for a given video.

Industrial Linear Camera

My source for more lines per sec­ond ended up be­ing the Basler ruL2048 – 19gm, de­signed to be pointed at fast-mov­ing con­veyor belts. The oddly-cap­i­tal­ized name comes from its abil­ity to read out its 1x2048 pixel im­age sen­sor just shy of 19,000 times per sec­ond.

These ca­pa­bil­i­ties come at a price, how­ever; brand new, the man­u­fac­tur­er’s low­est-spec cur­rent mod­els go for around US$700. Thankfully for my wal­let, I found mine on ebay for a tenth of that.

The price is also mea­sured in light. since it’s cap­tur­ing so quickly (the slow­est ex­po­sure time is 1/100s), it needs a lot of light. I can only shoot in the day­time, and all but the bright­est sta­tions and tun­nels are off lim­its to me.

To my sur­prise, hav­ing dealt with ven­dor­ware be­fore, Basler just let me down­load the SDK with­out a sup­port con­tract or proof of pur­chase. The most re­cent ver­sion also still sup­ports this cam­era from 2013, which is less sur­pris­ing but is still con­ve­nient.

The cam­era com­mu­ni­cates with the com­puter over a gi­ga­bit eth­er­net link and the soft­ware finds it au­to­mat­i­cally as long as the rel­e­vant in­ter­face is set up for APIPA ad­dresses (169.254.0.0/16). I could set sta­tic ad­dresses for both ends, but I’m only us­ing one cam­era at a time, so I haven’t been both­ered to change it.

With sur­pris­ingly lit­tle swear­ing at the SDK, apart from some com­plaints about their use of shut­ter time rather than shut­ter speed and what a frame” is on this cam­era, I put to­gether a pro­gram that grabbed buffers of pix­els and wrote them to disk.

This was my first im­age out of the cam­era us­ing my own code, and I think it looks pretty good for just mov­ing it free­hand.

The setup and me­chan­i­cal de­sign

In or­der to take it on a train with­out need­ing to have three hands to hold it, I needed a way to mount it to a tri­pod. I ended up de­sign­ing a rather util­i­tar­ian case with a heat-set in­sert in the bot­tom that my friend Brooke 3D-printed for me. Buying the parts for it gave me an ex­cuse to fi­nally make an or­der from McMaster-Carr and feel like a real en­gi­neer.

My first at­tempt did­n’t come out be­cause it turns out there’s these things called manufacturing tol­er­ances” that I com­pletely for­got about.

Oops, that’s a bit too small.

In ret­ro­spect, I prob­a­bly should’ve stuck the sen­sors on with some­thing other than blue painters’ tape, but it’s held on pretty well. Going clock­wise around it, the boards are:

6 de­gree of free­dom ac­celerom­e­ter/​gyro, which can be used with some maths to to get the speed

GPS, which did­n’t end up work­ing as well as I’d hoped be­cause the trains in Boston are a bit too good at block­ing GPS sig­nals

SAMD21 mi­cro­con­troller to shunt the data back off to the lap­top

The lens on the front is a Vivitar 28mm f/​2.8 that I al­ready had for a more nor­mal cam­era, with an adapter from Pentax K to the C-mount screw on the cam­era. Since some of the things I’m try­ing to shoot with it are kinda tall, its field of view worked out pretty well.

The whole thing is pow­ered off a USB-C bat­tery bank and there’s also eth­er­net and USB ca­bles run­ning to my lap­top, so it’s a bit of a ca­ble spaghetti mon­ster when in ac­tion.

With the sen­sors at­tached, I could fi­nally give them a try.

Both of these im­ages are the same cap­ture of wav­ing the cam­era back and forth out my win­dow, but the top one is the raw im­age and the bot­tom one is tak­ing ac­celerom­e­ter move­ment into ac­count. As you can see, us­ing the ac­celerom­e­ter makes every­thing look a lot closer to nor­mal and less stretched. (I’ll ex­plain more of how this works in a bit in the Postprocessing Hell sec­tion.)

Boston Attempts

Once I had every­thing as­sem­bled, it was time to take it on a train.

I started off on the MBTA Orange Line, since it’s the clos­est to me, but as you can see, the re­sults weren’t that good. Previewing what was com­ing out of the cam­era was a pain, so I kinda had to guess on the ex­po­sure, and I def­i­nitely guessed wrong. The post­pro­cess­ing code I wrote did­n’t work very well and every­thing was stretched and com­pressed a bit weirdly.

I went out again on a day with nicer weather and had some bet­ter luck with the ex­po­sure, al­though I think I messed up the fo­cus a bit. Unlike the at­tempt with my phone cam­era, the text on sta­tion signs is pretty leg­i­ble, so I’m def­i­nitely get­ting enough lines.

I’m par­tic­u­larly happy with how this one of the Longfellow Bridge from Boston to Cambridge came out. This one is in the gallery if you’d like to take a closer look.

Capture (in far too much de­tail)

When I was tak­ing these early pic­tures in Boston, I was us­ing a tool from the cam­era ven­dor called Pylon to pre­view. The black hor­i­zon­tal sec­tion was all I could see of the im­age at any one time, and it’s ro­tated 90° from how I’d like to see it. Dialing in the ex­po­sure in it, re­leas­ing its grip on the cam­era, and then start­ing my own code back up be­fore the train started mov­ing again was a right pain that I had to do some­thing about.

My first at­tempt at a GUI of my own used OpenCV high­gui, which did­n’t re­ally work for this. It re­quires a 1 ms de­lay af­ter each frame, which is fine for slower cam­eras, but would cause me to miss 4 en­tire lines (250 μs each at the shut­ter speeds I’m usu­ally us­ing) every dis­play frame (256 lines).

I ended up us­ing Dear ImGUI in­stead, which worked nicely with the frame ac­qui­si­tion loop I al­ready had. Out of the ap­prox­i­mately two dozen back­ends the li­brary sup­ports, I picked GLFW (“girl love for work­groups”, to quote a mes­sage from a friend at the time) and OpenGL3, prob­a­bly be­cause of the girl love” quip, al­though I’m not cer­tain.

I wrote most of the GUI in a sin­gle sleep­less night in Toronto where ro­tat­ing the im­age felt like the sin­gle hard­est prob­lem in com­puter sci­ence. (There’s def­i­nitely a few things I can do to im­prove the im­ple­men­ta­tion I set­tled on, but it runs well enough for the time be­ing.) Unfortunately, the pic­tures I took in Toronto did­n’t re­ally come out, but at least they were ex­posed cor­rectly.

I en­coun­tered some strange bugs while adding a his­togram for the im­age.

Getting the ac­celerom­e­ter data proved to be some­thing of a pain. my first ver­sion sent read­ings as text over se­r­ial, which turned out to be very com­pu­ta­tion­ally in­ten­sive on the mi­cro­con­troller. (Converting float­ing point num­bers to strings and then as­sem­bling strings is very ex­pen­sive, even on a rel­a­tively pow­er­ful SAMD21 mi­cro­con­troller that has thirty-two en­tire bits.) I de­cided to move the con­ver­sions over to my lap­top, which has the pro­cess­ing power to han­dle them with ease, but this came with prob­lems of its own.

A very frus­trat­ing de­bug­ging ses­sion.

The ac­celerom­e­ter mea­sure­ments were sent as raw float­ing point num­bers, but GPS data was still in NMEA sen­tences and switch­ing be­tween them re­quired send­ing fixed byte se­quences and hop­ing that noth­ing got mis­in­ter­preted as those se­quences. (Nothing in a NMEA sen­tence should come across as 0x11 0x11 0x11 0x11, my ac­celerom­e­ter data start se­quence, but it’s not com­pletely im­pos­si­ble for ac­celerom­e­ter data to con­tain 0x22 0x22 0x22 0x22, my NMEA string start se­quence.)

I also ran into is­sues where not flush­ing the se­r­ial port at the right time ru­ined an en­tire day’s shots. Thankfully, I was cap­tur­ing on the Mattapan Line in Boston, and I can pretty eas­ily go back and try again. That seam” in the im­age is where it lost all se­r­ial data for around half a sec­ond, which is an eter­nity in line cam­era time. The soft­ware kept wait­ing for an­other ac­celerom­e­ter sam­ple that never came be­cause the se­r­ial port buffer was full.

See It, Say It, Sorted

The fully as­sem­bled cam­era looks like a sus­pi­cious mess, and the witch us­ing it does­n’t look much less so.

The cam­era is­n’t usu­ally held to­gether with this much tape, but I’d for­got­ten to bring the tri­pod mount plate on that trip to Montréal. Would you trust her to bring strange equip­ment onto your train?

Despite Boston’s his­tory of po­lice over­re­ac­tion to harm­less elec­tron­ics pro­jects, I worry the least about be­ing ar­rested on the MBTA. People here tend to mind their own busi­ness and have never called the cops on me. The po­lice also don’t ride the trains much, pre­fer­ring to ha­rass peo­ple in sta­tions in­stead.

I’m less used to how things work in other cities, so I only take the cam­era out when rid­ing with a friend to look out for trou­ble (and some­times to lis­ten to dis­patch ra­dio).

So far, I’ve only been seen, not said or sorted. I’m cross­ing my fin­gers that this does­n’t change as I take the cam­era more places.

On my trip to Montréal, I was stopped by se­cu­rity in Gare Centrale and in­formed that tripods weren’t al­lowed and asked, au franglais, whether I was record­ing or tak­ing a pic­ture. Rather than try to an­swer that philo­soph­i­cal ques­tion in a lan­guage I don’t speak, I just said désolé” a few times and put away the tri­pod, which seemed to be suf­fi­cient.

The pic­tures I took in Montréal are here in the gallery (images 2 and 3) if you’d like to see them.

Postprocessing Hell

Capturing im­age and ac­celerom­e­ter data turned out to be the easy part com­pared to post­pro­cess­ing and mak­ing the im­ages ac­tu­ally look good.

The cam­era cap­tured some­where around 4,000 lines per sec­ond, so I had more lines than I needed in every cap­ture and had to pick which ones ac­tu­ally mat­ter.

What hap­pens if I take too few lines (Autoroute 10 in Brossard, Québec out of the win­dow of the REM A) Jumping be­tween lines too quickly looks ar­ti­fi­cial and wrong, like is vis­i­ble at the wa­ter­line in this al­bum cover edit of an early ver­sion of the Oakland ferry photo.

To de­cide which lines to use, I ended up us­ing the speed, as mea­sured by an ac­celerom­e­ter, but this came with sev­eral prob­lems.

Firstly, ac­celerom­e­ters don’t ac­tu­ally mea­sure speed. They mea­sure ac­cel­er­a­tion, the rate of change of ve­loc­ity. By tak­ing an in­te­gral, I can get ve­loc­ity, but that’s rel­a­tive to an ini­tial value. I can usu­ally as­sume that the start­ing speed is at a sta­tion and is thus zero, but I can’t be cer­tain of that. If it is­n’t zero, I have no good way of know­ing the cor­rect value and just have to guess un­til I find one that smells right.

Secondly, as shown in this di­a­gram, the ac­celerom­e­ter I’m us­ing is only mea­sur­ing so quickly. The cam­era is grab­bing lines maybe 4 times faster than it, so every few lines have to share a speed value. It also is­n’t very con­sis­tent be­cause my mi­cro­con­troller code is­n’t as fast as it could be, so this could cause ir­reg­u­lar­i­ties in the fi­nal im­age. How many ac­cel­er­a­tion mea­sure­ments there are or aren’t also changes how ac­cu­rate the in­te­gral is, which cre­ates more prob­lems.

You might re­mem­ber that I men­tioned putting a GPS re­ceiver on the cam­era ear­lier, and while I did do that, it was­n’t very use­ful. It did­n’t get a sig­nal on most of the trains I tried it on, and when it did man­age to get one, it only read 10 times a sec­ond, which cov­ers 400 en­tire lines out of the cam­era. If it worked a bit more con­sis­tently, it could be use­ful for cor­rect­ing for in­te­gra­tion er­ror us­ing a Kálmán fil­ter, but that’s a prob­lem for when I have bet­ter GPS data.

Even if my speed mea­sure­ment is per­fect, I still have the prob­lem of par­al­lax, where things closer to the cam­era ap­pear to move faster than things fur­ther away. This is in­de­pen­dent of op­ti­cal fo­cus, which I usu­ally set at in­fin­ity.

-u 10 dis­tance units per pixel

This prob­lem can be dealt with by chang­ing how much dis­tance each pixel rep­re­sents. Lower val­ues em­pha­size things closer to the cam­era more, while higher ones make the back­ground more vis­i­ble. You can give this a try by mov­ing the slider!

Each of these im­ages is same size (10,000 pix­els wide by 2048 tall, scaled to fit your browser) and each in­cludes every­thing from the pre­vi­ous by virtue of cov­er­ing more of the cap­ture. The units are ar­bi­trary and don’t mea­sure real dis­tance (I could make it ac­tual me­ters per pixel, but I don’t see a point to that.)

The cam­era and soft­ware have no idea what I want to focus” on, so I make the artis­tic de­ci­sion and man­u­ally pick that for each seg­ment of the im­age and stitch the seg­ments to­gether to get the pic­tures in the gallery. I tested dif­fer­ent val­ues for dis­tance per pixel and start­ing ve­loc­ity of each seg­ment and then stuck them to­gether in GNU IMP to pro­duce the fi­nal im­ages. The as­sem­bled im­ages of­ten be­came too big for the 65,535x65,535 max­i­mum size of a JPEG file, so I used the good old TIFF for­mat. (The PNG spec­i­fi­ca­tion al­lows sim­i­larly large im­ages in the­ory, but the soft­ware I had to hand seems to like big TIFFs bet­ter than big PNGs.)

Notes on dis­tance per pixel (u) and start­ing ve­loc­ity (v) val­ues for each part of a few im­ages

The pro­gram that takes the ac­celerom­e­ter data into ac­count for every line of the im­age is called grind­stone, since it grinds multi-gi­ga­byte raw cap­tures down into smaller us­able im­ages. My first ver­sion was loosely based on my very bad slit scan code from ear­lier and was ex­tremely slow, tak­ing hours to cap­ture a min­utes-long cap­ture. It would of­ten fail to save af­ter run­ning for hours be­cause the re­sult­ing im­age was too big for the JPEG for­mat, and de­bug­ging it was an ab­solute pain.

I ended up nerd­snip­ing my friend Maddie into rewrit­ing grind­stone in id­iomatic NumPy, to make the math­e­mat­i­cal op­er­a­tions that were go­ing on clearer (she in­sists that all the op­er­a­tions were al­ready in the orig­i­nal, and her changes were along the lines of transforming it into a mag­i­cal girl”). Maddie would later split this ver­sion into a perhaps slightly ov­erengi­neered” pipeline of sev­eral dif­fer­ent stages, mak­ing it eas­ier to ex­per­i­ment, and swap in dif­fer­ent op­er­a­tions, out­put strate­gies, and the like. Thanks to her help, I’ve been able to try dif­fer­ent com­bi­na­tions of pa­ra­me­ters much more eas­ily, and get re­sults I’m much hap­pier with.

Color Hell

In April, my friend Ari and I went for a ride on the Mattapan Line as the leaves were com­ing in on the trees. The pic­tures I took did­n’t come out due to a cap­ture soft­ware bug (see Capture) and I haven’t got­ten around to go­ing back yet, but it left us with the thought that color line cam pho­tos might look cool, es­pe­cially in au­tumn.

While brows­ing ebay late one night, I found a very good deal on a color line cam­era of the same gen­er­a­tion as the mono­chrome one I al­ready had (the Basler ruL2098 – 10gc, 3x2098 pix­els at around 10,000 lines per sec­ond). After a bit of dis­as­sem­bly (it came to me in the hous­ing it was used in on some fac­tory line) and swap­ping the lens mount over, the cam­era was ready me­chan­i­cally.

I ended up putting red, green, and blue stripes on it so I could tell the cam­eras apart with­out tak­ing the lens off or squint­ing at tiny text on the la­bel.

The cap­ture soft­ware side was­n’t that much harder, al­though I did have to fix a bunch of as­sump­tions about the size of each line and redo the ro­ta­tion for the GUI

Progress of get­ting the color cap­ture work­ing

Thanks to the very mod­u­lar way that Maddie rewrote grind­stone, adding sup­port for color im­ages was­n’t too dif­fi­cult, al­though we did have to fix some strange-look­ing bugs.

The train was mov­ing so slowly and in­con­sis­tently in this pic­ture that in­te­gra­tion er­ror piled up and grind­stone cal­cu­lated that the cam­era was mov­ing back­wards and jumped to var­i­ous pre­vi­ous points in the cap­ture.

With cap­tur­ing and pro­cess­ing im­ages mostly work­ing, more prob­lems be­came ap­par­ent. The most vis­i­ble one is that leaves are all far brighter than they should be. This hap­pens be­cause the color cam­era is sen­si­tive to in­frared light on all three chan­nels. (If it was only sen­si­tive to it on the red chan­nel, the leaves would look red­dish, but the com­bi­na­tion of all three chan­nels’ IR with the strong vis­i­ble green leads to the green­ish white in this pic­ture.) The mono­chrome cam­era is sen­si­tive to IR too, but it does­n’t mat­ter be­cause it’s just one chan­nel and vis­i­ble light com­pletely drowns it out.

(Diagram taken from the cam­er­a’s man­ual)

I will ad­mit the ef­fect does look pretty good in the right light. This pic­ture taken in Manchester-by-the-Sea, north of Boston, is both grayscale and col­or­ful at once. (Read on to learn what the color fringes in the back­ground are.)

I solved this with an UV and IR cut fil­ter that only passes light be­tween 400 and 700 nm, which is close enough to the hu­man vis­i­ble spec­trum that every­thing looks right. This is the first big cap­ture I took with it, and I only needed to ad­just the col­ors min­i­mally in post.

I also tried a fil­ter that only passes light longer than 720 nm (I’ve had quite in­ter­est­ing re­sults with it and IR-sensitive film), and I’m def­i­nitely go­ing to try tak­ing more pic­tures with it in the fu­ture.

The next prob­lem is that some things end up with weird red, green, and blue fringes, es­pe­cially sub­jects that are fur­ther from the cam­era or mov­ing faster. They turn out to be in­her­ent to how this cam­era sen­sor works. Red, green, and blue are each sep­a­rate ver­ti­cal lines (instead of a Bayer fil­ter), and thus can’t see ex­actly the same thing at the same time. The fringes come from when just one line sees some­thing, and it’s par­tic­u­larly no­tice­able with bright sub­jects. They’re di­ag­o­nal and not per­fectly ver­ti­cal be­cause the cam­era it­self is­n’t per­fectly ver­ti­cal. (I try to get it close, but there’s only so much I can do on a mov­ing train.)

From the cam­er­a’s man­ual; the man­u­fac­turer pro­vides for­mu­las that can be used with the op­ti­cal mag­ni­fi­ca­tion fac­tor of the lens and the ex­act speed to coun­ter­act it, but I don’t have (relative) speed es­ti­mates for the sub­ject.

I cor­rect for it for a given sub­ject by shift­ing the red and blue chan­nels to line up with the green chan­nel. Since the lines are evenly spaced, I can shift by the same amount in op­po­site di­rec­tions rather than hav­ing to mea­sure sep­a­rate off­sets for each chan­nel. In the­ory, I could de­cide how far to shift by cor­re­lat­ing bright­ness shifts across chan­nels, but at pre­sent, I do it man­u­ally. Separation be­tween chan­nels is vis­i­ble on the sail­boat’s masts, and I cor­rected for it by shift­ing the red chan­nel 10 pix­els right and the blue chan­nel 10 pix­els left. Color fringes are still vis­i­ble in the back­ground be­cause it’s much fur­ther away than the sail­boat and thus has a faster an­gu­lar ve­loc­ity; I could shift and cor­rect for it, but the sail­boat would look much worse.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.