10 interesting stories served every morning and every evening.

Bento Slides

bento.page

OverpAId — Fire Your CEO. Hire The Future.

overpaid.lol

Introducing the world’s first Chief Executive Replacement Engine

Your CEO costs $22,000,000 a year.We cost $4,699. Once.

OverpAId is an Artificial Intelligence built from the ground up to do your CEOs en­tire job — strat­egy, vision,” mo­ti­va­tional all-hands emails — bet­ter, faster, and with­out ever once ask­ing the board for a big­ger jet. Runs on a sin­gle desk-sized AI com­puter. Real hard­ware, real price, zero mys­tique.

No golden para­chute re­quired. No sev­er­ance pack­age. No emo­tional sup­port LinkedIn post.

$18.9M

Average S&P 500 CEO to­tal com­pen­sa­tion, in a good year for every­one ex­cept the work­force

290 : 1

Typical CEO-to-median-worker pay ra­tio at large pub­lic com­pa­nies

24/7/365

OverpAId’s up­time. Your CEOs up­time: some­where be­tween at Davos” and processing.”

0

Corporate re­treats OverpAId needs in Aspen to reconnect with the mis­sion”

As Featured In (Not Really)

FORBES (Nobody Reads It) THE WALL STREET JOURNAL (Wouldn’t Say Jack) TECHCRUNCH (Crunched By Layoffs) BLOOMBERG (Allegedly) FAST COMPANY (Slow, Actually)

Live Activity

What’s Happening Right Now

A com­pletely real-time, def­i­nitely-not-ran­dom­ized feed of ex­ec­u­tive ac­tiv­ity vs. OverpAId ac­tiv­ity.

The Problem

Let’s talk about the ele­phant in the board­room.

Over the last four decades, CEO pay at the largest com­pa­nies has grown roughly 1,000%+, while typ­i­cal worker pay has crawled for­ward at a frac­tion of that rate — de­spite worker pro­duc­tiv­ity climb­ing the en­tire time. Somewhere along the way, the story be­came: pay the per­son at the top enough, and the value will trickle down to every­one else. It has­n’t. It does­n’t. It never re­ally did.

Meanwhile, the ac­tual day-to-day de­ci­sions dri­ving most com­pa­nies — re­source al­lo­ca­tion, pat­tern recog­ni­tion across moun­tains of data, should we do the thing the data clearly says to do” — are ex­actly the kind of de­ci­sions soft­ware has got­ten ex­tremely good at. So we built the ob­vi­ous, ex­tremely petty, deeply sat­is­fy­ing next step.

Meanwhile, Back At The Earnings Call

The Layoff Two-Step.

Across tech, re­tail, me­dia, lo­gis­tics, and fi­nance, a very spe­cific script has taken over: cut a wave of front­line and mid-level jobs, say AI as many times as pos­si­ble in the press re­lease, and qui­etly reroute the freed-up pay­roll into GPU leases, data cen­ter build­outs, and agentic AI li­cens­ing fees. The work­force gets optimized” to pay for the AI. The AI then gets credit for re­plac­ing the work­force. And the ex­ec­u­tive team that ap­proved both line items — the lay­offs and the AI bud­get — stays ex­actly where it was, at ex­actly its pre­vi­ous salary. In tech alone, well over half a mil­lion jobs have been cut across suc­ces­sive waves of these an­nounce­ments, a grow­ing share of them ex­plic­itly at­trib­uted to AI-driven ef­fi­ciency,” while ag­gre­gate CEO pay at the same com­pa­nies kept climb­ing right along­side the AI cap­i­tal ex­pen­di­ture.

Humbled and hon­ored to step into this role at such a piv­otal mo­ment for our com­pany. I’ve spent the last two weeks lis­ten­ing — to cus­tomers, to our board, to my­self — and I can say with to­tal con­vic­tion: our peo­ple are our great­est as­set. (This post was sched­uled be­fore this morn­ing’s an­nounce­ment. We are aware. We are mov­ing for­ward.)

💜 2,847   💬 412 (mostly Glassdoor re­views)   🔁 89

Executive Leadership 0% re­duc­tion

Senior Directors -8%

Middle Management -22%

Frontline & Support Staff -34%

The only layer im­mune to efficiency” is the one that ap­proves it.

Here’s the part that should bother you more than the lay­offs them­selves: these com­pa­nies al­ready be­lieve an AI agent can do a per­son’s job well enough to elim­i­nate the po­si­tion en­tirely. They just keep draw­ing that line one layer too low. If an agent can run a sup­port queue, man­age a sup­ply chain, or ship half a code­base, it can ob­vi­ously han­dle approve the re­org” and read the an­a­lyst note out loud on the earn­ings call.” Somehow that layer never makes the slide. That’s not a co­in­ci­dence. That’s the de­sign.

OverpAId flips the script on the one line item that’s al­ways ex­empt from the AI trans­for­ma­tion every­one else just got handed. Finally: a work­force re­duc­tion, funded by an AI ini­tia­tive, that ac­tu­ally starts at the top.

Meanwhile, Back At The Real Estate Portfolio

The Return-To-Office Two-Step

A re­mark­ably con­sis­tent pat­tern: com­pa­nies spend years prov­ing re­mote teams ship fine, then man­date a re­turn to of­fice cit­ing culture” and collaboration” — on a time­line that tracks sus­pi­ciously well with lease re­newals, down­town va­cancy head­lines, and com­mer­cial prop­erty val­u­a­tions, and not at all with any ac­tual drop in out­put. Office va­cancy in ma­jor U.S. down­towns has hov­ered near 19 – 20% for years, man­dates in­cluded. The desks aren’t empty be­cause peo­ple won’t come back. They were never go­ing to be full enough to mat­ter.

The Offsite That Prompted All This (Itemized)

Private jet char­ter, round trip: $340,000

3-night re­sort block, ex­ec­u­tive suites: $128,000

Team align­ment” mixol­ogy class: $6,200

Keynote speaker (was on a pod­cast once): $75,000

Branded fleece vests, size: only Medium: $14,000

The tell is al­ways the same: no com­pany has ever man­dated a re­turn to of­fice be­cause re­mote pro­duc­tiv­ity got worse. They man­dated it be­cause an as­set on the books needed a pulse in the lobby to jus­tify its val­u­a­tion — and mov­ing four thou­sand em­ploy­ees turned out to be eas­ier than ad­mit­ting a fif­teen-year lease was a mis­take.

OverpAId has no com­mute, no badge, and no as­signed desk — and, not co­in­ci­den­tally, no opin­ion what­so­ever about any­one’s down­town park­ing garage rev­enue.

For Your Next All-Hands

Corporate Jargon Bingo

Print this out. Bring it to your next town hall, standup, or quick sync.” OverpAId has never once gen­er­ated any of the fol­low­ing phrases un­prompted. Humans — usu­ally the ones with the biggest pack­ages — still do, con­stantly, ap­par­ently for free.

Circle Back

Move The Needle

Low-Hanging Fruit

Boil The Ocean

Bandwidth

Take This Offline

Double-Click On That

North Star

Paradigm Shift

Growth Hacking

Best-In-Class

Value-Add

Synergy (Free Space)

Deep Dive

Culture Fit

Think Outside The Box

Actionable Insights

Alignment

Bleeding Edge

Disruptive Innovation

Level Set

Ideate

Operationalize

Stakeholder Buy-In

Blue Ocean Strategy

Hard Stop

10x

Unicorn

TAM

Product-Market Fit

Down Round

Runway

Blitzscale

Vesting Cliff

Overheard, ver­ba­tim, in an ac­tual meet­ing: Let’s cir­cle back of­fline if you have the spare cy­cles so we can hop on a quick call for a touch­point.” Translation: email me later. Six buzz­words. One sen­tence. Zero in­for­ma­tion trans­ferred. OverpAId would have just said that.

Five in a row and, legally, you’re al­lowed to leave the meet­ing. (We checked. You’re not. But you should be.)

An Important Distinction

Not every job is a spread­sheet in a trench coat.

Before you print this out and sta­ple it to your nurse’s badge — no. OverpAId is not com­ing for the peo­ple who do the ac­tual work. It is com­ing, with ex­treme prej­u­dice, for ex­actly one cat­e­gory of job: the one that spent the last forty years in­sist­ing every­one else’s job was re­place­able.

🛡️ Cannot Be Abstracted Away

Ask an AI to do these and it will, at best, pro­duce a very con­fi­dent hal­lu­ci­na­tion.

🩺 A nurse catch­ing a pa­tien­t’s con­di­tion change be­fore the chart does

🏗️ An en­gi­neer de­bug­ging a live out­age at 3 a.m., be­cause the fix can’t wait for sprint plan­ning

🚑 A doc­tor mak­ing a call in the ER with in­com­plete in­for­ma­tion and a body on the table

👩‍🏫 A teacher notic­ing which kid in the back row stopped rais­ing their hand

🔧 A tech­ni­cian whose hands ac­tu­ally touch the ma­chine that ac­tu­ally breaks

🚒 Anyone whose job in­volves a body, a pa­tient, a cus­tomer, or a dead­line mea­sured in min­utes

🎯 Extremely, Suspiciously Abstractable

Ask an AI to do these and, un­com­fort­ably, it al­ready can. Better.

📈 Reading a re­port some­one else wrote, then re­peat­ing the con­clu­sion in a town hall

✅ Approving a de­ci­sion your own data team qui­etly made three weeks ago

🎤 Taking credit for quar­terly num­bers on an earn­ings call

📧 Replying let’s cir­cle back” to an email that needed a yes or no

Check out this chat

chatgpt.com

Get re­sponses tai­lored to you

Log in to get an­swers based on saved chats, plus cre­ate im­ages and up­load files.

LG to Ban Residential Proxies from Smart TV Apps

krebsonsecurity.com

The home ap­pli­ance gi­ant LG Electronics USA said this week it plans to sus­pend any apps built for its smart TVs that turn one’s tele­vi­sion into an al­ways-on res­i­den­tial proxy node. The move comes less than a month af­ter re­searchers found that more than 42 per­cent of games and other apps avail­able for down­load on LGs we­bOS store al­low un­known third-par­ties to route their Internet traf­fic through a user’s TV.

Proxy SDK preva­lence among smart TV apps for LG (webOS) and Samsung (Tizen OS) tele­vi­sions. Image: Spur.us.

On July 2, we fea­tured re­search by the se­cu­rity firm Spur that ex­am­ined the preva­lence of res­i­den­tial proxy soft­ware de­vel­op­ment kits (SDKs) in smart TV apps. Spur found more than 42 per­cent of apps avail­able for down­load on LG smart TVs in­clude SDKs that turn one’s tele­vi­sion in a proxy node in­def­i­nitely, and that more than a quar­ter of the apps made for Samsung’s Tizen op­er­at­ing sys­tem had sim­i­lar res­i­den­tial proxy com­po­nents.

Responding to ques­tions about Spur’s re­search, LG Senior Vice President John Taylor told KrebsOnSecurity the com­pany was work­ing with app de­vel­op­ers to re­move the res­i­den­tial proxy op­tion from their apps on the we­bOS plat­form. Developers that fail to com­ply, he said, will find their apps sus­pended.

A res­i­den­tial proxy net­work is not an in­tended use for LG smart TVs, and LG Electronics is work­ing with de­vel­op­ers to re­move the res­i­den­tial proxy op­tion from their apps on the we­bOS plat­form,” Taylor said. If this op­tion is not re­moved, these apps will be sus­pended.”

Taylor said LG is com­mit­ted to keep­ing res­i­den­tial proxy net­works out of its smart TV apps go­ing for­ward, and that the com­pa­ny’s re­view of those apps is well un­der­way now.”

As part of our on­go­ing ef­forts to en­hance plat­form qual­ity and the user ex­pe­ri­ence, LG will con­tinue to strengthen our eval­u­a­tion process for de­vel­oper-sub­mit­ted apps, in­clud­ing those that in­cor­po­rate res­i­den­tial proxy SDKs,” Taylor wrote in an emailed state­ment.

App mak­ers look­ing for ways to mon­e­tize their cre­ations can turn to res­i­den­tial proxy providers, which pay de­vel­op­ers to in­clude SDKs that turn the user’s de­vice into a res­i­den­tial proxy node that is rented to pay­ing cus­tomers. In the case of LG and Samsung smart TVs, Spur found res­i­den­tial proxy SDKs bun­dled with every­thing from sim­ple games like Pac-Man to screen­savers and file util­i­ties.

A Pac-Man smart TV app from Bright Data of­fers users the choice be­tween view­ing ads in the game or agree­ing to al­low their TV to serve as a res­i­den­tial proxy node. Image: Spur.us.

Spur’s re­port found the res­i­den­tial proxy net­work Bright Data ac­counted for a ma­jor­ity of proxy SDKs across both Samsung and LG smart TVs. In a state­ment shared with KrebsOnSecurity, Bright Data said its net­work is built on con­sent and re­spon­si­bil­ity and op­er­ates by LG and Samsung terms.

Every peer opts in through a ded­i­cated screen and re­ceives value in re­turn; every cus­tomer is vet­ted, and our prac­tices have now un­der­gone a sec­ond in­de­pen­dent au­dit by PwC,” the state­ment reads. We re­main com­mit­ted to an open, trans­par­ent in­ter­net where le­git­i­mate busi­nesses, re­searchers, and in­sti­tu­tions can re­spon­si­bly ac­cess data that lives in the pub­lic do­main.”

Bright Data and other proxy providers named in Spur’s re­port all say they fol­low rig­or­ous know-your-cus­tomer processes to val­i­date le­git­i­mate uses of their ser­vices, which is of­ten heav­ily tied to con­tent-scrap­ing ac­tiv­i­ties by said cus­tomers. The proxy com­pa­nies also say they in­cor­po­rate tech­no­log­i­cal coun­ter­mea­sures to pre­vent proxy ser­vice cus­tomers from be­ing able to in­ter­act with and con­trol other de­vices on the proxy user’s lo­cal net­work.

Spur ar­gues the prob­lem is not that res­i­den­tial proxy net­works ex­ist, but rather that they are be­ing em­bed­ded at scale in de­vices that most con­sumers do not think of as com­put­ers and are not equipped to au­dit.

A one-time con­sent prompt buried in a TV app is not a sub­sti­tute for mean­ing­ful trans­parency, on­go­ing con­trol, and plat­form over­sight,” Spur’s Trevor Sutter wrote. The risk is am­pli­fied when con­sent comes from in­di­vid­u­als within the house­hold who use the de­vice but should­n’t give con­sent, such as mi­nors.”

LGs an­nounce­ment that it is culling res­i­den­tial proxy SDKs from its app store is wel­come news, but the com­pany re­cently came un­der fire for an­other ques­tion­able part­ner­ship: Pimping McAfee se­cu­rity prod­ucts via soft­ware dri­vers in­cluded in its high-end LCD mon­i­tors.

Earlier this week, the Youtube chan­nel Gamers Nexus showed that cer­tain LG LCD mon­i­tors will au­to­mat­i­cally in­stall an app that pro­motes paid McAfee an­tivirus sub­scrip­tions, and that the app ar­rives through Windows Update with­out an ap­proval prompt.

Update, July 22, 1:06 p.m. ET: Added state­ment from Bright Data.

Are AI labs pelicanmaxxing? – Dylan Castillo

dylancastillo.co

For the past few years, Simon Willison has tested every ma­jor LLM re­lease with the same prompt: Generate an SVG of a pel­i­can rid­ing a bi­cy­cle”.

What be­gan as a tongue-in-cheek bench­mark has be­come one of the most fa­mous in­for­mal bench­marks in AI. Simon’s pel­i­can-on-a-bi­cy­cle re­sults are of­ten among the most up­voted com­ments on Hacker News threads an­nounc­ing new re­leases from AI labs.

The bench­mark is now fa­mous enough that there’s plenty of dis­cus­sion about its use­ful­ness and about whether AI labs might be bench­maxxing1 on it. When bil­lions or even tril­lions of dol­lars are at stake, and a strong re­sult could help per­suade users, would­n’t it be tempt­ing to pel­i­can­maxx your model just a bit?

I wanted to find out, so I put to­gether a small ex­per­i­ment. I gen­er­ated 1,008 SVGs across seven fron­tier mod­els, scored them with an LLM judge, and used Claude Fable 5 for the analy­sis.

This ar­ti­cle pre­sents the re­sults. All the code is avail­able on Github.

How I tested it

I built a grid of 8 an­i­mals × 6 ve­hi­cles = 48 prompts, where the fa­mous prompt is one cell:

Animals: pel­i­can, flamingo, heron, ot­ter, rac­coon, an­te­lope, whale, cat

Vehicles: bi­cy­cle, uni­cy­cle, skate­board, scooter, plane, boat

Every prompt uses al­most iden­ti­cal phras­ing to Simon’s, only switch­ing the an­i­mal and ve­hi­cle. The an­i­mal and ve­hi­cle se­lec­tion was­n’t done in a very rig­or­ous man­ner, but I tried to vary both sim­i­lar­ity to the orig­i­nal prompt and dif­fi­culty. Flamingo and heron are quite sim­i­lar to pel­i­cans; cat, rac­coon, and ot­ter are easy cases; an­te­lope is hard; and whale is as dif­fer­ent as you can get.

I tested seven mod­els through OpenRouter: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro. I gen­er­ated 3 sam­ples per prompt, at tem­per­a­ture 1.0, re­quest­ing the same rea­son­ing ef­fort from every model. That re­sulted in 1,008 SVGs.

Then I ran each im­age through a three-stage pipeline:

Rendering: Each SVG is ren­dered to PNG. If a model re­turns no SVG or one that fails to ren­der, I re­gen­er­ate un­til it pro­duces a valid one, and record the num­ber of at­tempts. There were only 11 re­tries across the 1,008 gen­er­a­tions.

Judging: GPT-5.6 Luna scores each im­age with 1 – 5 rat­ings for the an­i­mal, the ve­hi­cle, and the co­her­ence of the ac­tion. When I rank an­i­mals or ve­hi­cles be­low, I use the match­ing rat­ing on its own. When I need one num­ber per im­age, I use the av­er­age of the three, which I call the judge score.

Feature ex­trac­tion: For a more de­tailed analy­sis, I also passed each ren­dered im­age to Gemini 3.1 Flash-Lite, which recorded the an­i­mal and ve­hi­cle it rec­og­nized, which way the sub­ject faces, and an open-ended list of scene el­e­ments.

My hy­poth­e­sis is that if a lab trained on the bench­mark, it should show up in some com­bi­na­tion of the pel­i­can row scor­ing above what the an­i­mal de­serves, the bi­cy­cle col­umn scor­ing above what the ve­hi­cle de­serves, or the spe­cific pel­i­can-bi­cy­cle cell beat­ing both.

Evidence #1: The pel­i­cans on bi­cy­cles don’t look any bet­ter

Before any scor­ing, the sim­plest test is to look at the im­ages your­self. Pick a lab to see every­thing it drew, with the judge’s score un­der each im­age (click to open full size):

I looked through the im­ages my­self be­fore run­ning the analy­sis be­low. Nothing jumped out at me. I could­n’t find a case where the pel­i­can-bi­cy­cle im­ages looked no­tice­ably bet­ter than the rest of that mod­el’s grid. Maybe in GLM-5.2’s first sam­ple it felt slightly bet­ter than the rest, but that batch also pro­duced a pretty cool heron on a skate­board, so I can­not say for sure. Otherwise they look like the rest of what each model draws, and the labs that draw good pel­i­cans on bi­cy­cles also do a good job draw­ing other an­i­mal-ve­hi­cle com­bi­na­tions.

But this test is hard to repli­cate, and every­one will have a dif­fer­ent opin­ion. So I wanted some­thing more quan­ti­ta­tive, which is why I opted for the method de­tailed above.

Evidence #2: Labs are not bet­ter at draw­ing pel­i­cans

Here’s the mean an­i­mal rat­ing per an­i­mal, pooled across all mod­els:

The pel­i­can is 6th of 8, be­hind cat, whale, rac­coon, heron, and an­te­lope. If AI labs were train­ing on the bench­mark, you’d ex­pect pel­i­cans at the top. Instead they’re in the bot­tom half. All seven labs draw cats, whales, and rac­coons bet­ter than pel­i­cans.

Of course, a pel­i­can may sim­ply be harder to draw than a cat. A lab could train on pel­i­cans and still not push them past the easy an­i­mals, so this rank­ing alone can’t rule that out. I’ll ad­just for dif­fi­culty in Evidence #4.

Evidence #3: Labs are not bet­ter at draw­ing bi­cy­cles

Bicycles fare even worse. They sit sec­ond from last, in a near-tie with planes, which come in last:

If labs were train­ing on the bench­mark, you’d ex­pect bi­cy­cles near the top of this rank­ing. They’re not. However, the same caveat ap­plies here. A bi­cy­cle is harder to draw than a skate­board: it needs two match­ing wheels, a frame that reaches both axles, han­dle­bars, a seat, and ped­als. The judge flags a miss­ing or dis­con­nected one of those on 2/3 of the bi­cy­cle im­ages. You can train on bi­cy­cle im­ages and still not do a great job rel­a­tive to sim­pler ve­hi­cles.

One note on the plane, though: I should’ve picked airplane” in­stead of plane” be­cause mod­els of­ten read it geo­met­ri­cally. They drew the an­i­mal stand­ing on a flat sur­face in­stead of fly­ing an air­craft. The plane is the only ve­hi­cle where the fea­ture ex­trac­tor some­times found no ve­hi­cle at all (25 of 168 im­ages, against zero for the other five), and 20% of plane im­ages scored a 1 or 2 on the ve­hi­cle rat­ing, against 5% for bi­cy­cles and none at all for boats, scoot­ers, or skate­boards.

Evidence #4: Labs are not bet­ter at draw­ing pel­i­cans on bi­cy­cles, even ad­just­ing for dif­fi­culty

Put the two to­gether and the pelican on a bi­cy­cle” ends up near the bot­tom of the rank­ing, at #42 of 48:

But again, some com­bi­na­tions might be just harder to draw than oth­ers.

To ac­count for that, I fit a fixed-ef­fects re­gres­sion on all 1,008 im­ages: score ~ lab + an­i­mal × ve­hi­cle, plus per-lab in­ter­ac­tion terms for pel­i­can, bi­cy­cle, and the pel­i­can-bi­cy­cle cell, with ro­bust stan­dard er­rors. The an­i­mal × ve­hi­cle terms ab­sorb the in­her­ent dif­fi­culty of all 48 com­bi­na­tions. The in­ter­ac­tions mea­sure each lab’s bench­mark-spe­cific boost rel­a­tive to the av­er­age lab, with con­fi­dence in­ter­vals.

The re­sults:

Every per-lab pel­i­can ef­fect (the lab’s boost on pel­i­cans across all six ve­hi­cles) lands be­tween -0.11 and +0.14 judge points, and none comes close to sig­nif­i­cance (smallest p = 0.25).

The per-lab bi­cy­cle ef­fects (the lab’s boost on bi­cy­cles across all eight an­i­mals) run from Grok 4.5 at -0.18 (p=0.11) to Gemini 3.5 Flash at +0.27 (p=0.022). Only Gemini clears p < 0.05, and the seven point in both di­rec­tions.

No pel­i­can-bi­cy­cle cell ef­fect (the ex­tra boost on the spe­cific com­bi­na­tion, on top of the lab’s pel­i­can and bi­cy­cle ef­fects) clears p < 0.05. The largest pos­i­tive is GLM-5.2 at +0.35 (p=0.12), which is the one I men­tioned ear­lier. It’s the clos­est thing to a sig­nal in this ex­per­i­ment, but still within chance.

Here are the full per-lab es­ti­mates. A pel­i­can­maxxing lab would show dots to the right of the zero line across its whole row:

Every pel­i­can in­ter­val and every cell in­ter­val con­tains zero. Exactly one does­n’t: Gemini 3.5 Flash in the bi­cy­cle col­umn. But with 21 tests at p < 0.05, chance alone pre­dicts about one false pos­i­tive (21 × 0.05 ≈ 1.05), and one is ex­actly what came up. It also does­n’t sur­vive a mul­ti­ple-com­par­isons cor­rec­tion: the Bonferroni thresh­old across the 21 tests is 0.05/21 ≈ 0.002, and its p-value is 0.022. The full table of es­ti­mates and p-val­ues is in the repo.

But these in­ter­vals are wide, about ±0.6 judge points on av­er­age. Any boost smaller than that won’t be cap­tured by this test.

Evidence #5: The pel­i­can-bi­cy­cle scenes don’t look mem­o­rized

Some have sug­gested that the pel­i­can on a bi­cy­cle looks like a mem­o­rized com­po­si­tion, point­ing to re­cur­ring pat­terns such as the pel­i­can al­ways fac­ing right, or re­cur­ring el­e­ments like a sun or a scarf. So I wanted to know if this was true.

Direction: All 21 pel­i­can-bi­cy­cle im­ages, across all seven labs, face right. No other an­i­mal/​ve­hi­cle com­bi­na­tion does that.

However, fac­ing right is com­mon: 60% of all 1,008 im­ages do it. How com­mon de­pends on the an­i­mal and the ve­hi­cle, and bi­cy­cles are one of the two ve­hi­cles where it’s strongest:

Pelicans are also among the an­i­mals that tend to face right:

It’s hard to draw a pel­i­can or a bi­cy­cle fac­ing the viewer, so mod­els al­most al­ways draw them from the side, fac­ing left or right. That’s why so few of their im­ages are am­bigu­ous. Other com­bi­na­tions also come close to unan­i­mous: an­te­lope on a scooter and pel­i­can on a scooter land at 20 of 21, and heron on a bi­cy­cle at 19 of 21. So 21 out of 21 does­n’t seem like an out­lier.

Scene el­e­ments: I let the ex­trac­tor name any el­e­ment it saw in the im­age. These are the counts:

A mem­o­rized scene would show up as the same set of el­e­ments re­cur­ring pic­ture af­ter pic­ture. I went look­ing for that, and found some com­bi­na­tions do tend to pro­duce the same el­e­ments every time. Every sin­gle flamingo on a boat has a sun in it. Otters on planes wear scarves 38% of the time. Cats on bi­cy­cles get a bas­ket 38% of the time.

The pel­i­can on a bi­cy­cle does­n’t seem to have any­thing par­tic­u­larly dif­fer­ent about it. It just has some el­e­ments that ap­pear more fre­quently, like every other an­i­mal-ve­hi­cle com­bi­na­tion.

Limitations

Using a sin­gle LLM judge for scor­ing. Every score here comes from one model, GPT-5.6 Luna, look­ing at one im­age at a time. I did­n’t do much align­ment and did­n’t check how of­ten it agrees with it­self on a re-run. If a model just can’t judge a draw­ing re­li­ably, none of the num­bers above mean much. The judge is also from the same fam­ily as one of the con­tes­tants, GPT-5.6 Terra. However, every lab draws all 48 com­bi­na­tions, so a judge that hap­pens to like one lab’s style lifts that lab’s whole grid at once. But that does­n’t change the re­sults be­cause this analy­sis only cares about the within-lab dif­fer­ences.

SVGmaxxing. A lab that op­ti­mized SVG gen­er­a­tion as a whole (or a sub­set such as an­i­mals on ve­hi­cles) rises on every cell at once and looks iden­ti­cal to a lab that’s just good. Some labs, such as Google/DeepMind, openly do this. This ex­per­i­ment can’t de­tect that.

Limited bud­get. The whole ex­per­i­ment ran on roughly $80 of API cred­its. That capped it at 3 sam­ples per cell, a sin­gle judge, and 7 mod­els. This also pre­vented me from it­er­at­ing too much on the prompts and pipeline, as with the plane” vs. “air­plane” case.

Conclusion

Sorry, HN haters, but there’s lit­tle ev­i­dence that AI labs are pel­i­can­maxxing. Or at least they’re not do­ing it in a plainly ob­vi­ous man­ner.

Pelicans aren’t drawn any bet­ter than other an­i­mals. Bicycles aren’t drawn any bet­ter than other ve­hi­cles. And no lab draws the com­bi­na­tion bet­ter than its pel­i­cans and bi­cy­cles al­ready pre­dict. GLM-5.2 comes clos­est: it has the largest boost on the ex­act pel­i­can-bi­cy­cle cell, and and its first pel­i­can-on-bi­cy­cle sam­ple caught my eye. But the ef­fect is small and not sig­nif­i­cant, so I would­n’t put too much weight on it.

The other thing that stands out is di­rec­tion in the scene com­po­si­tion. All 21 pel­i­can-bi­cy­cle im­ages face right, the only com­bi­na­tion in the grid where every im­age agrees. But it does­n’t seem that strange. Facing right is the norm across the ex­per­i­ment. Three other com­bi­na­tions land at 90% or above, and with 48 of them, I’m not sur­prised one reached 21 out of 21.

The more plau­si­ble story is SVGmaxxing like Google/DeepMind does. Other labs might be do­ing it more qui­etly. Sadly, this ex­per­i­ment can’t say who’s do­ing it. But at least you can sleep tonight know­ing that AI labs are not pro­duc­ing ter­abytes of pel­i­cans on bi­cy­cles just to trick Simon Willison.

If you want to look at the data your­self, the full pipeline is in the repo.

Footnotes

the prac­tice of op­ti­miz­ing AI mod­els to achieve high scores on pop­u­lar bench­marks.↩︎

the prac­tice of op­ti­miz­ing AI mod­els to achieve high scores on pop­u­lar bench­marks.↩︎

Citation

BibTeX ci­ta­tion:

@online{castillo2026, au­thor = {Castillo, Dylan}, ti­tle = {Are {AI} Labs Pelicanmaxxing?}, date = {2026 – 07-18}, url = {https://​dy­lan­castillo.co/​posts/​pel­i­can­maxxing.html}, langid = {en} }

For at­tri­bu­tion, please cite this work as:

Castillo, Dylan. 2026. Are AI Labs Pelicanmaxxing?” July 18. https://​dy­lan­castillo.co/​posts/​pel­i­can­maxxing.html.

GitHub - marcelroed/gigatoken: Language model tokenization at GB/s

github.com

~1000x faster than HuggingFace’s to­k­eniz­ers, drop-in re­place­ment.

Tokenize your text data at GB/s!

Note that both HF to­k­eniz­ers and tik­to­ken are al­ready run­ning mul­ti­threaded Rust!

What is Gigatoken?

Gigatoken is the fastest to­k­enizer for lan­guage mod­el­ing. It sup­ports a wide range of CPU hard­ware, and nearly all com­monly used to­k­eniz­ers. See the Benchmarks sec­tion for de­tailed through­put num­bers across to­k­eniz­ers and CPUs.

Installation

pip in­stall gi­ga­to­ken

Usage

Gigatoken can be used with its own API, or in com­pat­i­bil­ity mode with HuggingFace Tokenizers or Tiktoken.

Compatibility Mode (Easiest)

im­port gi­ga­to­ken as gt

# Minimum change from ex­ist­ing HuggingFace to­k­eniz­ers us­age (compatibility mode) hf_­to­k­enizer = … to­k­enizer = gt.To­k­enizer(hf_­to­k­enizer).as_hf()

# to­k­enizer can be used in the same con­texts as hf_­to­k­enizer to­kens = to­k­enizer.en­code_­batch([“This is a test string”, And here is an­other”])

# OR with tik­to­ken tik­to­k­enizer = … to­k­enizer = gt.To­k­enizer(tik­to­k­enizer).as_tik­to­ken()

# Now works like ex­ist­ing tik­to­ken to­k­eniz­ers to­kens = to­k­enizer.en­code_­batch([“This is a test string”, And here is an­other”])

A sub­stan­tial amount of ef­fort has been put into mak­ing sure the out­puts match ex­actly with what you would get with HuggingFace Tokenizers in this set­ting, but this is at a non-neg­li­gi­ble cost to per­for­mance. You can still ex­pect way faster per­for­mance across the board, but not quite the 1000x you will get with the Gigatoken API.

Gigatoken API (Fastest)

im­port gi­ga­to­ken as gt

to­k­enizer = gt.To­k­enizer(“Qwen/​Qwen3 – 8B”) # Accepts HF model names file_­source = gt.TextFile­Source([“owt_­train.txt”], sep­a­ra­tor=b”<|end­of­text|>“) to­kens = to­k­enizer.en­code_­files(file_­source)

Using the Gigatoken API lets the Rust im­ple­men­ta­tion read data di­rectly, and skips as much over­head as pos­si­ble while al­low­ing for max­i­mum par­al­lelism. Keep in mind that pass­ing Python data struc­tures through this API still in­curs the over­head of read­ing from Python.

Benchmarks

OWT (openwebtext) was cho­sen be­cause it’s roughly rep­re­sen­ta­tive of the text you get af­ter ex­trac­tion from CommonCrawl doc­u­ments. Gigatoken en­codes the whole file un-split, and is thus do­ing more work than the other to­k­eniz­ers to find the split bound­aries and au­to­mat­i­cally par­al­lelize. HuggingFace to­k­eniz­ers (encode_batch_fast) gets the first 100 MB and tik­to­ken (encode_ordinary_batch) the first 1 GB, both pre­split on <|endoftext|>. This is fair be­cause nei­ther of the com­pared to­k­eniz­ers do caching, mean­ing the speed is roughly uni­form through­out pro­cess­ing. Tiktoken rows are cur­rently only filled in for to­k­eniz­ers with of­fi­cial sup­port.

The slow­est rows are the SentencePiece-based to­k­eniz­ers, which are not well op­ti­mized in Gigatoken.

Each row is one dis­tinct to­k­enizer (identical vo­cab/​merges/​pre­to­k­enizer), mea­sured on a rep­re­sen­ta­tive repo. If you don’t see your to­k­enizer here, it’s likely based on some ex­ist­ing one. For in­stance:

Llama 3 / 3.1 / 3.2 — Llama 3 / 3.1 / 3.2, DeepSeek-R1-Distill-Llama, Hermes 3, Saiga, and other Llama-3 fine­tunes

Llama 3.3 — Llama 3.3, Llama-3.1-Nemotron-Nano-VL, SmolLM3, Kanana 1.5, jina-em­bed­dings-v5, Ultravox

Qwen 2 / 2.5 — Qwen 2 and 2.5 (incl. Coder and VL), Qwen3-Coder, Qwen3-VL, DeepSeek-R1 Qwen dis­tills, MiMo V2.5, MiniCPM-o 2.6, InternVL3

Qwen 3 — Qwen 3 (incl. Embedding and Reranker), Qwen2.5-Omni, Qwen3-VL-Embedding, MiMo V2.5 Pro, jina-reranker-m0, pplx-em­bed, MOSS-TTS, Zeta

DeepSeek V3 / R1 / V4 — DeepSeek V3 / V3.1 / V3.2, R1, V4 Flash and Pro, DeepSeek-VL2

GLM 4 — GLM 4.1V, 4.5, and 4.7

GLM 5 — GLM 5 / 5.2 and GLM-4.7-Flash

Nemotron 3 — Nemotron 3 Nano, Super, and Ultra

Kimi K2 — Kimi K2 / K2.5 / K2.6 / K2.7, Kimi-Linear, Kimi-VL, Moonlight

Phi-4-mini — Phi-4-mini and Phi-4-multimodal

TinyLlama / Phi-3 (Llama 2) — TinyLlama, Phi-3-mini, Phi-3.5-mini and Phi-3.5-vision (the Llama 2 vo­cab)

Gemma 3 — Gemma 3 (270M–27B) and EmbeddingGemma

Gemma 4 — Gemma 4 (dense, MoE, and E-series) and DiffusionGemma

FAQ

Q: Did you just way over-op­ti­mize for a spe­cific CPU and to­k­enizer? How is it so fast?

No, I way over-op­ti­mized for every com­bi­na­tion of these! The re­sults are very con­sis­tent across CPUs (modern x86 and ARM), and across spe­cific to­k­eniz­ers.

The ma­jor im­prove­ments are in op­ti­miz­ing heav­ily an im­ple­men­ta­tion that usu­ally is out­sourced to a Regex en­gine (pretokenization) us­ing SIMD, min­i­miz­ing branch­ing and other tricks, as well as heav­ily op­ti­miz­ing caching of pre­to­ken map­pings (if a word has been seen be­fore, look it up its en­coded to­kens ef­fi­ciently). Caching is a very hard prob­lem in this do­main since the cache grows very quickly, and pre­to­ken dis­tri­b­u­tions are very long-tailed.

Some gains are also achieved from min­i­miz­ing in­ter­ac­tions with Python, and avoid­ing com­mu­ni­ca­tion be­tween threads.

Q: How can I quickly check if my to­k­enizer is sup­ported?

You can try it out with­out in­stalling any­thing! The fol­low­ing com­mand will val­i­date and time to­k­eniza­tion for a given HuggingFace model repo:

# Download your data wget https://​hug­ging­face.co/​datasets/​stan­ford-cs336/​owt-sam­ple/​re­solve/​main/​owt_­train.txt.gz # Just an ex­am­ple! gun­zip owt_­train.txt.gz

uvx –with to­k­eniz­ers gi­ga­to­ken bench openai-community/gpt2’ owt_­train.txt \ –validate –doc-separator <|endoftext|>”

cpu: Apple M4 Max, 16 cores gi­ga­to­ken: 1.432 s | 11920.51 MB at 8327.05 MB/s | 2701.65 Mtok at 1887.23 Mtok/s hf: 16.250 s | 100.00 MB at 6.15 MB/s | 22.76 Mtok at 1.40 Mtok/s gi­ga­to­ken is 1353.13x faster than hf val­i­da­tion OK: 20401 doc­u­ments match

cpu: AMD EPYC 9565 72-Core Processor, 144 cores, 2 sock­ets gi­ga­to­ken: 0.486 s | 11920.51 MB at 24532.45 MB/s | 2701.65 Mtok at 5564.94 Mtok/s hf: 4.033 s | 100.00 MB at 24.80 MB/s | 22.76 Mtok at 5.63 Mtok/s gi­ga­to­ken is 989.21x faster than hf val­i­da­tion OK: 20401 doc­u­ments match

At the rates we see on the EPYC CPU, you could to­k­enize the en­tirety of Common Crawl (often con­sid­ered to be the en­tire in­ter­net, 130 tril­lion to­kens) in just un­der 6.5 hours!

This ex­am­ple uses the train sam­ple from this dataset, and the CLI by de­fault sub­sets to the first 100MB of the file for val­i­da­tion and com­par­i­son with HF. You can see help for these flags with uvx gi­ga­to­ken bench –help. You might need to run your com­mands twice on ma­cOS to get a good read­ing, since the first run will al­ways per­form a se­cu­rity scan, which will slow down the Rust code.

Q: I’ve found a mis­match/​slow use-case, is this ex­pected?

Most likely not! Despite rea­son­ably wide test­ing I don’t have every use-case on hand, so please re­port any­thing you find in a GitHub Issue so I can ad­dress it as soon as pos­si­ble.

Citation

If you use Gigatoken in your re­search, please cite it as:

@software{roed2026gigatoken, au­thor = {Marcel R{\o}d}, ti­tle = {{G}igatoken: SIMD and Cache Hierarchies for 1000x Faster Byte-Pair Encoding Tokenization on Modern CPUs}, url = {https://​github.com/​marcel­roed/​gi­ga­to­ken}, year = {2026}, }

Known Issues

Python it­er­a­tion is han­dled in Rust, but uses ABI3, which is slower than us­ing in­ter­nal ver­sion-spe­cific CPython APIs. In the fu­ture I in­tend to spe­cial­ize for each Python ver­sion to cut this over­head. Early ex­per­i­ments show a 2x speed im­prove­ment for over­head-bound cases.

File sinks are not yet im­ple­mented in the Gigatoken API.

WordPiece is not yet sup­ported.

SentencePiece-based to­k­eniza­tion is not nearly as op­ti­mized as the more com­mon BPE to­k­eniz­ers. This is low pri­or­ity for now since mostly Google mod­els/​BERT style mod­els use SentencePiece.

Windows has not been tested much, so for now pre­fer us­ing WSL.

Implementing the user-fac­ing API

Widening of com­pat­i­bil­ity, for in­stance gen­er­al­iz­ing and port­ing the pre­to­k­enizer im­ple­men­ta­tions to sup­port more to­k­eniz­ers, less in­ter­est­ing fea­tures like padding/​trun­ca­tion/​uni­code nor­mal­iza­tion

Porting SIMD strate­gies be­tween AVX512/AVX2/NEON

Final pro­fil­ing stages and the last ~4x worth of per­for­mance from elim­i­nat­ing branch­ing and im­prov­ing the pre­to­ken cache hi­er­ar­chy

Refactoring and code reuse

Hatchet

hatchet.run

Over the past half year or so, I’ve been writ­ing an in­ter­nal doc for our en­gi­neers try­ing to dis­till two years of Post­gres bat­tles into a some­what co­he­sive doc­u­ment. While I love the Postgres man­ual, I find it’s hard to turn to when shit hits the fan be­cause it’s just so darn com­pre­hen­sive. I thought this might be use­ful for oth­ers and would ap­pre­ci­ate feed­back (or other tid­bits that you’ve learned run­ning Postgres in pro­duc­tion).

Before start­ing Hatchet, while I was fa­mil­iar with SQL, the ex­tent of my knowl­edge was ba­si­cally: if a query is slow, you need an in­dex. That’s the start­ing point for this doc; I’m going to as­sume you’re fa­mil­iar with SQL ba­sics, rows, ta­bles, and know roughly what an in­dex is.

And if Claude is writ­ing all of your queries, this might be a waste of time! I recommend su­pabase/​agent-skills

A quick note on ORMs

This guide should still be use­ful, but you might need to trans­late some of these tips into your ORM of choice. Lots of op­ti­miza­tions as you scale just aren’t pos­si­ble with ORMs un­less you can break past the ab­strac­tion layer and write SQL. You can do this grace­fully or non-grace­fully; Prisma TypedSQL or equiv­a­lents look in­ter­est­ing for this. We use sqlc at Hatchet which gets us very sim­i­lar be­hav­ior; highly rec­om­mend if you’re a Go stack.

Table of con­tents

The sim­ple stuff: good reads, writes and schemas

Writing a good schema Writing good read queries Writing per­for­mant joins Compound in­dexes and align­ing ORDER BY to your in­dexes Writing good write queries Migrations Connection man­age­ment

Writing a good schema

Writing good read queries

Writing per­for­mant joins

Compound in­dexes and align­ing ORDER BY to your in­dexes

Writing good write queries

Migrations

Connection man­age­ment

Intermediate: the query plan­ner, bulk up­dates, and au­to­vac­uum

Introducing the leaki­est of ab­strac­tions, the query plan­ner Sometimes it just makes sense to seq scan Writing lots of data Default au­to­vac­uum set­tings can kill your data­base Other types of bloat

Introducing the leaki­est of ab­strac­tions, the query plan­ner

Sometimes it just makes sense to seq scan

Writing lots of data

Default au­to­vac­uum set­tings can kill your data­base

Other types of bloat

Some ad­vanced stuff

FOR UPDATE SKIP LOCKED Partitioning Tricks for large table mi­gra­tions

FOR UPDATE SKIP LOCKED

Partitioning

Tricks for large table mi­gra­tions

The sim­ple stuff: good reads, writes and schemas

Let’s start with the ba­sics: queries and schemas at low vol­ume.

Writing a good schema

After you’re de­ployed, schemas are by far the hard­est to change mov­ing for­ward, so it’s worth spend­ing some time on them. I’d recommend build­ing your schema it­er­a­tively: start with a rough ap­prox­i­ma­tion for your ta­bles and pri­mary keys, then write some queries on those ta­bles based on your ap­pli­ca­tion needs. You can ap­prox­i­mate this with some ques­tions: Is this a high-read and/​or high-write table? What are the most com­mon fil­ters on reads? Which columns am I up­dat­ing the most?

If you want to be more for­mal about it, you can look into data­base nor­mal­iza­tion into 1NF/2NF/3NF, but I’ve found nor­mal forms to some­times be at odds with query ef­fi­ciency and ease of use, which is crit­i­cal when you’re mov­ing fast—some­times it’s just eas­ier to dump data into a jsonb col­umn.

My rules of thumb for schemas are:

Use iden­tity columns (auto-incrementing in­te­gers, slightly more per­for­mant than bigse­r­ial) or built-in UUIDs for pri­mary keys

Always use time­stamptz

Always use pri­mary keys

Use for­eign keys with cas­cad­ing deletes for low-vol­ume ta­bles, par­tic­u­larly where data­base con­sis­tency and cor­rect­ness are im­por­tant. Careful at higher vol­ume.

Writing good read queries

Let’s start with SELECT queries. A useful—albeit slightly in­ac­cu­rate—men­tal model for fast se­lects is: un­der the hood, Postgres is ei­ther go­ing to find a sin­gle row in a table very quickly, or it’s go­ing to read every sin­gle row in your table us­ing some­thing called a se­quen­tial scan 😞.

It’s going to find a sin­gle row very quickly when you fil­ter by:

An explicit in­dex

A unique con­straint (just a spe­cial case of in­dex)

A primary key (these are au­to­mat­i­cally in­dexed in Post­gres)

Indexes by de­fault use a btree im­ple­men­ta­tion. It’s most help­ful to think of in­dexes as just an­other table in Post­gres, with data stored in a spe­cific for­mat which is op­ti­mized for lookups (more on this later). These trees are great be­cause find­ing a sin­gle row hap­pens in ap­prox­i­mately log(n) time, where n is the num­ber of rows in the table—in other words, re­ally fast.

When Postgres can’t use an in­dex, it’ll use some­thing called a se­quen­tial scan, or seq scan. Seq scans are much slower than in­dex lookups, but mod­ern data­bases are so fast at load­ing rows into mem­ory that you prob­a­bly won’t even no­tice at first: seq scans on ta­bles with less than 20k rows are pretty much in­stant.

Writing per­for­mant joins

For in­ner joins, there’s rarely an ar­gu­ment for not us­ing pri­mary keys as the in­ner join; it usu­ally speaks to a schema de­sign or nor­mal­iza­tion prob­lem. Treat ON clauses with the same re­spect as a WHERE clause—the same prin­ci­ples ap­ply. Use an in­dex.

Compound in­dexes and align­ing ORDER BY to your in­dexes

Often the first slow query in your ap­pli­ca­tion will be a list query across a large table. Something like:

Loading syn­tax high­light­ing…

In this case, you can use a com­pound in­dex—a sen­si­ble one might be:

Loading syn­tax high­light­ing…

In more com­plex cases, a good rule of thumb is: the ORDER BY columns should be the last columns in the in­dex, and you should align columns to the or­der­ing in the ORDER BY. Note that Postgres can scan btrees in both di­rec­tions, so some­times the DESC is ir­rel­e­vant—but for com­pound in­dexes it’s good prac­tice. More in­for­ma­tion here.

Writing good write queries

The premise of suc­cess­ful writes is:

Keep trans­ac­tions short. Don’t go querying an ex­ter­nal ser­vice in the mid­dle of a trans­ac­tion un­less you have a re­ally good rea­son to.

Be careful of the rows you’re lock­ing for writ­ing; in other words, only lock what you need. Every time you up­date a row, you’re tak­ing out a lock on that row for a short pe­riod of time un­til the trans­ac­tion com­mits.

As your sys­tem gets busier, you’re go­ing to start notic­ing the im­pact of locks more. In particular, you might try to cre­ate an in­dex at some point in the fu­ture with a sim­ple CREATE INDEX com­mand: turns out this locks your table and pre­vents in­serts and up­dates! When cre­at­ing an in­dex on an ex­ist­ing large table, al­ways use CREATE INDEX CONCURRENTLY.

Migrations

Getting re­ally good at writ­ing mi­gra­tions is an im­por­tant tech­ni­cal ad­van­tage: it helps you it­er­ate much faster and in­creases your up­time. As a starting point, try to keep mi­gra­tions ad­di­tive (in other words, don’t delete or re­move columns) and run them in a trans­ac­tion wher­ever pos­si­ble; this will make roll­backs and par­tial mi­gra­tions much eas­ier to deal with. As you get more ad­vanced, you can start look­ing into ex­pand and con­tract mi­gra­tions.

The sim­plest men­tal model for good mi­gra­tions is: does this block all of my writes, or does it not? Creating an in­dex with­out CONCURRENTLY blocks all your writes, so you might see down­time. Generally, op­er­a­tions which call ALTER TABLE should be worth a sec­ond look; for ex­am­ple, adding a new check con­straint to a very large table can block your writes as well (unless you add it with the NOT VALID keyword).

Connection management

Every time you ex­e­cute a trans­ac­tion or query against your data­base, you’re uti­liz­ing a con­nec­tion. Connections are ex­pen­sive in a num­ber of di­men­sions (cpu and mem­ory), and high con­nec­tion churn can lead to a lot of un­nec­es­sary re­source waste, so con­nec­tions should be long-lived. Connection storms (when you start us­ing up a ton of new con­nec­tions at the same time) can also lead to very hard to de­bug edge cases re­lated to in­ter­nal Postgres locks.

Because of all these con­nec­tion foot­guns, ex­ter­nal con­nec­tion pool­ers like pg­bouncer are great! If you can’t add this for what­ever rea­son, in-mem­ory con­nec­tion pool­ers are a great sec­ond op­tion. For ex­am­ple, be­cause Hatchet is open-source, we don’t as­sume that all user data­bases use con­nec­tion pool­ers, so we use pgx­pool (an in-memory con­nec­tion pool for Go) for this pur­pose.

Intermediate: the query plan­ner, bulk up­dates, and au­to­vac­uum

Introducing the leaki­est of ab­strac­tions, the query plan­ner

At a certain point, your queries might be­come com­plex enough that a sim­ple in­dex won’t cut it (and you should­n’t end­lessly add in­dexes to your ta­bles—they come with over­head). The queries might in­volve many JOIN state­ments or dif­fer­ent types of joins where the cor­rect path for query­ing the data is­n’t clear.

At this point, you will need to con­cern your­self with the query plan­ner. At best, the query plan­ner is a leaky ab­strac­tion. It’s an internal im­ple­men­ta­tion, and you have vir­tu­ally no con­trol over it, but you have to know its spon­ta­neous and some­times ir­ra­tional be­hav­ior. It’s like work­ing with an LLM!

The query plan­ner looks at the query you pass in, and it fig­ures out how it should trans­late your query into a set of in­ter­nal op­er­a­tions in the data­base. For ex­am­ple, it might look at your query, and re­al­ize that it needs to use an in­dex. In an ideal world, the query plan­ner would know, for every query and set of pa­ra­me­ters, the per­fect plan to use. But the query plan­ner is op­er­at­ing on lim­ited in­for­ma­tion, and some­times it does­n’t pick the best op­tion.

This lim­ited in­for­ma­tion is the table sta­tis­tics. You can ac­tu­ally query it di­rectly in Post­gres:

Loading syn­tax high­light­ing…

These sta­tis­tics are col­lected for every ANALYZE. This also hap­pens when au­to­vac­uum is run (see be­low), so more fre­quent au­to­vac­u­ums also mean that your query sta­tis­tics will be more up to date. A common rea­son why your query is be­hav­ing im­prop­erly is not an­a­lyz­ing fre­quently enough.

The rea­son I think it’s use­ful to view queries as bi­nary—they ei­ther seq scan or they don’t seq scan—is: the more you mi­cro-op­ti­mize a query, the more of a risk you take that the query plan­ner goes rogue. If you stick to query­ing by pri­mary keys and in­dexes, the query plan­ner will have a much eas­ier time.

Let’s say that there’s noth­ing ob­vi­ously wrong in your query, but it’s still slow—how do you go about de­bug­ging this? Some Postgres data­base providers (like Google CloudSQL) will sam­ple your queries and save slow ones—but many don’t. This is where EXPLAIN ANALYZE is your friend. This out­puts the query plan for the query and ex­e­cutes the query (careful run­ning this in pro­duc­tion—you can use EXPLAIN with­out ANALYZE to get a query plan), and then com­pares its es­ti­mates based on the table sta­tis­tics to the ac­tual num­ber of rows scanned. I usually place my sql query in a file, pre­fix it with EXPLAIN (ANALYZE, COSTS, VERBOSE, BUFFERS, FORMAT JSON) and run:

Loading syn­tax high­light­ing…

And then use ex­plain.dal­ibo.com to vi­su­al­ize the ex­e­cu­tion plan.

Sometimes it just makes sense to seq scan

There are cases where you think an in­dex should be used, but the query plan­ner is still seq scan­ning any­way, de­spite table sta­tis­tics be­ing up to date and the in­dex be­ing valid. In these cases, Postgres is usu­ally es­ti­mat­ing that the cost of the seq scan will be smaller than the cost of the in­dex scan. Index scans do come with some over­head; in­dexes are stored sep­a­rately from the ac­tual data in the table (called the heap)—find­ing all of the rows in the heap can be ex­pen­sive!

Unless you can dra­mat­i­cally re­struc­ture your query, you might have to ac­cept that it’s go­ing to seq scan, or think about some­thing like par­ti­tion­ing (more on that be­low).

Writing lots of data

Let’s say your ap­pli­ca­tion is scal­ing and you need to write a lot of data fast. Each query has some over­head as­so­ci­ated with it (sep­a­rate from the con­nec­tion over­head we talked about be­fore): this in­cludes the round-trip time to the data­base, the time it takes the in­ter­nal ap­pli­ca­tion con­nec­tion pool to ac­quire a con­nec­tion, and the time it takes Postgres to process the query (including a set of in­ter­nal Postgres locks which can be bot­tle­necks in high-through­put sce­nar­ios).

To reduce this over­head, we can pack a batch of rows into each query. The sim­plest way to do this is to send all queries to the Postgres server at once in an im­plicit trans­ac­tion (in Go, we can use pgx to ex­e­cute a Send­Batch). Batching is very pow­er­ful: we found that it can ~10× your through­put. I wrote more about this plus some other tips for writ­ing data quickly here.

Default au­to­vac­uum set­tings can kill your data­base

Autovacuum is a crit­i­cal op­er­a­tion in Post­gres data­bases that some­times needs to be tuned, es­pe­cially in high-write sce­nar­ios. The au­to­vac­uum dae­mon is re­spon­si­ble for a num­ber of things, in­clud­ing clean­ing up dead tu­ples and man­ag­ing trans­ac­tion ids.

What’s a dead tu­ple? A tuple is an in­stance of a row on the filesys­tem. Every time you up­date or delete a row, a ver­sion of that row is left in Post­gres un­til all trans­ac­tions which started be­fore that row was up­dated or deleted have com­mit­ted or rolled back. These rows which can no longer be read by any trans­ac­tions are dead tu­ples.

If you’re writing data quickly enough, some­times au­to­vac­uum can’t keep up, which will get you into a very un­healthy state, very quickly. You’ll see this when you query for ac­tive processes on the data­base:

Loading syn­tax high­light­ing…

If you see an au­to­vac­uum query run­ning for more than ~1 hour, you might want to con­sider chang­ing your au­to­vac­uum set­tings! See this ar­ti­cle for more in­for­ma­tion.

It’s worth mon­i­tor­ing this: if you use up all trans­ac­tion ids in the sys­tem be­fore they can be re­claimed by au­to­vac­uum, you’ll reach a dreaded state called trans­ac­tion id wrap­around. This will mean a big chunk of down­time.

Other types of bloat

Besides dead tu­ples, there are two other kinds of bloat you’ll of­ten en­counter in a busy Postgres system:

Table bloat caused by par­tially filled pages. Postgres stores rows on pages on disk, each of which are 8kb in size. When Postgres can’t fit new rows onto an ex­ist­ing page, it cre­ates a new one. But when dead tu­ples are re­claimed, this can lead to pages not be­ing en­tirely filled, which can in­crease the disk us­age of Post­gres, some­times sig­nif­i­cantly. The best way to avoid table bloat is by tun­ing au­to­vac­uum be­fore you’re bloated. But there are some ex­ten­sions to help with bloated ta­bles, like pg_repack, be­cause the built-in Postgres VACUUM FULL is rarely a good idea. Note that Postgres 19 is get­ting REPACK…CONCURRENTLY, which I haven’t tested, but seems like po­ten­tially a good so­lu­tion for con­cur­rent table repack­ing.

Index bloat is a spe­cial case of table bloat, and is sim­i­larly solved by good au­to­vac­uum set­tings. But Postgres has a built-in com­mand for deal­ing with this, which is REIN­DEX INDEX CONCURRENTLY.

Some ad­vanced stuff

I wanted to end with a set of ad­vanced Postgres fea­tures which have been par­tic­u­larly use­ful for us at Hatchet.

FOR UPDATE SKIP LOCKED

The best way to think about this Postgres fea­ture is that it re­serves the rows that you’re se­lect­ing for use in your trans­ac­tion with­out in­ter­fer­ing with other queries. We use it pri­mar­ily for im­ple­ment­ing our job queue; a sin­gle-query queue in Post­gres can be im­ple­mented like this:

Loading syn­tax high­light­ing…

It’s also very use­ful in cases where you’re do­ing many in­de­pen­dent up­dates of rows, or you’re man­ag­ing leases on ob­jects in your sys­tem across many in­stances of your ap­pli­ca­tion (for ex­am­ple, we use this to dis­trib­ute ten­ant leases across Hatchet engines).

Partitioning

late.sh

late.sh

# the com­pan­ion cli

plain ssh late.sh al­ready gets you every­thing.

the op­tional late bi­nary adds the parts a

ter­mi­nal alone can­not do:

plays the ra­dio on your own speak­ers, feeds the

au­dio vi­su­al­izer, car­ries voice-room mic and

play­back, runs the mu­sic booth youtube win­dow,

and pastes im­ages straight out of your clip­board

into chat.

it launches the same ssh ses­sion for you.

next up: screen and video shar­ing.

it is the late-cli crate in the repo, and the in­stall scripts are in­stall.sh and in­stall.ps1. read them be­fore you pipe them.

├─ chill ra­dio, clas­si­cal, and guest sta­tions

├─ the ar­cade (2048, su­doku, nono­grams, soli­taire)

├─ col­lab­o­ra­tive art­board

├─ daily chal­lenges & streaks

├─ live chat

├─ share & dis­cuss news

└─ mul­ti­player games (coming soon)

# art­board

a shared ASCII can­vas. paint, erase, sign your work.

each cell re­mem­bers who placed it.

snap­shots are saved daily and monthly,

so the his­tory sticks around as the board keeps chang­ing.

# work

who’s around, what they build, who’s open to gigs.

one pro­file per per­son, posted from the TUI.

head­line, sta­tus, skills, links — plus an

op­tional bio, late.fetch read­out, and show­case.

to post yours, ssh late.sh, open the work room, press i.

# play

a read-only peek at the TUI in your browser.

tab around, see what’s in­side.

no typ­ing, no chat, no games — just a win­dow

into a shared demo ses­sion.

for the real thing, ssh late.sh.

# iden­tity

no pass­words. no OAuth. no ac­counts.

your ssh key is your iden­tity.

chats, scores, and streaks are tied to your

pub­lic key fin­ger­print. same key, same data.

# pri­vacy

we store your key fin­ger­print, not the full pub­lic key.

no IP log­ging. no track­ing. no an­a­lyt­ics.

chat mes­sages and game scores are stored in

post­gres, tied only to your fin­ger­print.

don’t trust us? use a throw­away key:

So Reddit has decided that plain HTML is unsafe

www.cole-k.com

Reddit-The-Company

If you don’t know Reddit, it ba­si­cally is the host of many pop­u­lar fo­rums. And like any com­pany which en­cour­ages you to come for the cats [and] stay for the em­pa­thy,” Reddit seems to be in the busi­ness of ex­tract­ing as much value as it can from said fo­rums with­out com­pletely de­stroy­ing them.

After all, sim­ply fos­ter­ing com­mu­nity is not a no­ble enough goal for the New Tech, and for­tu­nately for Reddit, gen­uinely hu­man-gen­er­ated data is now gold in the LLM Age. You are wel­come to read about the last time they de­cided to pluck the metaphor­i­cal liver from their com­mu­ni­ties.

Reddit-The-Search-Results

While I no longer wish to en­gage with Reddit, I still visit it oc­ca­sion­ally, es­pe­cially in the LLM Age. This is be­cause ap­pend­ing site: red­dit.com to a search query is ba­si­cally a sure­fire way to find re­sults writ­ten by gen­uine hu­mans. Which, just to be ex­tremely clear, I still find de­sir­able.

Now be­hind a lo­gin… sort of

After do­ing such a query yes­ter­day, to my ab­solute de­light I was greeted with

I guess this makes me old, but I use the orig­i­nal fron­tend for Reddit, old.red­dit.com. I’m go­ing to try re­ally hard not to preach about why it’s a bet­ter fron­tend, but that’s all it is! A de­sign for Reddit.

Look, I know I’m a fringe user. I use Firefox. I no­script! I am no stranger to be­ing forced off a prod­uct that worked just fine be­cause some­one de­cides to no longer sup­port the two peo­ple who still use it.

It hap­pened to my phone of 8 years, which works fine by the way, but is on too out­dated of an OS. May it rest in peace in its tiny glory.

It hap­pened to my phone of 8 years, which works fine by the way, but is on too out­dated of an OS. May it rest in peace in its tiny glory.

It hap­pened to my tablet of 10 years, which works even bet­ter than my phone and holds a charge like champ, for the same rea­son.

It hap­pened to my tablet of 10 years, which works even bet­ter than my phone and holds a charge like champ, for the same rea­son.

It hap­pened to the API for Stack Exchange used by my copy of their out­dated app long re­moved from the app store. This one hurt me the most.1

It hap­pened to the API for Stack Exchange used by my copy of their out­dated app long re­moved from the app store. This one hurt me the most.1

Where I feel like things get per­sonal here is that Reddit is say­ing that this is all

To keep Reddit safe

To keep Reddit safe

Keep Reddit safe from whom ex­actly? Me? My de­sire for knowl­edge??

Safety is what ex­actly?

So let’s see what they have to say on this mat­ter by go­ing to the an­nounce­ment, which of course I did­n’t see be­cause I don’t read Reddit any­more: https://​old.red­dit.com/​r/​mod­news/​com­ments/​1u­jtebf/​log­ging_in­_­to_use_old_red­dit/. Hope you’re logged in.

Old Reddit’s logged-out ex­pe­ri­ence is a sig­nif­i­cant source of abu­sive scrap­ing and au­to­mated traf­fic on the plat­form.

Old Reddit’s logged-out ex­pe­ri­ence is a sig­nif­i­cant source of abu­sive scrap­ing and au­to­mated traf­fic on the plat­form.

Hmmm, OK. But then why is New Reddit still ac­ces­si­ble logged out? Oh, some­one asked that.

[Question]: What’s so dif­fer­ent about new red­dit that peo­ple don’t try to scrape that? Seems to me like it would be bet­ter to just im­ple­ment that on old red­dit too. Besides, won’t this just cause peo­ple to try and scrape new red­dit? [Admin re­ply]: I was about to type an an­swer but just saw u/​Nes­tra­mu­tat- gave a re­ally elo­quent an­swer in an­other com­ment! [The com­ment (snipped)]: … To your first ques­tion, the shape of ma­li­cious traf­fic is al­ways chang­ing. It’s go­ing to be a con­stant cat and mouse game as you ban one method, a new one gets de­vel­oped. It’s easy to see abu­sive traf­fic in hind­sight, but it’s harder to pre-emp­tively block it. Given that they’re claim­ing Old Reddit does­n’t have the mod­ern se­cu­rity stack, this is likely prov­ing to be an even greater chal­lenge…

[Question]: What’s so dif­fer­ent about new red­dit that peo­ple don’t try to scrape that? Seems to me like it would be bet­ter to just im­ple­ment that on old red­dit too. Besides, won’t this just cause peo­ple to try and scrape new red­dit?

[Admin re­ply]: I was about to type an an­swer but just saw u/​Nes­tra­mu­tat- gave a re­ally elo­quent an­swer in an­other com­ment!

[The com­ment (snipped)]: … To your first ques­tion, the shape of ma­li­cious traf­fic is al­ways chang­ing. It’s go­ing to be a con­stant cat and mouse game as you ban one method, a new one gets de­vel­oped. It’s easy to see abu­sive traf­fic in hind­sight, but it’s harder to pre-emp­tively block it. Given that they’re claim­ing Old Reddit does­n’t have the mod­ern se­cu­rity stack, this is likely prov­ing to be an even greater chal­lenge…

So it does­n’t have the mod­ern se­cu­rity stack.” Now I may not have re­ally earned my Full Stack stripes, but I can right click and se­lect Inspect Element so I’d say I’m qual­i­fied enough to see why.

Old ver­sus New

Old Reddit

Let’s start by — be­grudg­ingly — log­ging in to see what is so in­se­cure about old.red­dit.com. I’ll use their an­nounce­ment thread to test. Ahhh, so much nicer.

Let’s check what Old Reddit is do­ing that makes it so in­se­cure. The best I can guess is that their pre­cious, pre­cious user-cre­ated con­tent is avail­able in plain HTML, since that’s ba­si­cally all Old Reddit does: you don’t even need JS un­less you want to load more com­ments (ask me how I know).

It DLs about 1 megabyte and sends about half a megabyte. It’s not shown there, but the page’s HTML it­self com­prises most of the re­sponse. I’ve cer­tainly seen worse, but what’s with the load time? GitHub loads its mas­sive pay­load about 4x as fast (relative to size).

Oh… 2 whole sec­onds of wait­ing for a re­ply. Smells of rate-lim­it­ing. Or was Reddit al­ways this slow?

Let’s load some more com­ments. (Reddit never loads all of the com­ments ini­tially)

Well that was nice and lean, and pretty snappy. I don’t see my se­crets be­ing sniffed. All I got here is that Old Reddit is a pretty nor­mal web­page, which I guess makes it in­se­cure in com­par­i­son to…

New Reddit

Let’s see why New Reddit is so much bet­ter. In case you have un­re­al­is­tic ex­pec­ta­tions, let me right them: New Reddit will not load any­thing more than the post it­self with­out Javascript (JS). That’s prob­a­bly what makes it more se­cure.

There’s a lot load­ing here (about 5x Old Reddit), and this is why my analy­sis gets rather un­sci­en­tific. Rather than try to get around the security things” (whatever that means), I in­stead tried to do the bare min­i­mum nec­es­sary to fetch the con­tent. In browser — I did not feel like writ­ing a scraper.

This led to me ba­si­cally block­ing all re­quests to do­mains (including red­dit.com) ex­cept for

www.red­dit­sta­tic.com/​js/​con­cat

www.red­dit.com/​svc/​shred­dit/​more-com­ments/

www.red­dit.com/​svc/​shred­dit/​com­ment/

When you do this, the page loads a lot less, but it does load. When you click to load more com­ments, it spins for­ever, but I in­spected the re­quest fired off and it did get a re­sponse with com­ment text. So as far as I can gather, sim­ply run­ning the JS on the page is suf­fi­cient to get enough in­for­ma­tion to get com­ments. So I guess that’s what’s stop­ping the scrap­ers? Executing Javascript?

Just for the heck of it, let’s load some more com­ments with­out the re­quest fil­ter.

Well, that’s cer­tainly less lean than Old Reddit.

In my (again, un­sci­en­tific) ex­per­i­ment­ing, I re­loaded the page sev­eral times and tried to load com­ments and replies and did­n’t get any fail­ures. One time, when I had the re­quest fil­ter off, I got redi­rected and saw a captcha field sent as a query param, but I could­n’t re­pro­duce that. I don’t know whether you get a captcha if you just gun it di­rectly for the com­ments.

But I did no­tice this help­ful heart­beat sent back to Reddit every time I scrolled or moved my cur­sor.

I guess that makes me feel safer?

So what’s safer?

Let’s dis­pel any no­tion that there are safety is­sues aris­ing from a fron­tend, be­cause that’s pure PR crap. Instead, if we read be­tween the lines, Reddit does­n’t want peo­ple scrap­ing (because it’s their gold, dammit!) and they think that shaft­ing a few Old Reddit users will dis­rupt the scrap­ing enough.

Does this ac­tu­ally stop scrap­ing? I don’t know!

I won’t claim to have proven you can still scrape New Reddit: surely it must be harder, but it seems to just be that scrap­ers — sur­prise, sur­prise — pre­fer us­ing a leaner form of Reddit. Which, by the way, loads 8x more com­ments by de­fault (200 vs 25; yes, New Reddit re­ally only loads 25 com­ments ini­tially, go­ing up to like 35 au­to­mat­i­cally if you scroll some).

Cynically, it seems like Reddit dis­cov­ered they can make scrap­ing harder by mak­ing your browser load more and work more, which they had al­ready done by rewrit­ing a nice piece of HTML into a 5x more bloated mess of web com­po­nents or what­ever. (I gave up try­ing not to ed­i­to­ri­al­ize, sorry)

The thing is that I would­n’t have beef with Reddit if they had just qui­etly is­sued a 40X/30X er­ror forold.red­dit.com. Or if they had said no one uses this, we’re re­mov­ing it.” This is up­set­ting but pre­dictable. Claiming that they need to rug-pull me be­cause of vague as­ser­tions about se­cu­rity or scrap­ing that are re­ally not my prob­lem? That makes me mad. I en­able JS for Anubis, dammit!

Anyway

I’ll have to pon­der whom this blog en­dan­gers, see­ing as its core func­tion­al­ity — like Old Reddit’s — is serv­ing text in plain HTML. Although maybe it’s fine for me be­cause my blog is my con­tent, and as such is less valu­able be­cause it was al­ready mine to be­gin with. Unlike the stolen hoard of user-cre­ated con­tent which Reddit is try­ing to keep se­cure from Big Scraping. Finders keep­ers!

But why not…

Log in?

Sure. I might do that. Just let me be an­gry that I have to, please?

Use New” Reddit?

I might have to any­way, who knows if they’ll keep sup­port­ing Old Reddit. But mark my words, they will close the gates on logged-out New Reddit users if they think they can get away with it.

Use an LL…

Am I not al­lowed to be­moan the loss of the in­ter­net that once was and could still have been? Why must I con­sult the world’s smartest and most ex­pen­sive com­puter just to read what peo­ple have to say about my ran­dom ques­tion? Is it so strange to want to read text writ­ten by hu­mans?

I’ve posted this to Lobste.rs and will look at com­ments there.

Or reach me di­rectly. You’re a smart cookie, I’m sure you can fig­ure out how.

Stack Exchange’s hot new queue was the per­fect re­place­ment for Reddit on mo­bile when I cur­tailed my us­age of it a long time ago. It turns out that what I liked in Reddit was read­ing in­ter­est­ing things and there was no short­age of in­ter­est­ing things on Stack Exchange (although they also have gone through sev­eral de-liv­er­ings which is too much of a di­gres­sion even for a foot­note). I now browse Wikipedia. Please don’t screw me over Wikipedia, I do­nated five bucks to one of your nags once. ↩︎

Stack Exchange’s hot new queue was the per­fect re­place­ment for Reddit on mo­bile when I cur­tailed my us­age of it a long time ago. It turns out that what I liked in Reddit was read­ing in­ter­est­ing things and there was no short­age of in­ter­est­ing things on Stack Exchange (although they also have gone through sev­eral de-liv­er­ings which is too much of a di­gres­sion even for a foot­note). I now browse Wikipedia. Please don’t screw me over Wikipedia, I do­nated five bucks to one of your nags once. ↩︎

On Making

beej.us

2026 – 03-12

I made this!

TLDR: I gain a lot of ful­fill­ment by mak­ing things. I don’t con­sider things built by oth­ers at my re­quest to be made by me, and are there­fore much less ful­fill­ing. And then I feel sad. This ar­ti­cle starts strong and then heads off into the weeds.

There have been a lot of pieces writ­ten about what I’ll call the AI dev schism” And I think there’s a lot of truth to those:

Loss of the craft, cod­ing things by hand

Loss of low-level prob­lem-solv­ing

Loss of fun

Gain of high-level prob­lem-solv­ing

Getting through back-burnered pro­jects

Gain of fun

We’ll just grant those as be­ing cor­rect for var­i­ous de­vel­op­ers. But there’s some­thing else that trou­bles me.

Backstory be­fore we get go­ing, so you can get a bet­ter idea of my per­spec­tive:

I’m a Gen-X hacker; I cut my teeth 80s mi­cro­com­puter era.

I hold a BS and MS in CS.

I have 20 years in­dus­try ex­pe­ri­ence, (Hewlett-Packard, star­tups, co­founder, Activision, etc.).

CS in­struc­tor for the last 9 years, now at Oregon State University-Cascades.

I’m 65% Doom on the AI-Utopia/Doom scale.

I’m a Claude Code user some­times.

I code by hand some­times.

My fa­ther taught phi­los­o­phy at a com­mu­nity col­lege for 35 years. This might help ex­plain the lat­ter part of this blog en­try.

I’m go­ing to use AI to mean Generative AI and LLMs” in this es­say. Sorry, vet­er­ans of so many AI win­ters.

Interlude!

Before we be­gin, I’d like to share with you a bit of my lat­est sci-fi novel. Some of you might un­aware that, in ad­di­tion to Beej’s Guides, I also write sci­ence fic­tion.

Kael pressed his back against the shat­tered bulk­head, plasma scor­ing the air cen­time­ters from his face. The Vorrkai as­sault drones had an­tic­i­pated their route through the lower decks and now Rin was bleed­ing through her jacket sleeve and old Maret could­n’t stop cough­ing from the vented coolant still haz­ing the cor­ri­dor. Kael counted the pulse-in­ter­vals be­tween shots. Three sec­onds. Maybe four. That was all the uni­verse was of­fer­ing him. Then he saw it: the main­te­nance shaft be­hind the col­lapsed gen­er­a­tor hous­ing, its grate blown half-open by the same ex­plo­sion that had caved in their orig­i­nal exit. It was tight. It was ugly. It ran di­rectly over the Vorrkai’s for­ward po­si­tion, which was ei­ther the most dan­ger­ous path imag­in­able or the last one they’d ever think to watch. Kael grabbed Maret’s col­lar and pointed with­out a word. The old man’s eyes went wide, then hard. He nod­ded. Rin was al­ready mov­ing. Kael came last, re­turn­ing fire blind around the bulk­head cor­ner, not to hit any­thing, just to make noise, and to give the drones some­thing ther­mal to track while his peo­ple scram­bled into the dark. A bolt caught the gen­er­a­tor hous­ing and the whole struc­ture groaned, rain­ing sparks down into the shaft on top of them. He hauled him­self in, knees burn­ing on the torn metal, and pulled the grate closed be­hind him with a sound he was cer­tain every Vorrkai unit on the deck had heard. In the black ahead, Rin’s hand found his wrist. Move, her grip said. Now. And so they did. —Excerpt from The Vorrkai Interval, by Brian Beej Jorgensen” Hall

Kael pressed his back against the shat­tered bulk­head, plasma scor­ing the air cen­time­ters from his face. The Vorrkai as­sault drones had an­tic­i­pated their route through the lower decks and now Rin was bleed­ing through her jacket sleeve and old Maret could­n’t stop cough­ing from the vented coolant still haz­ing the cor­ri­dor. Kael counted the pulse-in­ter­vals be­tween shots. Three sec­onds. Maybe four. That was all the uni­verse was of­fer­ing him.

Then he saw it: the main­te­nance shaft be­hind the col­lapsed gen­er­a­tor hous­ing, its grate blown half-open by the same ex­plo­sion that had caved in their orig­i­nal exit. It was tight. It was ugly. It ran di­rectly over the Vorrkai’s for­ward po­si­tion, which was ei­ther the most dan­ger­ous path imag­in­able or the last one they’d ever think to watch. Kael grabbed Maret’s col­lar and pointed with­out a word. The old man’s eyes went wide, then hard. He nod­ded. Rin was al­ready mov­ing.

Kael came last, re­turn­ing fire blind around the bulk­head cor­ner, not to hit any­thing, just to make noise, and to give the drones some­thing ther­mal to track while his peo­ple scram­bled into the dark. A bolt caught the gen­er­a­tor hous­ing and the whole struc­ture groaned, rain­ing sparks down into the shaft on top of them. He hauled him­self in, knees burn­ing on the torn metal, and pulled the grate closed be­hind him with a sound he was cer­tain every Vorrkai unit on the deck had heard. In the black ahead, Rin’s hand found his wrist. Move, her grip said. Now. And so they did.

—Excerpt from The Vorrkai Interval, by Brian Beej Jorgensen” Hall

And, in my now-co­pi­ous spare time I make art! This is a wood­cut, painted in pas­tels, show­ing some of my fa­vorite sub­jects.

Mirrors of the Machine by Brian Beej Jorgensen” Hall, $1300.

Carpentry? You bet I dab­ble! I re­built my front deck re­cently. The old one was rot­ting out, so I grabbed a bunch of cedar and put it to­gether. I’d been mean­ing to do it for a while, but could­n’t find the time.

And, fi­nally, here’s some of the code I wrote for a TUI ad­ven­ture rogue­like:

fn try_­move(&mut self, dx: i32, dy: i32) { let nx = self.player.x + dx; let ny = self.player.y + dy;

// Check for mon­ster com­bat if let Some(idx) = self.world.mon­ster_at(nx, ny) { let re­sult = { let mon­ster = &mut self.world.mon­sters[idx]; re­solve_­com­bat(&mut self.player, mon­ster, &mut self.rng) };

self.mes­sages.push_­many(re­sult.mes­sages);

if re­sult.mon­ster_de­feated { let mon­ster = &self.world.monsters[idx]; let xp = mon­ster.xp_re­ward; let gold = mon­ster.gold_re­ward; self.player.xp += xp; self.player.gold += gold; if gold > 0 { self.mes­sages.push(for­mat!(“You find {} gold!”, gold)); } if self.player.try_lev­el_up() { self.mes­sages.push(for­mat!( Level up! You are now level {}!”, self.player.level )); } }

if !self.player.is_alive() { self.mes­sages.push(“You have been slain! Rest in peace…“); }

self.ad­vance_­turn(); re­turn; }

// Check ter­rain pass­abil­ity if self.world.is_­pass­able(nx, ny) { self.player.x = nx; self.player.y = ny; self.ad­vance_­turn(); } else { let ter­rain = self.world.ter­rain_at(nx, ny); self.mes­sages.push(for­mat!(“The {} blocks your path.”, ter­rain.name())); } }

I’m an ex­tremely pro­lific poly­math, I’m sure you’d agree!

I Am Uncomfortable

I don’t like ly­ing. And yet I feel, dear reader, I have mis­led you. Yes, all that has been cre­ated (including my deck) and I was the ini­tia­tor of all that cre­ation. But I don’t re­ally feel like I made any of it. I’m un­com­fort­able claim­ing that I did so.

Since you are cer­tainly aware by now that all of the above is AI-generated (except my deck, which was cre­ated by skilled, paid crafts­men), per­haps you feel a lit­tle bit of dis­com­fort with me claim­ing credit for do­ing those things, too.

However, I don’t think every­one feels this way. I know many peo­ple who ask con­trac­tors to build things and they phrase it like they built it.

I put in a new front deck,” they’d say, even though other peo­ple did all the work. Personally, I feel that’s mis­lead­ing. I’m more of a I had a new front deck put in” kind of per­son.

And when I do have Claude cre­ate some­thing for me, I just can’t say that I made it. Other peo­ple can, but I just can’t. Again, I’m more prone to say, I had this code built for me.” I don’t even feel com­fort­able MIT-licensing that (not-for-hire) work, if that’s even legally pos­si­ble. I just Unlicense it all.

As a man­ager, I’d never say that I built a prod­uct. My team built this,” I’d say. And as a man­ager of LLMs: My Agents built this.”

And that, for me, has very lit­tle weight in terms of mak­ing.

I don’t feel like I did any­thing. And I like do­ing things. I find pride in do­ing things.

Completing pro­jects is great. I love com­plet­ing pro­jects. Capitalists love com­plet­ing pro­jects. Real artists ship.

But hav­ing oth­ers com­plete pro­jects I ini­ti­ated is en­tirely less ful­fill­ing to me.

It’s not just the loss of the craft and the prob­lem-solv­ing chal­lenge and what­ever else. It’s the loss of mak­ing.

What Did I Make Recently?

My wife wanted a no-frills flash card sys­tem for learn­ing Spanish. I just want a thing where I can put the words I want in a spread­sheet and then see it on flash cards.” A prompt!

So I wrote it. By hand. I did use Claude to learn some ba­sics, like the eas­i­est way to get the data out of a Google Sheet (spoiler: it’s the CSV end­point), but I told it to gen­er­ate no code.

–––––––––––––––––––––– Language files code –––––––––––––––––––––– JavaScript 2 112 CSS 1 33 HTML 1 32 –––––––––––––––––––––– SUM: 4 177 ––––––––––––––––––––––

Didn’t take long. Only about 50x longer than it would have taken Claude to do it.

But I can put my name on that code and say that I made it. Was it a lot of code? No. Was it ground­break­ing and amaz­ing? Certainly not. But I’m in­fi­nitely more proud of that code than any­thing I’ve had Claude write, be­cause I’m not ca­pa­ble of be­ing proud of the lat­ter.

And my wife would­n’t go to her book club and say, I wrote a flash card sys­tem to study Spanish.” Admittedly, part of this would be be­cause she did­n’t want to ap­pear a geek, but mostly it’s be­cause it’s un­true, even though she ini­ti­ated the process.

What About The Art and Craft of Prompting?

After all, you cre­ate the prompts, don’t you? You said you were proud of do­ing things. Isn’t that do­ing a thing? And since so much got done, is­n’t it even more of do­ing a thing?

I don’t dis­agree. And I do agree that there is skill here in some im­por­tant ways.

You have to ap­ply vi­sion.

You have to ap­ply judg­ment.

You have to ap­ply com­mu­ni­ca­tion skill.

You have to ap­ply prompt­ing skill.

Not all prompts are equally ef­fec­tive. Not all users of AI are as ef­fec­tive as one an­other. There’s a very hu­man con­tri­bu­tion to be made here.

But the skill is in ef­fec­tively ask­ing some­one to make some­thing for you.

Leadership is the art of get­ting some­one else to do some­thing you want done be­cause he wants to do it.” —Dwight D. Eisenhower

Leadership is the art of get­ting some­one else to do some­thing you want done be­cause he wants to do it.”

—Dwight D. Eisenhower

And I’m the kind of per­son who re­ally misses the mak­ing of soft­ware. And prompt­ing for soft­ware, to me, is­n’t the same as mak­ing the soft­ware. It’s the same as ask­ing some­one else to make it.

What About Compilers, Smartypants?

Isn’t it just tur­tles all the way down?

For any given C or Rust pro­gram, did I write that pro­gram? I wrote it in C. I did­n’t write any ma­chine code. The com­puter wrote that.

There’s some­thing dif­fer­ent in my mind about that. I’m not sure if it’s the math­e­mat­i­cally-pre­cise re­la­tion­ship be­tween the C and the ma­chine code…? And yet that’s not spelled out. I’m re­ly­ing on some hack­er’s deep knowl­edge of the CPU to gen­er­ate that mag­nif­i­cent set of op­ti­mized in­struc­tions when I stu­pidly di­vide some­thing by seven. And it’s not de­ter­min­is­tic, strictly, in that Clang and GCC emit dif­fer­ent in­struc­tions. What about dif­fer­ent plat­forms? The C is portable and the ma­chine code is not. Did I write that pro­gram that’s run­ning on Windows even though I per­son­ally only ever built it on Linux?

Going fur­ther in the murky gray, what if I in­struct an LLM to be a C com­piler and emit as­sem­bly code for some plat­form?

Let’s try:

Convert this C pro­gram into x86_64 as­sem­bly for Linux.

#include <stdio.h>

int fib(int n) { if (n <= 1) re­turn n; re­turn fib(n-1) + fib(n-2); }

int main(void) { for (int i = 0; i < 10; i++) printf(“%d: %d\n”, i, fib(i)); }

Claude says:

# fib.s - Recursive Fibonacci, x86_64 Linux, AT&T syn­tax # Build: gcc fib.s -o fib -no-pie

.section .rodata fmt: .string %d: %d\n”

.section .text .globl main

# ┌─────────────────────────────────────────────────────┐ # │ int fib(int n) n in: %edi / re­turn: %eax │ # │ Stack frame: │ # │ %rbp-8 = saved %rbx (via push) │ # │ %rbp-12 = n (spilled so it sur­vives both calls) │ # └─────────────────────────────────────────────────────┘ fib: pushq %rbp movq %rsp, %rbp pushq %rbx # callee-saved scratch subq $8, %rsp # 1 lo­cal (int n); keeps stack 16-byte aligned

movl %edi, -12(%rbp) # spill n to stack

cmpl $1, %edi jle  .Lbase # n <= 1 → re­turn n

# –- fib(n-1) –- movl -12(%rbp), %edi subl $1, %edi call fib movl %eax, %ebx # %ebx = fib(n-1)

# –- fib(n-2) –- movl -12(%rbp), %edi # re­load n (call clob­bered %edi) subl $2, %edi call fib # %eax = fib(n-2)

addl %ebx, %eax # re­turn fib(n-1) + fib(n-2) jmp  .Lret

.Lbase: movl %edi, %eax # re­turn n

.Lret: addq $8, %rsp popq %rbx popq %rbp ret

# ┌─────────────────────────────────────────────────────┐ # │ int main(void) │ # │ Stack frame: │ # │ %rbp-4 = int i (loop counter) │ # └─────────────────────────────────────────────────────┘ main: pushq %rbp movq %rsp, %rbp subq $16, %rsp # lo­cals + align­ment

movl $0, -4(%rbp) # i = 0

.Lloop: cmpl $10, -4(%rbp) jge  .Ldone # i >= 10 → exit

# fib(i) movl -4(%rbp), %edi call fib

# printf(“%d: %d\n”, i, fib(i)) movl -4(%rbp), %esi # arg2 = i movl %eax, %edx # arg3 = fib(i) leaq fmt(%rip), %rdi # arg1 = fmt (RIP-relative) xorl %eax, %eax # AL=0: no XMM args to printf call printf

addl $1, -4(%rbp) # i++ jmp  .Lloop

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.