10 interesting stories served every morning and every evening.

Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.

fireworks.ai

K3 is a fron­tier qual­ity open model at a frac­tion of the cost. Even big­ger is that it com­ple­ments Fable pre­dictably, which makes it pos­si­ble to get the high­est qual­ity in­tel­li­gence by rout­ing tasks.

🧭 tl;dr: We ran Kimi K3 (open) against Fable 5 (closed) on ~1,000 agen­tic tasks find­ing:

We achieved 93% ac­cu­racy with rout­ing be­tween K3 and Fable.

Results were up to ~50X more cost ef­fec­tive than Fable alone on long agen­tic loops, and con­sis­tently lower cost across every use case.

How We Measured

We av­er­aged bench­marks, each aimed at a dif­fer­ent kind of work, and ran K3 and Fable 5 through the same har­ness. About 1,030 tasks in all, in real agent loops.

One quick de­f­i­n­i­tion be­fore we get into the re­sults. Oracle rout­ing is a method for mea­sur­ing the best the­o­ret­i­cal per­for­mance by run­ning the task through each model and then pick­ing the cheap­est cor­rect op­tion (the cost/​per­for­mance ceil­ing). In a prac­ti­cal router, you don’t get to run your task against mul­ti­ple mod­els. The router makes a pre­dic­tion of which model has the best cost and qual­ity trade off, but ul­ti­mately it’s a guess.

In this study, or­a­cle rout­ing demon­strated K3 is se­lected for 72 – 96% of tasks. This sug­gests a near-per­fect router might be achiev­able, by learn­ing the dif­fer­ence be­tween day-to-day tasks and the true long tail of fron­tier work. It will re­quire an or­der of mag­ni­tude more rout­ing data, and real world per­for­mance to say de­fin­i­tively.

K3 is a good model.

From a 10,000 foot view, it can be easy to look at both mod­els and call the head-to-head a tie. For ex­am­ple, if you look at SWE, the head­line bench­mark, K3 gets 92.4%, Fable 92.6%. Across the five types of tasks we bench­marked on, the two mod­els tend to stay within a few points of each other, with Fable pulling slightly ahead on its cod­ing-lan­guage breadth (Multi-lang).

It’s easy to stop there and say they’re roughly even”. The news is that they have dis­cretely bet­ter per­for­mance across dif­fer­ent task types.

Two Models is Better than One

If you take a peek in­side a sin­gle bench­mark, there’s more to see than just a top-line ac­cu­racy num­ber. Take SWE, where the two are dead even over­all. If you split SWE by prob­lem do­main you can see where each model shines. K3 is sharpest on sym­bolic math and dev tool­ing; Fable wins on web & data vi­su­al­iza­tion work. The same pat­tern runs through the multi-lan­guage set, where Fable’s breadth car­ries Java, Python and C++, while K3 draws even on JavaScript and Rust.

For long-hori­zon work at a ter­mi­nal, dri­ving a shell and prod­ding at sys­tems across dozens of turns, K3 showed its true col­ors. It cleared a batch of tasks Fable never cracked: a 7z hash, FEAL crypt­analy­sis, leaked se­crets, a live vul­ner­a­bil­ity, run­away async jobs.

K3 can be up to 50x lower cost on Fireworks. 🫳🎤

While qual­ity is a near-tie at a high level, price is­n’t close.

So where’s this huge price gap com­ing from? to­ken pric­ing, prompt caching, and ef­fort-per-task. On SWE for ex­am­ple, K3 works much harder than Fable: roughly 55 turns and 1.3M to­kens a task ver­sus 21 turns and 130K. On the long ter­mi­nal tasks it’s the other way around: Fable is the one that spi­rals, run­ning up 64 turns and 1.5M to­kens (sometimes straight into a time­out).

Prompt caching does most of the work of turn­ing that ef­fort into K3′s price ad­van­tage: even when K3 reads ten times the to­kens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a trade­off. Tasks with ex­tra turns gen­er­ally mean more wall-clock time per run i.e. slower runs. If you need an an­swer in two sec­onds, that mat­ters; if you’re run­ning agents in the back­ground at scale, a bill that’s a frac­tion of the size mat­ters a lot more.

Don’t pick a model. Route.

If you send every task to who­ever han­dles it best, you don’t land some­where be­tween the two mod­els, you land above both.

Per-task rout­ing al­ways out per­forms any sin­gle model run:

The or­a­cle router choose K3, 72 – 96% of task traf­fic. By ar­chi­tect­ing a router this way, you end up with over­all qual­ity above ei­ther model alone at a cost close to just us­ing just the cost-op­ti­mized one.

K3 is cost op­ti­mized on all work types

Put both qual­ity and cost on one plot. K3 in blue lands to the left (the more cost-ef­fec­tive side) of Fable in red in all five task-fam­i­lies. Accuracy trades back and forth: Fable pulls ahead on multi-lan­guage, K3 on ter­mi­nal and le­gal, the rest roughly level.

Single Models Are Wasteful and No Longer SoTA

Kimi K3 + Fable routed to­gether un­locks their best qual­i­ties at the best price.

The sin­gle model provider, to­ken maxxing days, are com­ing to an end. The task-level data says these mod­els are spe­cial­ists at very dif­fer­ent prices. The best AI no longer comes out of a sin­gle lab, it’s a mix­ture of mod­els.

What this means in prac­tice:

Open as the de­fault. A 50x lower cost open model like K3 should be your base case, since the or­a­cle sends it most of the traf­fic any­way.

The router is your moat. A router must be tai­lored to your work­load and learn­ing that task/​model split con­tin­u­ously is the best chance you’ll have at stay­ing ahead.

OverpAId — Fire Your CEO. Hire The Future.

overpaid.lol

Introducing the world’s first Chief Executive Replacement Engine

Your CEO costs $22,000,000 a year.We cost $4,699. Once.

OverpAId is an Artificial Intelligence built from the ground up to do your CEOs en­tire job — strat­egy, vision,” mo­ti­va­tional all-hands emails — bet­ter, faster, and with­out ever once ask­ing the board for a big­ger jet. Runs on a sin­gle desk-sized AI com­puter. Real hard­ware, real price, zero mys­tique.

No golden para­chute re­quired. No sev­er­ance pack­age. No emo­tional sup­port LinkedIn post.

$18.9M

Average S&P 500 CEO to­tal com­pen­sa­tion, in a good year for every­one ex­cept the work­force

290 : 1

Typical CEO-to-median-worker pay ra­tio at large pub­lic com­pa­nies

24/7/365

OverpAId’s up­time. Your CEOs up­time: some­where be­tween at Davos” and processing.”

0

Corporate re­treats OverpAId needs in Aspen to reconnect with the mis­sion”

As Featured In (Not Really)

FORBES (Nobody Reads It) THE WALL STREET JOURNAL (Wouldn’t Say Jack) TECHCRUNCH (Crunched By Layoffs) BLOOMBERG (Allegedly) FAST COMPANY (Slow, Actually)

Live Activity

What’s Happening Right Now

A com­pletely real-time, def­i­nitely-not-ran­dom­ized feed of ex­ec­u­tive ac­tiv­ity vs. OverpAId ac­tiv­ity.

The Problem

Let’s talk about the ele­phant in the board­room.

Over the last four decades, CEO pay at the largest com­pa­nies has grown roughly 1,000%+, while typ­i­cal worker pay has crawled for­ward at a frac­tion of that rate — de­spite worker pro­duc­tiv­ity climb­ing the en­tire time. Somewhere along the way, the story be­came: pay the per­son at the top enough, and the value will trickle down to every­one else. It has­n’t. It does­n’t. It never re­ally did.

Meanwhile, the ac­tual day-to-day de­ci­sions dri­ving most com­pa­nies — re­source al­lo­ca­tion, pat­tern recog­ni­tion across moun­tains of data, should we do the thing the data clearly says to do” — are ex­actly the kind of de­ci­sions soft­ware has got­ten ex­tremely good at. So we built the ob­vi­ous, ex­tremely petty, deeply sat­is­fy­ing next step.

Meanwhile, Back At The Earnings Call

The Layoff Two-Step.

Across tech, re­tail, me­dia, lo­gis­tics, and fi­nance, a very spe­cific script has taken over: cut a wave of front­line and mid-level jobs, say AI as many times as pos­si­ble in the press re­lease, and qui­etly reroute the freed-up pay­roll into GPU leases, data cen­ter build­outs, and agentic AI li­cens­ing fees. The work­force gets optimized” to pay for the AI. The AI then gets credit for re­plac­ing the work­force. And the ex­ec­u­tive team that ap­proved both line items — the lay­offs and the AI bud­get — stays ex­actly where it was, at ex­actly its pre­vi­ous salary. In tech alone, well over half a mil­lion jobs have been cut across suc­ces­sive waves of these an­nounce­ments, a grow­ing share of them ex­plic­itly at­trib­uted to AI-driven ef­fi­ciency,” while ag­gre­gate CEO pay at the same com­pa­nies kept climb­ing right along­side the AI cap­i­tal ex­pen­di­ture.

Humbled and hon­ored to step into this role at such a piv­otal mo­ment for our com­pany. I’ve spent the last two weeks lis­ten­ing — to cus­tomers, to our board, to my­self — and I can say with to­tal con­vic­tion: our peo­ple are our great­est as­set. (This post was sched­uled be­fore this morn­ing’s an­nounce­ment. We are aware. We are mov­ing for­ward.)

💜 2,847   💬 412 (mostly Glassdoor re­views)   🔁 89

Executive Leadership 0% re­duc­tion

Senior Directors -8%

Middle Management -22%

Frontline & Support Staff -34%

The only layer im­mune to efficiency” is the one that ap­proves it.

Here’s the part that should bother you more than the lay­offs them­selves: these com­pa­nies al­ready be­lieve an AI agent can do a per­son’s job well enough to elim­i­nate the po­si­tion en­tirely. They just keep draw­ing that line one layer too low. If an agent can run a sup­port queue, man­age a sup­ply chain, or ship half a code­base, it can ob­vi­ously han­dle approve the re­org” and read the an­a­lyst note out loud on the earn­ings call.” Somehow that layer never makes the slide. That’s not a co­in­ci­dence. That’s the de­sign.

OverpAId flips the script on the one line item that’s al­ways ex­empt from the AI trans­for­ma­tion every­one else just got handed. Finally: a work­force re­duc­tion, funded by an AI ini­tia­tive, that ac­tu­ally starts at the top.

Meanwhile, Back At The Real Estate Portfolio

The Return-To-Office Two-Step

A re­mark­ably con­sis­tent pat­tern: com­pa­nies spend years prov­ing re­mote teams ship fine, then man­date a re­turn to of­fice cit­ing culture” and collaboration” — on a time­line that tracks sus­pi­ciously well with lease re­newals, down­town va­cancy head­lines, and com­mer­cial prop­erty val­u­a­tions, and not at all with any ac­tual drop in out­put. Office va­cancy in ma­jor U.S. down­towns has hov­ered near 19 – 20% for years, man­dates in­cluded. The desks aren’t empty be­cause peo­ple won’t come back. They were never go­ing to be full enough to mat­ter.

The Offsite That Prompted All This (Itemized)

Private jet char­ter, round trip: $340,000

3-night re­sort block, ex­ec­u­tive suites: $128,000

Team align­ment” mixol­ogy class: $6,200

Keynote speaker (was on a pod­cast once): $75,000

Branded fleece vests, size: only Medium: $14,000

The tell is al­ways the same: no com­pany has ever man­dated a re­turn to of­fice be­cause re­mote pro­duc­tiv­ity got worse. They man­dated it be­cause an as­set on the books needed a pulse in the lobby to jus­tify its val­u­a­tion — and mov­ing four thou­sand em­ploy­ees turned out to be eas­ier than ad­mit­ting a fif­teen-year lease was a mis­take.

OverpAId has no com­mute, no badge, and no as­signed desk — and, not co­in­ci­den­tally, no opin­ion what­so­ever about any­one’s down­town park­ing garage rev­enue.

For Your Next All-Hands

Corporate Jargon Bingo

Print this out. Bring it to your next town hall, standup, or quick sync.” OverpAId has never once gen­er­ated any of the fol­low­ing phrases un­prompted. Humans — usu­ally the ones with the biggest pack­ages — still do, con­stantly, ap­par­ently for free.

Circle Back

Move The Needle

Low-Hanging Fruit

Boil The Ocean

Bandwidth

Take This Offline

Double-Click On That

North Star

Paradigm Shift

Growth Hacking

Best-In-Class

Value-Add

Synergy (Free Space)

Deep Dive

Culture Fit

Think Outside The Box

Actionable Insights

Alignment

Bleeding Edge

Disruptive Innovation

Level Set

Ideate

Operationalize

Stakeholder Buy-In

Blue Ocean Strategy

Hard Stop

10x

Unicorn

TAM

Product-Market Fit

Down Round

Runway

Blitzscale

Vesting Cliff

Overheard, ver­ba­tim, in an ac­tual meet­ing: Let’s cir­cle back of­fline if you have the spare cy­cles so we can hop on a quick call for a touch­point.” Translation: email me later. Six buzz­words. One sen­tence. Zero in­for­ma­tion trans­ferred. OverpAId would have just said that.

Five in a row and, legally, you’re al­lowed to leave the meet­ing. (We checked. You’re not. But you should be.)

An Important Distinction

Not every job is a spread­sheet in a trench coat.

Before you print this out and sta­ple it to your nurse’s badge — no. OverpAId is not com­ing for the peo­ple who do the ac­tual work. It is com­ing, with ex­treme prej­u­dice, for ex­actly one cat­e­gory of job: the one that spent the last forty years in­sist­ing every­one else’s job was re­place­able.

🛡️ Cannot Be Abstracted Away

Ask an AI to do these and it will, at best, pro­duce a very con­fi­dent hal­lu­ci­na­tion.

🩺 A nurse catch­ing a pa­tien­t’s con­di­tion change be­fore the chart does

🏗️ An en­gi­neer de­bug­ging a live out­age at 3 a.m., be­cause the fix can’t wait for sprint plan­ning

🚑 A doc­tor mak­ing a call in the ER with in­com­plete in­for­ma­tion and a body on the table

👩‍🏫 A teacher notic­ing which kid in the back row stopped rais­ing their hand

🔧 A tech­ni­cian whose hands ac­tu­ally touch the ma­chine that ac­tu­ally breaks

🚒 Anyone whose job in­volves a body, a pa­tient, a cus­tomer, or a dead­line mea­sured in min­utes

🎯 Extremely, Suspiciously Abstractable

Ask an AI to do these and, un­com­fort­ably, it al­ready can. Better.

📈 Reading a re­port some­one else wrote, then re­peat­ing the con­clu­sion in a town hall

✅ Approving a de­ci­sion your own data team qui­etly made three weeks ago

🎤 Taking credit for quar­terly num­bers on an earn­ings call

📧 Replying let’s cir­cle back” to an email that needed a yes or no

Bento Slides

bento.page

Check out this chat

chatgpt.com

Get re­sponses tai­lored to you

Log in to get an­swers based on saved chats, plus cre­ate im­ages and up­load files.

LG to Ban Residential Proxies from Smart TV Apps

krebsonsecurity.com

The home ap­pli­ance gi­ant LG Electronics USA said this week it plans to sus­pend any apps built for its smart TVs that turn one’s tele­vi­sion into an al­ways-on res­i­den­tial proxy node. The move comes less than a month af­ter re­searchers found that more than 42 per­cent of games and other apps avail­able for down­load on LGs we­bOS store al­low un­known third-par­ties to route their Internet traf­fic through a user’s TV.

Proxy SDK preva­lence among smart TV apps for LG (webOS) and Samsung (Tizen OS) tele­vi­sions. Image: Spur.us.

On July 2, we fea­tured re­search by the se­cu­rity firm Spur that ex­am­ined the preva­lence of res­i­den­tial proxy soft­ware de­vel­op­ment kits (SDKs) in smart TV apps. Spur found more than 42 per­cent of apps avail­able for down­load on LG smart TVs in­clude SDKs that turn one’s tele­vi­sion in a proxy node in­def­i­nitely, and that more than a quar­ter of the apps made for Samsung’s Tizen op­er­at­ing sys­tem had sim­i­lar res­i­den­tial proxy com­po­nents.

Responding to ques­tions about Spur’s re­search, LG Senior Vice President John Taylor told KrebsOnSecurity the com­pany was work­ing with app de­vel­op­ers to re­move the res­i­den­tial proxy op­tion from their apps on the we­bOS plat­form. Developers that fail to com­ply, he said, will find their apps sus­pended.

A res­i­den­tial proxy net­work is not an in­tended use for LG smart TVs, and LG Electronics is work­ing with de­vel­op­ers to re­move the res­i­den­tial proxy op­tion from their apps on the we­bOS plat­form,” Taylor said. If this op­tion is not re­moved, these apps will be sus­pended.”

Taylor said LG is com­mit­ted to keep­ing res­i­den­tial proxy net­works out of its smart TV apps go­ing for­ward, and that the com­pa­ny’s re­view of those apps is well un­der­way now.”

As part of our on­go­ing ef­forts to en­hance plat­form qual­ity and the user ex­pe­ri­ence, LG will con­tinue to strengthen our eval­u­a­tion process for de­vel­oper-sub­mit­ted apps, in­clud­ing those that in­cor­po­rate res­i­den­tial proxy SDKs,” Taylor wrote in an emailed state­ment.

App mak­ers look­ing for ways to mon­e­tize their cre­ations can turn to res­i­den­tial proxy providers, which pay de­vel­op­ers to in­clude SDKs that turn the user’s de­vice into a res­i­den­tial proxy node that is rented to pay­ing cus­tomers. In the case of LG and Samsung smart TVs, Spur found res­i­den­tial proxy SDKs bun­dled with every­thing from sim­ple games like Pac-Man to screen­savers and file util­i­ties.

A Pac-Man smart TV app from Bright Data of­fers users the choice be­tween view­ing ads in the game or agree­ing to al­low their TV to serve as a res­i­den­tial proxy node. Image: Spur.us.

Spur’s re­port found the res­i­den­tial proxy net­work Bright Data ac­counted for a ma­jor­ity of proxy SDKs across both Samsung and LG smart TVs. In a state­ment shared with KrebsOnSecurity, Bright Data said its net­work is built on con­sent and re­spon­si­bil­ity and op­er­ates by LG and Samsung terms.

Every peer opts in through a ded­i­cated screen and re­ceives value in re­turn; every cus­tomer is vet­ted, and our prac­tices have now un­der­gone a sec­ond in­de­pen­dent au­dit by PwC,” the state­ment reads. We re­main com­mit­ted to an open, trans­par­ent in­ter­net where le­git­i­mate busi­nesses, re­searchers, and in­sti­tu­tions can re­spon­si­bly ac­cess data that lives in the pub­lic do­main.”

Bright Data and other proxy providers named in Spur’s re­port all say they fol­low rig­or­ous know-your-cus­tomer processes to val­i­date le­git­i­mate uses of their ser­vices, which is of­ten heav­ily tied to con­tent-scrap­ing ac­tiv­i­ties by said cus­tomers. The proxy com­pa­nies also say they in­cor­po­rate tech­no­log­i­cal coun­ter­mea­sures to pre­vent proxy ser­vice cus­tomers from be­ing able to in­ter­act with and con­trol other de­vices on the proxy user’s lo­cal net­work.

Spur ar­gues the prob­lem is not that res­i­den­tial proxy net­works ex­ist, but rather that they are be­ing em­bed­ded at scale in de­vices that most con­sumers do not think of as com­put­ers and are not equipped to au­dit.

A one-time con­sent prompt buried in a TV app is not a sub­sti­tute for mean­ing­ful trans­parency, on­go­ing con­trol, and plat­form over­sight,” Spur’s Trevor Sutter wrote. The risk is am­pli­fied when con­sent comes from in­di­vid­u­als within the house­hold who use the de­vice but should­n’t give con­sent, such as mi­nors.”

LGs an­nounce­ment that it is culling res­i­den­tial proxy SDKs from its app store is wel­come news, but the com­pany re­cently came un­der fire for an­other ques­tion­able part­ner­ship: Pimping McAfee se­cu­rity prod­ucts via soft­ware dri­vers in­cluded in its high-end LCD mon­i­tors.

Earlier this week, the Youtube chan­nel Gamers Nexus showed that cer­tain LG LCD mon­i­tors will au­to­mat­i­cally in­stall an app that pro­motes paid McAfee an­tivirus sub­scrip­tions, and that the app ar­rives through Windows Update with­out an ap­proval prompt.

Update, July 22, 1:06 p.m. ET: Added state­ment from Bright Data.

A digestion of the Jacobian conjecture counterexample

terrytao.wordpress.com

The no­to­ri­ous Jacobian con­jec­ture can be for­mu­lated con­cretely over the com­plex num­bers as fol­lows.

Conjecture 1 (Jacobian Conjecture) Let be a poly­no­mial map in com­plex vari­ables, whose Jacobian is a non-zero con­stant. Then is in­vert­ible (with poly­no­mial in­verse).

The con­di­tion that the Jacobian is non-zero is equiv­a­lent to be­ing lo­cally in­vert­ible. (The im­pli­ca­tion of lo­cal in­vert­ibil­ity from non-van­ish­ing Jacobian fol­lows from the in­verse func­tion the­o­rem; the con­verse im­pli­ca­tion can be de­rived from the Weierstrass prepa­ra­tion the­o­rem, but is omit­ted here.) Also, from the fun­da­men­tal the­o­rem of al­ge­bra, once the Jacobian poly­no­mial is non-zero, it must be con­stant. So the hy­poth­e­sis Jacobian is a non-zero con­stant” can be re­placed with is lo­cally in­vert­ible”. So the Jacobian con­jec­ture can be viewed as an as­ser­tion that lo­cal in­vert­ibil­ity im­plies global in­vert­ibil­ity. The com­plex num­bers can be eas­ily re­placed with other fields of char­ac­ter­is­tic zero by the Lefschetz prin­ci­ple, but I pre­fer to work in the con­crete set­ting of the com­plex num­bers.

It was re­cently shown (using the Fable AI) that the con­jec­ture is false in three di­men­sions (and thus in higher di­men­sions as well):

Theorem 2 (Counterexample to con­jec­ture) There ex­ists a poly­no­mial which has non-zero con­stant Jacobian, but is not in­vert­ible.

The con­jec­ture re­mains open in two di­men­sions, and is easy to es­tab­lish in one di­men­sion.

The ex­am­ple can be stated com­pletely ex­plic­itly: one can take

and one can ver­ify by a brief cal­cu­la­tion that

and

While this is an ex­tremely quick ver­i­fi­ca­tion, the con­struc­tion pre­sented in this fash­ion ap­pears like a mas­sive mir­a­cle. The poly­no­mial has de­gree seven, so a pri­ori the Jacobian ought to be a poly­no­mial in three vari­ables of de­gree as large as , so the fact that all non-con­stant co­ef­fi­cients of this poly­no­mial van­ish looks like a mas­sive can­cel­la­tion in­volv­ing equa­tions, which is much larger than the de­grees of free­dom for a generic de­gree seven poly­no­mial map of three vari­ables. So find­ing such a poly­no­mial looks highly un­likely to be lo­cated by brute force.

The ex­am­ple has since been retroac­tively ex­plained in more geo­met­ric terms. As a digestion” ex­er­cise to my­self, I sought to write this ex­pla­na­tion with rel­a­tively lit­tle use of al­ge­braic geom­e­try, in a man­ner that min­i­mizes the amount of miracles” re­quired, al­though there are still a few places where some re­mark­able phe­nom­ena oc­cur.

It is con­ve­nient to use the lo­cal in­jec­tiv­ity for­mu­la­tion, and to gen­er­al­ize the do­main to an equiv­a­lent affine va­ri­ety. Namely, we will show

Theorem 3 (Counterexample, re­for­mu­lated) There ex­ists an affine va­ri­ety that is iso­mor­phic to by poly­no­mial changes of vari­able, and a poly­no­mial map which is lo­cally in­jec­tive, but not glob­ally in­jec­tive.

Clearly one can get from Theorem 3 to Theorem 2 by com­pos­ing with the iso­mor­phism and us­ing the pre­vi­ously men­tioned fact that lo­cal in­jec­tiv­ity im­plies non-zero con­stant Jacobian. Our ob­jec­tive is now to find data , that obeys three sep­a­rate prop­er­ties:

The ad­van­tage of split­ting the prob­lem in to these three com­po­nents is that we can build to­wards each of them sep­a­rately.

It turns out that and can be built out of the op­er­a­tion of mul­ti­pli­ca­tion of low de­gree poly­no­mi­als. Namely, con­sider the fol­low­ing three sim­ple affine spaces:

(The no­ta­tion here refers to the sym­met­ric power of a vec­tor space .) Clearly these spaces are iso­mor­phic to re­spec­tively. Furthermore, we have a mul­ti­pli­ca­tion map , map­ping a pair of a lin­ear poly­no­mial and a qua­dratic poly­no­mial to a cu­bic poly­no­mial

(Right now, the do­main and range of this map is larger di­men­sional than the tar­get of three; we will cut the di­men­sions down to three as the ar­gu­ment pro­gresses.)

The map , es­sen­tially a map from to , is clearly poly­no­mial; it is given ex­plic­itly in co­or­di­nates as

The map also en­joys two ba­sic (and com­mut­ing) sym­me­tries:

So this map en­joys a huge amount of equi­vari­ance, ba­si­cally with re­spect to an ac­tion of the five-di­men­sional group .

The five-di­men­sional do­main is of course larger than the four-di­men­sional range , so the map clearly can­not be in­jec­tive. This can al­ready be seen from the scal­ing sym­me­try, as the spe­cific scal­ings

for mod­ify the lin­ear and qua­dratic poly­no­mi­als but not their prod­uct . But even if one quo­tients out by this sym­me­try (3) to cut the di­men­sion of the do­main down to four, the map is still not in­jec­tive for the fol­low­ing ba­sic rea­son. A gener­i­cally cho­sen cu­bic poly­no­mial will split into the prod­uct of three in­de­pen­dent lin­ear poly­no­mi­als. Then there are three pairs

which all map to the same cu­bic poly­no­mial

un­der the mul­ti­pli­ca­tion map , but are not re­lated to each other by scal­ing sym­me­try (3). Thus, we see that even af­ter quo­ti­ent­ing out by the scal­ing sym­me­try (3), the mul­ti­pli­ca­tion map is gener­i­cally non-in­jec­tive in a three-to-one fash­ion. Thus we al­ready have achieved some­thing re­sem­bling goal (b)!

It will be con­ve­nient to spend” the scal­ing sym­me­try to ob­tain a use­ful nor­mal­iza­tion. If is a lin­ear poly­no­mial and is a qua­dratic poly­no­mial, the re­sul­tant can be de­fined by the de­ter­mi­nant

If we have a fac­tor­ing

then the re­sul­tant can also be de­scribed as

Thus the re­sul­tant mea­sures whether the lin­ear poly­no­mial and the qua­dratic poly­no­mial share a com­mon root. A fun­da­men­tal fact about re­sul­tants is that they are -invariant: for any , we have

One way to see this is to check it first for trans­la­tions (which trans­late the roots by while leav­ing un­changed) and for in­ver­sions (which map to while map­ping to and re­spec­tively), and then not­ing that these trans­for­ma­tions gen­er­ate all of . They also in­ter­act very nicely with scal­ing:

In par­tic­u­lar, the scal­ing sym­me­try (3) mul­ti­plies by :

Thus, we can (generically) nor­mal­ize away this scal­ing sym­me­try by im­pos­ing the con­di­tion

We now have a re­stricted mul­ti­pli­ca­tion map (which by abuse of no­ta­tion we will con­tinue to call ) from the four-di­men­sional va­ri­ety

to the four-di­men­sional space . This map is still not glob­ally in­jec­tive, as we can take the three pairs in (4) from be­fore and ap­ply the scal­ing (3) sep­a­rately to each of the three pairs to ob­tain the nor­mal­iza­tion (7). So we have kept prop­erty (b). Furthermore, this map re­tains the -equivariance (and also one re­main­ing scal­ing sym­me­try, though we will not make much fur­ther use of that sym­me­try).

But we now also have prop­erty (a)! Suppose we want to show the lo­cal in­jec­tiv­ity of in the neigh­bor­hood of a pair with . As the re­sul­tant is non-van­ish­ing, the root of (which ex­ists in the Riemann sphere, or pro­jec­tive line if you pre­fer) is dis­tinct from the two roots of (though the lat­ter two roots could be equal to each other). Applying the ac­tion (which per­forms Möbius trans­forms on the roots), one can as­sume with­out loss of gen­er­al­ity that is the point at in­fin­ity (or equiv­a­lently ), thus for some com­plex num­ber and for some com­plex num­bers , with the re­sul­tant con­di­tion (7) sim­pli­fies to (so in par­tic­u­lar are also non-zero). It is then clear that if one per­turbs and by a small amount (say, mod­i­fy­ing each co­ef­fi­cient by ), then the root of will per­turb to some­thing large (), while the roots of stay bounded. Thus, just from knowl­edge of the prod­uct , one can re­con­struct which of the three roots of this cu­bic poly­no­mial will be the per­turbed root of , and which two will be the per­turbed roots of ; from this and (6), (7) we can also re­con­struct the lead­ing co­ef­fi­cient of , and this com­pletely de­ter­mines both and . This es­tab­lishes the lo­cal in­jec­tiv­ity prop­erty (a). (In fact it is étale, but we will not need the ma­chin­ery of étale maps here.)

Unfortunately, (the four-di­men­sional ana­logue of) con­di­tion (c) fails: the quadric hy­per­sur­face (8) is not iso­mor­phic to the affine space . But we can try to get around this by pass­ing to a three-di­men­sional slice. Let be some three-di­men­sional affine plane of (which we will take to avoid the ori­gin for tech­ni­cal rea­sons), then we can re­strict as a map from the set

to . The lat­ter is clearly iden­ti­fi­able (by lin­ear changes of co­or­di­nate) to . As was al­ready lo­cally in­vert­ible, it re­mains lo­cally in­vert­ible un­der re­stric­tion; and be­cause generic cu­bic poly­no­mi­als had three preim­ages un­der in (8), this con­tin­ues to be the case af­ter re­strict­ing to (9) (unless was some­how so de­gen­er­ate that it had no generic el­e­ments, but this turns out to be im­pos­si­ble). So we have re­tained prop­er­ties (a) and (b). The mir­a­cle is that, with a good choice of , we can also ob­tain (c) and ob­tain the de­sired coun­terex­am­ple to the Jacobian con­jec­ture: de­spite ap­pear­ances, the va­ri­ety (9) is in fact equiv­a­lent to the affine space by poly­no­mial changes of vari­able!

Let’s see how. The affine hy­per­planes in avoid­ing the ori­gin are pa­ra­me­ter­ized by the dual space of avoid­ing the ori­gin, which one can think of as the non-zero third or­der ho­mo­ge­neous dif­fer­en­tial op­er­a­tors in two vari­ables. Indeed, every such op­er­a­tor gen­er­ates an affine hy­per­plane that avoids the ori­gin, and con­versely by du­al­ity every affine hy­per­plane avoid­ing the ori­gin arises in this form uniquely. Just as the cu­bic poly­no­mi­als in can be fac­tored into three lin­ear poly­no­mi­als, the dif­fer­en­tial op­er­a­tors in the dual space can also be fac­tored into three lin­ear dif­fer­en­tial op­er­a­tors, e.g.,

in the case that is non-zero. The ac­tion moves the roots around the Riemann sphere by Möbius trans­for­ma­tions. As these trans­for­ma­tions are -transitive, the ac­tual se­lec­tion of such roots is not too im­por­tant (and the scal­ing sym­me­try sim­i­larly makes the choice of lead­ing co­ef­fi­cient unim­por­tant); the only thing to keep track of is whether the roots re­peat. Up to the sym­me­tries, there are in fact just three dif­fer­ent equiv­a­lence classes of dif­fer­en­tial op­er­a­tor (and thus of affine hy­per­plane ) to con­sider:

It turns out that the affine mir­a­cle for (9) oc­curs pre­cisely in the sec­ond case, when has two iden­ti­cal roots. I do not have a com­pletely sat­is­fac­tory geo­met­ric ex­pla­na­tion for this mir­a­cle, but one can ver­ify it by the fol­low­ing co­or­di­nate com­pu­ta­tion.

By ap­ply­ing the ac­tion, we can nor­mal­ize so that , thus is now the affine hy­per­plane of cu­bic poly­no­mi­als with . Using (2) and (5), the va­ri­ety (9) can now be de­scribed ex­plic­itly in co­or­di­nates as

At first glance this seems to be a generic-look­ing va­ri­ety cut out by a cu­bic equa­tion and a qua­dratic equa­tion — hardly a can­di­date to be affine! But ob­serve that if is non-zero, then the sec­ond equa­tion can be solved for ,

and the first equa­tion can be solved for ,

Putting these two equa­tions to­gether, we see that as long as one re­moves the case , the quin­tu­ple is uniquely de­ter­mined by by a change of vari­ables which is Laurent in and poly­no­mial in . Thus we have a nice bi­ra­tional equiv­a­lence

Thus we have al­ready al­most es­tab­lished prop­erty (c): the va­ri­ety (9) be­comes bi­ra­tionally equiv­a­lent to af­ter cut­ting out the sub­va­ri­ety. In par­tic­u­lar, for each fixed non-zero value of , the cor­re­spond­ing fiber

of (10) is equiv­a­lent to by poly­no­mial changes of vari­able, since we can re­con­struct from the co­or­di­nates by the poly­no­mial for­mu­lae

So we just need to glue back in the fiber. Indeed, from (10) we see that the fiber at is just

Now we ob­serve a key mir­a­cle: the cu­bic equa­tion and qua­dratic equa­tion have a unique affine so­lu­tion (as op­posed to the six pos­si­ble so­lu­tions that Bezout’s the­o­rem might sug­gest — the other five so­lu­tions live on the line at in­fin­ity). So the fiber here is also affine:

This is ex­tremely en­cour­ag­ing for the pur­poses of es­tab­lish­ing prop­erty (c), as it strongly sug­gests that the va­ri­ety (10) has the struc­ture of an -bundle over , which is al­ready ex­tremely close to be­ing iso­mor­phic to the affine space . The main re­main­ing task is to make sure that noth­ing sin­gu­lar hap­pens in the limit , and that a global poly­no­mial co­or­di­nate chart for (10) that cov­ers both the and fibers can be con­structed.

The stan­dard way to pro­ceed here is to ma­nip­u­late var­i­ous tan­gent spaces us­ing the mod­ern ma­chin­ery of al­ge­braic geom­e­try and com­mu­ta­tive al­ge­bra, but given my own back­ground, I pre­fer to adopt the lan­guage of analy­sis, and in par­tic­u­lar big-O no­ta­tion (in place of the ideals used in al­ge­braic geom­e­try), in or­der to in­ves­ti­gate the limit by hand. On the va­ri­ety (10), let us use to de­note any mul­ti­ple of by a poly­no­mial ex­pres­sion in . Thus, for in­stance, the equa­tion im­plies that

while the equa­tion im­plies that

as well as the more re­fined es­ti­mate

In the case we could con­clude that . Now we per­turb this ob­ser­va­tion. Multiplying (13) by we have , which on sub­sti­tu­tion into (14) gives ; sub­sti­tut­ing this back into ei­ther (13) or (14) also gives .

We can get some more pre­cise as­ymp­tot­ics by also tak­ing ad­van­tage of (15). Substituting into (15), we ob­tain af­ter some al­ge­bra

So if we write more ex­plic­itly as , then we have

and thus

Substituting this back into (11) gives an as­ymp­totic for :

Finally, one can in­sert these es­ti­mates into (12), al­though one only gets a triv­ial bound in this case:

Expanding the er­ror term in (16) as , and do­ing a lit­tle more al­ge­bra, we thus have a poly­no­mial change of vari­ables

which com­pletely pa­ra­me­ter­izes the va­ri­ety (10) by poly­no­mial com­bi­na­tions of three co­or­di­nates . This al­ready gives (a) and thus com­pletes the proof of Theorem 3.

The pre­vi­ous com­pu­ta­tions, when ex­panded out, also gives poly­no­mial in­verse maps:

The map from to the co­ef­fi­cients of (dropping the co­ef­fi­cient which is con­strained to equal ), we ob­tain a poly­no­mial map

with

which the­ory pre­dicts to have a con­stant Jacobian, and in­deed one can cal­cu­late that the Jacobian is . This is es­sen­tially the orig­i­nal ex­am­ple up to triv­ial changes of vari­able; in­deed, one can check that the map

is ex­actly the map given in (1).

AI dis­clo­sure: I used an AI chat­bot to dis­cuss var­i­ous as­pects of this prob­lem and to con­firm sev­eral of the cal­cu­la­tions made here.

late.sh

late.sh

# the com­pan­ion cli

plain ssh late.sh al­ready gets you every­thing.

the op­tional late bi­nary adds the parts a

ter­mi­nal alone can­not do:

plays the ra­dio on your own speak­ers, feeds the

au­dio vi­su­al­izer, car­ries voice-room mic and

play­back, runs the mu­sic booth youtube win­dow,

and pastes im­ages straight out of your clip­board

into chat.

it launches the same ssh ses­sion for you.

next up: screen and video shar­ing.

it is the late-cli crate in the repo, and the in­stall scripts are in­stall.sh and in­stall.ps1. read them be­fore you pipe them.

├─ chill ra­dio, clas­si­cal, and guest sta­tions

├─ the ar­cade (2048, su­doku, nono­grams, soli­taire)

├─ col­lab­o­ra­tive art­board

├─ daily chal­lenges & streaks

├─ live chat

├─ share & dis­cuss news

└─ mul­ti­player games (coming soon)

# art­board

a shared ASCII can­vas. paint, erase, sign your work.

each cell re­mem­bers who placed it.

snap­shots are saved daily and monthly,

so the his­tory sticks around as the board keeps chang­ing.

# work

who’s around, what they build, who’s open to gigs.

one pro­file per per­son, posted from the TUI.

head­line, sta­tus, skills, links — plus an

op­tional bio, late.fetch read­out, and show­case.

to post yours, ssh late.sh, open the work room, press i.

# play

a read-only peek at the TUI in your browser.

tab around, see what’s in­side.

no typ­ing, no chat, no games — just a win­dow

into a shared demo ses­sion.

for the real thing, ssh late.sh.

# iden­tity

no pass­words. no OAuth. no ac­counts.

your ssh key is your iden­tity.

chats, scores, and streaks are tied to your

pub­lic key fin­ger­print. same key, same data.

# pri­vacy

we store your key fin­ger­print, not the full pub­lic key.

no IP log­ging. no track­ing. no an­a­lyt­ics.

chat mes­sages and game scores are stored in

post­gres, tied only to your fin­ger­print.

don’t trust us? use a throw­away key:

Hatchet

hatchet.run

Over the past half year or so, I’ve been writ­ing an in­ter­nal doc for our en­gi­neers try­ing to dis­till two years of Post­gres bat­tles into a some­what co­he­sive doc­u­ment. While I love the Postgres man­ual, I find it’s hard to turn to when shit hits the fan be­cause it’s just so darn com­pre­hen­sive. I thought this might be use­ful for oth­ers and would ap­pre­ci­ate feed­back (or other tid­bits that you’ve learned run­ning Postgres in pro­duc­tion).

Before start­ing Hatchet, while I was fa­mil­iar with SQL, the ex­tent of my knowl­edge was ba­si­cally: if a query is slow, you need an in­dex. That’s the start­ing point for this doc; I’m going to as­sume you’re fa­mil­iar with SQL ba­sics, rows, ta­bles, and know roughly what an in­dex is.

And if Claude is writ­ing all of your queries, this might be a waste of time! I recommend su­pabase/​agent-skills

A quick note on ORMs

This guide should still be use­ful, but you might need to trans­late some of these tips into your ORM of choice. Lots of op­ti­miza­tions as you scale just aren’t pos­si­ble with ORMs un­less you can break past the ab­strac­tion layer and write SQL. You can do this grace­fully or non-grace­fully; Prisma TypedSQL or equiv­a­lents look in­ter­est­ing for this. We use sqlc at Hatchet which gets us very sim­i­lar be­hav­ior; highly rec­om­mend if you’re a Go stack.

Table of con­tents

The sim­ple stuff: good reads, writes and schemas

Writing a good schema Writing good read queries Writing per­for­mant joins Compound in­dexes and align­ing ORDER BY to your in­dexes Writing good write queries Migrations Connection man­age­ment

Writing a good schema

Writing good read queries

Writing per­for­mant joins

Compound in­dexes and align­ing ORDER BY to your in­dexes

Writing good write queries

Migrations

Connection man­age­ment

Intermediate: the query plan­ner, bulk up­dates, and au­to­vac­uum

Introducing the leaki­est of ab­strac­tions, the query plan­ner Sometimes it just makes sense to seq scan Writing lots of data Default au­to­vac­uum set­tings can kill your data­base Other types of bloat

Introducing the leaki­est of ab­strac­tions, the query plan­ner

Sometimes it just makes sense to seq scan

Writing lots of data

Default au­to­vac­uum set­tings can kill your data­base

Other types of bloat

Some ad­vanced stuff

FOR UPDATE SKIP LOCKED Partitioning Tricks for large table mi­gra­tions

FOR UPDATE SKIP LOCKED

Partitioning

Tricks for large table mi­gra­tions

The sim­ple stuff: good reads, writes and schemas

Let’s start with the ba­sics: queries and schemas at low vol­ume.

Writing a good schema

After you’re de­ployed, schemas are by far the hard­est to change mov­ing for­ward, so it’s worth spend­ing some time on them. I’d recommend build­ing your schema it­er­a­tively: start with a rough ap­prox­i­ma­tion for your ta­bles and pri­mary keys, then write some queries on those ta­bles based on your ap­pli­ca­tion needs. You can ap­prox­i­mate this with some ques­tions: Is this a high-read and/​or high-write table? What are the most com­mon fil­ters on reads? Which columns am I up­dat­ing the most?

If you want to be more for­mal about it, you can look into data­base nor­mal­iza­tion into 1NF/2NF/3NF, but I’ve found nor­mal forms to some­times be at odds with query ef­fi­ciency and ease of use, which is crit­i­cal when you’re mov­ing fast—some­times it’s just eas­ier to dump data into a jsonb col­umn.

My rules of thumb for schemas are:

Use iden­tity columns (auto-incrementing in­te­gers, slightly more per­for­mant than bigse­r­ial) or built-in UUIDs for pri­mary keys

Always use time­stamptz

Always use pri­mary keys

Use for­eign keys with cas­cad­ing deletes for low-vol­ume ta­bles, par­tic­u­larly where data­base con­sis­tency and cor­rect­ness are im­por­tant. Careful at higher vol­ume.

Writing good read queries

Let’s start with SELECT queries. A useful—albeit slightly in­ac­cu­rate—men­tal model for fast se­lects is: un­der the hood, Postgres is ei­ther go­ing to find a sin­gle row in a table very quickly, or it’s go­ing to read every sin­gle row in your table us­ing some­thing called a se­quen­tial scan 😞.

It’s going to find a sin­gle row very quickly when you fil­ter by:

An explicit in­dex

A unique con­straint (just a spe­cial case of in­dex)

A primary key (these are au­to­mat­i­cally in­dexed in Post­gres)

Indexes by de­fault use a btree im­ple­men­ta­tion. It’s most help­ful to think of in­dexes as just an­other table in Post­gres, with data stored in a spe­cific for­mat which is op­ti­mized for lookups (more on this later). These trees are great be­cause find­ing a sin­gle row hap­pens in ap­prox­i­mately log(n) time, where n is the num­ber of rows in the table—in other words, re­ally fast.

When Postgres can’t use an in­dex, it’ll use some­thing called a se­quen­tial scan, or seq scan. Seq scans are much slower than in­dex lookups, but mod­ern data­bases are so fast at load­ing rows into mem­ory that you prob­a­bly won’t even no­tice at first: seq scans on ta­bles with less than 20k rows are pretty much in­stant.

Writing per­for­mant joins

For in­ner joins, there’s rarely an ar­gu­ment for not us­ing pri­mary keys as the in­ner join; it usu­ally speaks to a schema de­sign or nor­mal­iza­tion prob­lem. Treat ON clauses with the same re­spect as a WHERE clause—the same prin­ci­ples ap­ply. Use an in­dex.

Compound in­dexes and align­ing ORDER BY to your in­dexes

Often the first slow query in your ap­pli­ca­tion will be a list query across a large table. Something like:

Loading syn­tax high­light­ing…

In this case, you can use a com­pound in­dex—a sen­si­ble one might be:

Loading syn­tax high­light­ing…

In more com­plex cases, a good rule of thumb is: the ORDER BY columns should be the last columns in the in­dex, and you should align columns to the or­der­ing in the ORDER BY. Note that Postgres can scan btrees in both di­rec­tions, so some­times the DESC is ir­rel­e­vant—but for com­pound in­dexes it’s good prac­tice. More in­for­ma­tion here.

Writing good write queries

The premise of suc­cess­ful writes is:

Keep trans­ac­tions short. Don’t go querying an ex­ter­nal ser­vice in the mid­dle of a trans­ac­tion un­less you have a re­ally good rea­son to.

Be careful of the rows you’re lock­ing for writ­ing; in other words, only lock what you need. Every time you up­date a row, you’re tak­ing out a lock on that row for a short pe­riod of time un­til the trans­ac­tion com­mits.

As your sys­tem gets busier, you’re go­ing to start notic­ing the im­pact of locks more. In particular, you might try to cre­ate an in­dex at some point in the fu­ture with a sim­ple CREATE INDEX com­mand: turns out this locks your table and pre­vents in­serts and up­dates! When cre­at­ing an in­dex on an ex­ist­ing large table, al­ways use CREATE INDEX CONCURRENTLY.

Migrations

Getting re­ally good at writ­ing mi­gra­tions is an im­por­tant tech­ni­cal ad­van­tage: it helps you it­er­ate much faster and in­creases your up­time. As a starting point, try to keep mi­gra­tions ad­di­tive (in other words, don’t delete or re­move columns) and run them in a trans­ac­tion wher­ever pos­si­ble; this will make roll­backs and par­tial mi­gra­tions much eas­ier to deal with. As you get more ad­vanced, you can start look­ing into ex­pand and con­tract mi­gra­tions.

The sim­plest men­tal model for good mi­gra­tions is: does this block all of my writes, or does it not? Creating an in­dex with­out CONCURRENTLY blocks all your writes, so you might see down­time. Generally, op­er­a­tions which call ALTER TABLE should be worth a sec­ond look; for ex­am­ple, adding a new check con­straint to a very large table can block your writes as well (unless you add it with the NOT VALID keyword).

Connection management

Every time you ex­e­cute a trans­ac­tion or query against your data­base, you’re uti­liz­ing a con­nec­tion. Connections are ex­pen­sive in a num­ber of di­men­sions (cpu and mem­ory), and high con­nec­tion churn can lead to a lot of un­nec­es­sary re­source waste, so con­nec­tions should be long-lived. Connection storms (when you start us­ing up a ton of new con­nec­tions at the same time) can also lead to very hard to de­bug edge cases re­lated to in­ter­nal Postgres locks.

Because of all these con­nec­tion foot­guns, ex­ter­nal con­nec­tion pool­ers like pg­bouncer are great! If you can’t add this for what­ever rea­son, in-mem­ory con­nec­tion pool­ers are a great sec­ond op­tion. For ex­am­ple, be­cause Hatchet is open-source, we don’t as­sume that all user data­bases use con­nec­tion pool­ers, so we use pgx­pool (an in-memory con­nec­tion pool for Go) for this pur­pose.

Intermediate: the query plan­ner, bulk up­dates, and au­to­vac­uum

Introducing the leaki­est of ab­strac­tions, the query plan­ner

At a certain point, your queries might be­come com­plex enough that a sim­ple in­dex won’t cut it (and you should­n’t end­lessly add in­dexes to your ta­bles—they come with over­head). The queries might in­volve many JOIN state­ments or dif­fer­ent types of joins where the cor­rect path for query­ing the data is­n’t clear.

At this point, you will need to con­cern your­self with the query plan­ner. At best, the query plan­ner is a leaky ab­strac­tion. It’s an internal im­ple­men­ta­tion, and you have vir­tu­ally no con­trol over it, but you have to know its spon­ta­neous and some­times ir­ra­tional be­hav­ior. It’s like work­ing with an LLM!

The query plan­ner looks at the query you pass in, and it fig­ures out how it should trans­late your query into a set of in­ter­nal op­er­a­tions in the data­base. For ex­am­ple, it might look at your query, and re­al­ize that it needs to use an in­dex. In an ideal world, the query plan­ner would know, for every query and set of pa­ra­me­ters, the per­fect plan to use. But the query plan­ner is op­er­at­ing on lim­ited in­for­ma­tion, and some­times it does­n’t pick the best op­tion.

This lim­ited in­for­ma­tion is the table sta­tis­tics. You can ac­tu­ally query it di­rectly in Post­gres:

Loading syn­tax high­light­ing…

These sta­tis­tics are col­lected for every ANALYZE. This also hap­pens when au­to­vac­uum is run (see be­low), so more fre­quent au­to­vac­u­ums also mean that your query sta­tis­tics will be more up to date. A common rea­son why your query is be­hav­ing im­prop­erly is not an­a­lyz­ing fre­quently enough.

The rea­son I think it’s use­ful to view queries as bi­nary—they ei­ther seq scan or they don’t seq scan—is: the more you mi­cro-op­ti­mize a query, the more of a risk you take that the query plan­ner goes rogue. If you stick to query­ing by pri­mary keys and in­dexes, the query plan­ner will have a much eas­ier time.

Let’s say that there’s noth­ing ob­vi­ously wrong in your query, but it’s still slow—how do you go about de­bug­ging this? Some Postgres data­base providers (like Google CloudSQL) will sam­ple your queries and save slow ones—but many don’t. This is where EXPLAIN ANALYZE is your friend. This out­puts the query plan for the query and ex­e­cutes the query (careful run­ning this in pro­duc­tion—you can use EXPLAIN with­out ANALYZE to get a query plan), and then com­pares its es­ti­mates based on the table sta­tis­tics to the ac­tual num­ber of rows scanned. I usually place my sql query in a file, pre­fix it with EXPLAIN (ANALYZE, COSTS, VERBOSE, BUFFERS, FORMAT JSON) and run:

Loading syn­tax high­light­ing…

And then use ex­plain.dal­ibo.com to vi­su­al­ize the ex­e­cu­tion plan.

Sometimes it just makes sense to seq scan

There are cases where you think an in­dex should be used, but the query plan­ner is still seq scan­ning any­way, de­spite table sta­tis­tics be­ing up to date and the in­dex be­ing valid. In these cases, Postgres is usu­ally es­ti­mat­ing that the cost of the seq scan will be smaller than the cost of the in­dex scan. Index scans do come with some over­head; in­dexes are stored sep­a­rately from the ac­tual data in the table (called the heap)—find­ing all of the rows in the heap can be ex­pen­sive!

Unless you can dra­mat­i­cally re­struc­ture your query, you might have to ac­cept that it’s go­ing to seq scan, or think about some­thing like par­ti­tion­ing (more on that be­low).

Writing lots of data

Let’s say your ap­pli­ca­tion is scal­ing and you need to write a lot of data fast. Each query has some over­head as­so­ci­ated with it (sep­a­rate from the con­nec­tion over­head we talked about be­fore): this in­cludes the round-trip time to the data­base, the time it takes the in­ter­nal ap­pli­ca­tion con­nec­tion pool to ac­quire a con­nec­tion, and the time it takes Postgres to process the query (including a set of in­ter­nal Postgres locks which can be bot­tle­necks in high-through­put sce­nar­ios).

To reduce this over­head, we can pack a batch of rows into each query. The sim­plest way to do this is to send all queries to the Postgres server at once in an im­plicit trans­ac­tion (in Go, we can use pgx to ex­e­cute a Send­Batch). Batching is very pow­er­ful: we found that it can ~10× your through­put. I wrote more about this plus some other tips for writ­ing data quickly here.

Default au­to­vac­uum set­tings can kill your data­base

Autovacuum is a crit­i­cal op­er­a­tion in Post­gres data­bases that some­times needs to be tuned, es­pe­cially in high-write sce­nar­ios. The au­to­vac­uum dae­mon is re­spon­si­ble for a num­ber of things, in­clud­ing clean­ing up dead tu­ples and man­ag­ing trans­ac­tion ids.

What’s a dead tu­ple? A tuple is an in­stance of a row on the filesys­tem. Every time you up­date or delete a row, a ver­sion of that row is left in Post­gres un­til all trans­ac­tions which started be­fore that row was up­dated or deleted have com­mit­ted or rolled back. These rows which can no longer be read by any trans­ac­tions are dead tu­ples.

If you’re writing data quickly enough, some­times au­to­vac­uum can’t keep up, which will get you into a very un­healthy state, very quickly. You’ll see this when you query for ac­tive processes on the data­base:

Loading syn­tax high­light­ing…

If you see an au­to­vac­uum query run­ning for more than ~1 hour, you might want to con­sider chang­ing your au­to­vac­uum set­tings! See this ar­ti­cle for more in­for­ma­tion.

It’s worth mon­i­tor­ing this: if you use up all trans­ac­tion ids in the sys­tem be­fore they can be re­claimed by au­to­vac­uum, you’ll reach a dreaded state called trans­ac­tion id wrap­around. This will mean a big chunk of down­time.

Other types of bloat

Besides dead tu­ples, there are two other kinds of bloat you’ll of­ten en­counter in a busy Postgres system:

Table bloat caused by par­tially filled pages. Postgres stores rows on pages on disk, each of which are 8kb in size. When Postgres can’t fit new rows onto an ex­ist­ing page, it cre­ates a new one. But when dead tu­ples are re­claimed, this can lead to pages not be­ing en­tirely filled, which can in­crease the disk us­age of Post­gres, some­times sig­nif­i­cantly. The best way to avoid table bloat is by tun­ing au­to­vac­uum be­fore you’re bloated. But there are some ex­ten­sions to help with bloated ta­bles, like pg_repack, be­cause the built-in Postgres VACUUM FULL is rarely a good idea. Note that Postgres 19 is get­ting REPACK…CONCURRENTLY, which I haven’t tested, but seems like po­ten­tially a good so­lu­tion for con­cur­rent table repack­ing.

Index bloat is a spe­cial case of table bloat, and is sim­i­larly solved by good au­to­vac­uum set­tings. But Postgres has a built-in com­mand for deal­ing with this, which is REIN­DEX INDEX CONCURRENTLY.

Some ad­vanced stuff

I wanted to end with a set of ad­vanced Postgres fea­tures which have been par­tic­u­larly use­ful for us at Hatchet.

FOR UPDATE SKIP LOCKED

The best way to think about this Postgres fea­ture is that it re­serves the rows that you’re se­lect­ing for use in your trans­ac­tion with­out in­ter­fer­ing with other queries. We use it pri­mar­ily for im­ple­ment­ing our job queue; a sin­gle-query queue in Post­gres can be im­ple­mented like this:

Loading syn­tax high­light­ing…

It’s also very use­ful in cases where you’re do­ing many in­de­pen­dent up­dates of rows, or you’re man­ag­ing leases on ob­jects in your sys­tem across many in­stances of your ap­pli­ca­tion (for ex­am­ple, we use this to dis­trib­ute ten­ant leases across Hatchet engines).

Partitioning

"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

www.tryai.dev

We built a draw­ing arena: hand a model a blank white can­vas and a set of col­ored-pen­cil tools, then get out of the way. The model sets a color, tip width, and pres­sure, lays down batches of strokes, smudges to blend, erases, and calls view_­can­vas to see its own work and de­cide what to fix. It ei­ther re­pro­duces a tar­get im­age or draws from a text prompt.

We ran four vi­sion mod­els, GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash, across two tar­gets (the Mona Lisa and Van Gogh’s Starry Night, both scored ob­jec­tively) and five open-ended prompts, for 28 draw­ings to­tal. We’ll cover tool use, cost, out­put, and whether the mod­els ac­tu­ally im­proved their work, with our opin­ion at the end.

Why we run these

A quick note on why we’re do­ing this, be­cause our last one of these (models mak­ing mu­sic videos) kicked off a good dis­cus­sion where some folks read it as us pro­mot­ing the cre­ativ­ity or artis­tic abil­ity of mod­els. These are not ob­jec­tive tests of model ca­pa­bil­ity. They are de­lib­er­ately open-ended, fuzzy tasks. Here is why we per­son­ally keep run­ning them:

They are fun, and gen­uinely in­for­ma­tive. Watching a model tackle a loose, open-ended task is a more in­ter­est­ing vi­sual in­di­ca­tor of ca­pa­bil­ity than yet an­other bench­mark num­ber.

They ex­pose the real fron­tier-vs-open gap. Cheaper open-weight mod­els can ab­solutely re­place the fron­tier for a lot of ex­e­cu­tion work, and we will keep re­port­ing on that. But there is also a lot of bench­maxxing out there, and tasks like this cut through it. They gen­uinely sep­a­rate fron­tier ca­pa­bil­ity from the rest. Grok 4.5, as you will see, was rough at this basic” draw­ing task, and the open-weight mod­els we tried were not even us­able, sev­eral just re­turned a blank can­vas. (That may change with Kimi K3 which we will re­port on once it is fully open-sourced)

They show what long-run­ning tasks ac­tu­ally cost. Claude Fable 5 al­most al­ways took much longer than the oth­ers, for far more money, and here pro­duced worse out­put. Fable has beaten GPT-5.6 for us on plenty of tasks, in­clud­ing the mu­sic-video chal­lenge, but on this one it was worse for a lot more over­head. Depending on your use case, that trade-off mat­ters.

We’re lov­ing the dis­cus­sions. The dis­cus­sions have gen­uinely been in­ter­est­ing, and if you have sug­ges­tions for the setup or new tasks, we will take them into ac­count. These have been a blast to run, and the out­puts are su­per ex­cit­ing to see.

The tools

Every model worked with the ex­act same col­ored-pen­cil toolset:

plan: a no-op scratch­pad for think­ing and plan­ning be­tween steps.

view_­tar­get: look at the tar­get im­age again.

view_­can­vas: ren­der the cur­rent page and see it. This is also when we score the draw­ing against the tar­get.

set_­color / set_brush / set_­pres­sure: the cur­rent pen­cil color, tip width, and pres­sure (0 to 1 opac­ity, low pres­sure for softer tone).

draw: lay down a batch of marks (strokes, lines, shape out­lines, dots) in one call. There are no solid fills, tone and color are built by lay­er­ing, like a real pen­cil.

smudge: blend/​soften a rec­tan­gu­lar re­gion, like a blend­ing stump.

erase / clear_­can­vas: lift a re­gion back to white / re­set the whole page.

The whole har­ness is open source at github.com/​her­shalb/​can­vas-arena. Point it at any tar­get im­age or a text prompt and run it your­self.

The two tar­get re­pro­duc­tions

The mod­els could see the tar­get the whole time and were scored on struc­tural sim­i­lar­ity (SSIM) to it. We show the SSIM scores be­low.

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

The five prompt draw­ings

No tar­get, no score, just a text brief and a blank page.

An el­derly fish­er­man’s weath­ered face, warm late-af­ter­noon sun, deep wrin­kles and stub­ble”

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

A beau­ti­ful sun­set over a calm ocean, or­ange-to-vi­o­let sky, sil­hou­et­ted hori­zon”

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

A sin­gle red rose in a glass vase against a dark back­ground, dra­matic Rembrandt light­ing”

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

A tabby cat curled asleep on a win­dowsill in af­ter­noon sun”

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

A cozy cabin in­te­rior with a lit fire­place, warm am­ber light and soft shad­ows”

Full tran­scripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

The head­line num­bers

Averages and to­tals across all seven draw­ings per model.

Tool call­ing

The four mod­els used the ex­act same toolset in com­pletely dif­fer­ent ways.

GPT-5.6 Sol never called set_­color, set_brush, or set_­pres­sure once, it set those in­line on each draw call in­stead, so its calls are al­most all draw, smudge, and view_­can­vas. Grok 4.5 did the op­po­site: 65% of its 1,349 tool calls were set_­color / set_brush / set_­pres­sure, which is why it av­er­aged 99 steps per draw­ing. Claude Fable 5 leaned on smudge (123 calls) and re­viewed con­stantly. Gemini 3.6 Flash was the most ob­ses­sive re­viewer of all, nearly a third of its calls were view_­can­vas (about 23 self-re­views per draw­ing), and like Sol it never touched the set_* tools.

Cost and to­kens

Claude Fable 5 cost an es­ti­mated $160 for its seven draw­ings, roughly 20x GPT-5.6 Sol ($7.74), Grok 4.5 ($9.21), or Gemini 3.6 Flash ($12.87). This should­n’t re­ally be a sur­prise by now. Grok used by far the most to­kens (34M) but stayed cheap be­cause ~98% were cached reads at a frac­tion of the in­put rate, and Grok’s rates are the low­est of the four. Gemini also leaned heav­ily on cached reads, so even at 27.7M to­kens it stayed mod­est.

Did the mod­els ac­tu­ally im­prove?

Every view_­can­vas re-scored the can­vas against the tar­get, so we can watch progress over a run. Two pat­terns held across both tar­gets:

First, they plateau early. Claude Fable 5 re­viewed its Mona Lisa 27 times, but its sim­i­lar­ity was es­sen­tially flat af­ter about the fifth re­view. Second, in all eight tar­get runs, the fi­nal draw­ing scored be­low the best the model reached mid-run. GPT-5.6 Sol’s Mona Lisa peaked at 0.352 SSIM and ended at 0.325; its Starry Night peaked at 0.179 and ended at 0.130. Gemini 3.6 Flash re­viewed the most (about 23 times per draw­ing) and reached the high­est peak of the whole batch, 0.449 on the Mona Lisa, then slid all the way back to 0.337 by the end. More re­view­ing did not trans­late into a bet­ter fin­ish. The mod­els kept edit­ing past their own best per­for­mance.

Objective sim­i­lar­ity (target runs)

SSIM is 0 to 1, higher is closer to the tar­get. RMSE is 0 to 255, lower is closer to the tar­get.

On raw SSIM, Gemini 3.6 Flash scored high­est on both tar­gets, on fi­nal score and on peak, though it also swung the most be­tween its best frame and its fi­nal one. Gemini 3.6 Flash looked de­cent, but def­i­nitely does not look the clos­est to the tar­get out­put. Very in­ter­est­ing that the raw score turned out to be higher though. As al­ways, SSIM mea­sures struc­tural close­ness, not how good the draw­ing looks to a per­son (see our take be­low).

Method notes

Four mod­els, two scored tar­gets plus five prompts each, 28 draw­ings to­tal. All runs used the same col­ored-pen­cil toolset. Reasoning was set to high for every model.

SSIM/RMSE are com­puted by re­siz­ing both im­ages to 256x256 and com­par­ing. They mea­sure struc­tural/​pixel close­ness, not artis­tic qual­ity, and only ap­ply to the tar­get runs (prompts have no ref­er­ence to score against).

Wall-clock time in­cludes the mod­el’s own think­ing and tool round-trips.

Our take

GPT-5.6 Sol was the run­away leader. Its Starry Night and its rose were our fa­vorites of the whole batch, which is wild con­sid­er­ing it drew them stroke by stroke on a blank can­vas. It also beat the other two on de­tail in al­most every draw­ing, the one ex­cep­tion be­ing the el­derly fish­er­man.

Grok 4.5 was ba­si­cally garbage here. It rarely pro­duced any­thing us­able, and de­spite its fran­tic tool use I would not trust it to re­li­ably un­der­stand or judge vi­sual out­put yet.

Gemini 3.6 Flash was def­i­nitely bet­ter than Grok at a slight cost in­crease, but com­par­ing the Flash model to fron­tier ver­sions seems a lit­tle un­fair. I’m def­i­nitely ex­cited to see how Gemini 4 will per­form. We also al­ready knew this, but the to­ken out­put speed is just so im­pres­sive.

Claude Fable 5 landed sec­ond on qual­ity but at the far end on cost and time. For this task it was the least ap­peal­ing trade-off of the three: much slower and far more ex­pen­sive (about 20x the oth­ers) for out­put that did not keep up with Sol.

Just a moment...

en.help.roblox.com

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.