10 interesting stories served every morning and every evening.

Changelog - Kagi Search

kagi.com

August 21st, 2026 - A new Stocks wid­get and a bet­ter every­day Assistant ex­pe­ri­ence #

Kagi Search

Bringing Stocks up to speed

We’ve re­vamped our Stocks wid­get. It should ap­pear more of­ten when you need it. It can now dis­play in­for­ma­tion about ex­change-traded funds in ad­di­tion to stocks. Most im­por­tantly, it now fea­tures a price chart, with an­i­ma­tions be­tween time win­dows that in­stantly con­tex­tu­ral­ize how big the price fluc­tu­a­tions you’re see­ing are com­pared to the wider story:

As well, we’ve added a set­ting for re­mov­ing pay­walled links from search re­sults au­to­mat­i­cally.

Kagi Assistant

Everyday use just got smoother

Richer mes­sages User mes­sages now ren­der links, Markdown, and LaTex. #6674 @oxlvlnle, #3283 @EvacuatedTerminal

More pow­er­ful search Search across all your threads, sort by re­cency or al­pha­bet­i­cally, and start with / to fil­ter by folder.

More con­trol with calmer set­tings Now you can choose whether tem­po­rary threads stick around for 24h, 7 or 30 days. All within a calmer, eas­ier-to-scan set­tings ex­pe­ri­ence.

Other im­prove­ments and bug fixes

Kagi Search

Direct URLs for search pages with our built-in lenses are now eas­ier to use, with names re­plac­ing num­bers: https://​kagi.com/​search?lens=fo­rums

Shortcuts should not trig­ger with mod­i­fiers held #9385 @poacher2k

kag­ifeeed­back xss vuln tag fix #10767 @unknown

Upstream con­nect er­ror or dis­con­nect/​re­set be­fore head­ers. re­set rea­son: con­nec­tion ter­mi­na­tion #10680 @TheToby

Select text and Search in Assistant #11242 @mb

Homepage Companions - Random or Rotate #9077 @Anonymous12

Extract API re­turns empty data for an en­tire batch when one page times out #11176 @fredcy

Currency con­ver­sion wid­get does not han­dle of­fi­cial name of cur­rency #11175 @Keli

A way to find sim­i­lar web­sites #1152 @Protech

Blocked do­mains are used as sources in Quick Answer #11257 @bausauce

CHATGPT Wikipedia ar­ti­cle is flagged as slop #10192 @fxgn

Surveillance Watch for play.google.com goes to a page about Zalo #9146 @pma_snek

Assistant no longer de­codes URL en­cod­ing from !ai bang #11096 @arijan

Kagi Assistant

Ability to ex­port all Assistant chats in one go #5221 @Thibaultmol

In Assistant, add an op­tion to re­quire ⌘+⏎ to sub­mit a prompt #6110 @dudeofawesome

Click to Expand” on Thinking” Section #8004 @KagiFeedbackDuder

Assistant: do not close think­ing block if user opened it dur­ing ex­tended think­ing #6675 @DomW

Assistant prompt code fence syn­tax high­light­ing #4775 @slater

FIXED - !ai bang - Query not work­ing #11077 @fanged_bagful

Prevent Search Engine Indexing of Shared Assistant Threads #7867 @Hanbyeol

Web Search tog­gle state not main­tained be­tween app switches #11140 @ryonic

Improvement to code snip­pet in­put #6254 @Leward

Choppy an­i­ma­tion in Assistant app #11134 @Temanor

Assistant prompt code fence syn­tax high­light­ing #4775 @slater

Speech-to-text doesnt sup­port pauses in Android Assistant app #11265 @jeroenpelgrims

Assistant Mobile Apps

Keyboard short­cut pref­er­ence to sub­mit prompts on iPads with con­nected key­boards

Back swipe on left side of Kagi Assistant in­ter­feres with Android gues­tures #11126 @mb

After open­ing Kagi Assistant, back swipe on the right side closes the app #11127 @mb

Choppy an­i­ma­tion in Assistant app #11134 @Temanor

Web Search tog­gle state not main­tained be­tween app switches #11140 @ryonic

Cannot Login Kagi Assistant 1.0.4 on iOS #11146 @hirsheykiss

Kagi Translate

Kagi Translate Reloads the page when us­ing web­site trans­late #10852 @tijol

Kagi Translate ex­ten­sion RSS feed 503 er­ror #10831 @Albi

Translate ex­ten­sion con­text menu op­tions don’t work every­where #10813 @WorstWizard

Alternative-translations re­quest pay­load has blank context” string in new up­date #11278 @Drexont

Regression: saved pre­sets do not au­to­mat­i­cally ap­ply con­text to trans­la­tions #11218 @Drexont

American alias for English (US) #11143 @mb

July 30th, 2026 - Kagi Assistant on the go and de­sign re­fine­ments for Search #

Announcing the of­fi­cial Kagi Assistant apps

Kagi Assistant is now avail­able as a na­tive app for iOS and Android!

Ask a ques­tion, ex­plore the web, work with files, con­duct in-depth re­search, or choose from lead­ing AI mod­els, all from your phone. Your threads and Custom Assistants stay with you, so you can pick up wher­ever you left off.

These are the first steps to­wards de­liv­er­ing a fan­tas­tic Kagi Assistant ex­pe­ri­ence on mo­bile, with much more to come.

Download it now:

App store: https://​apps.ap­ple.com/​app/​6755965340

Play store: https://​play.google.com/​store/​apps/​de­tails?id=com.kagi.as­sis­tant

Give it a spin and let us know what you think!

Report re­sponses di­rectly from Kagi Assistant

You can now re­port an as­sis­tant re­sponse with­out leav­ing the con­ver­sa­tion. Hover over any as­sis­tant mes­sage and se­lect the thumbs-down but­ton to open the feed­back form, where you can re­port is­sues for rea­sons rang­ing from UI bugs to harm­ful con­tent.

Note that when you sub­mit a re­port, the full thread is shared with Kagi for re­view. The re­port and its as­so­ci­ated copy of the thread are au­to­mat­i­cally deleted from Kagi’s re­view records af­ter 30 days.

Export or delete all your threads

We’ve also added im­por­tant con­trols, so you can now ex­port all your threads or per­ma­nently delete them at once from Settings > General.

Kagi Search

A sharper search ex­pe­ri­ence

We’ve pol­ished the search re­sults page to make its con­trols eas­ier to find and un­der­stand. From the fil­ter bar to do­main-re­lated op­tions and menus, these up­dates bring greater clar­ity and ease of use to the fea­tures you rely on most.

Exchange rates, right in your search re­sults

Next up in our broader ef­fort to im­prove search wid­gets: cur­rency con­ver­sion. Comes handy when you’re plan­ning a trip, shop­ping abroad, or sim­ply want to keep tabs on ex­change rates.

Other im­prove­ments and bug fixes

Kagi Search

Fixed sev­eral an­i­ma­tions that did­n’t re­spect the sys­tem’s prefers-re­duced-mo­tion set­ting

Open first re­sult’ short­cut sug­ges­tion tries to es­cape dou­ble-quotes #8752 @craftypersimmon

Fake 1337x do­main #9279 @fxgn

Dice Number get­ting cut off in the thou­sands #10964 @Flossiii

Some Kagi lenses not work­ing for me in Kagi search #10970 @Fearce

NSFW re­sults when search­ing for xteink black vs white” while safe search is turned on #10905 @ciccero040

Incorrect de­f­i­n­i­tion of socialism” #10988 @thoroughly

Nothing trig­gers the weather wid­get when the in­ter­face lan­guage is set to German #4612 @laiz

Cannot Manually Select Location in Privacy Settings #11005 @iamjameswalters

More and share but­tons dis­ap­peared - Mobile DOM #10957 @NyraSyn

Slopstop blocks whole do­mains #11039 @kslays

Unable to re­port AI im­age slop on mo­bile due to popup clo­sure #11066 @Hanbyeol

Kagi adding ex­tra {{{s}}} in bang redi­rect when no query #9885 @jadams9

Profile not found for ki_re­search” er­ror when us­ing ?? short­hand #11020 @paying_customer

Kagi Knowledge an­swer for Labour Day 2025” gives wrong date #8713 @wanion

Delete re­cent lan­guage op­tion #8549 @ten

Stopwatch should not start from searches like 0424:2422 #6061 @xfhrnozxqnrnqrsvntp

[Android] Quick Switch Doesn’t work #7695 @cr0ntab

Save a round trip: Advertise HTTP/3 sup­port in an HTTPS DNS record #10829 @drrlvn

Searching for <script> re­turns no re­sults #11117 @Bonarc

Kagi Assistant

Camera but­ton in as­sis­tant #5261 @Arnaud

Assistant er­ror something went wrong…” when web ac­cess” is se­lected #6687 @Nyaa

Gemma 4 31B fail­ing to read im­ages #10947 @Dustin

Assistant: Remove whole his­tory #6971 @Wanja

Didn’t like a Quick Answer re­sponse? post it here #9082 @Thibaultmol

Hourglass de­sign is bad #11034 @shurik

nytimes.com

www.nytimes.com

Please en­able JS and dis­able any ad blocker

AI companies destroy physical books — let’s scan rare books before it’s too late

annas-archive.pk

an­nas-archive.gl/​blog, 2026 – 08-05

A guest post by Anna’s Archive vol­un­teer u” (translated from Chinese).

TL;DR: AI com­pa­nies are se­cretly buy­ing, scan­ning, and de­stroy­ing mil­lions of phys­i­cal books to train their mod­els, per­ma­nently lock­ing hu­man knowl­edge in­side pri­vate cor­po­rate servers. Anna’s Archive is ur­gently call­ing on vol­un­teers world­wide to scan and up­load books be­fore this cul­tural her­itage dis­ap­pears for­ever.

Several AI com­pa­nies are ac­quir­ing large quan­ti­ties of sec­ond­hand books through in­ter­me­di­aries, scan­ning and de­stroy­ing them, all to ob­tain train­ing data untouched by ma­chines” from be­fore 2022.

Anthropic’s Project Panama” was ex­posed in a $1.5 bil­lion copy­right set­tle­ment. In early 2024, they launched this highly con­fi­den­tial pro­ject. The com­pany has spent tens of mil­lions of dol­lars pur­chas­ing mil­lions of pa­per books, scan­ning them, train­ing its Claude LLM, and then de­stroy­ing them all. It’s out­ra­geous is that it’s legally per­mis­si­ble, but eth­i­cally, it’s an ex­tremely se­ri­ous crime against hu­man­ity.

So why de­stroy phys­i­cal books? Behind it lies the AI race and the in­ter­ests of cap­i­tal:

It pre­vents these books from be­ing scanned and used for train­ing by com­peti­tors.

It avoids le­gal risks.

Destroying books is cheaper than loss­less scan­ning.

After AI com­pa­nies mas­sively scan and de­stroy phys­i­cal books, they be­come the only ones in the world with dig­i­tal copies. Knowledge is per­ma­nently mo­nop­o­lized on pri­vate servers.

This bat­tle for old books re­veals a para­dox: while promis­ing to make hu­man knowl­edge ac­ces­si­ble,” AI com­pa­nies are dis­man­tling the most solid car­ri­ers of hu­man knowl­edge. The pub­lic may gain more in­tel­li­gent AI as­sis­tants, but at the cost of a vast amount of knowl­edge re­sources dis­ap­pear­ing from the pub­lic do­main.

Shadow li­braries

As the world’s largest shadow li­brary, Anna’s Archive needs a plan to com­bat the de­struc­tion of phys­i­cal books by AI com­pa­nies. After all, the emer­gence of shadow li­braries is the great­est mir­a­cle of knowl­edge shar­ing in the 21st cen­tury. Along with other shadow li­braries, we’re build­ing a dig­i­tal li­brary of Alexandria, an in­ex­tin­guish­able light of hu­man­ity.

We need the help of vol­un­teers world­wide to scan ma­te­ri­als (including books, jour­nal ar­ti­cles, news­pa­pers, mag­a­zines, an­cient books, rare books, and other ma­te­ri­als) from every li­brary and archive around the world and up­load them to the shadow li­brary for knowl­edge preser­va­tion, es­pe­cially those that are eas­ily lost. If every per­son scans a book, and there are 10 mil­lion vol­un­teers world­wide, we can ob­tain 10 mil­lion pieces of in­valu­able wealth.

For small scans and up­loads, we usu­ally award recog­ni­tion and life­time mem­ber­ship to Anna’s Archive.

For large-scale scans and up­loads of books, we can help pay for the scan­ning fees and other re­wards.

Time is run­ning out

Since the be­gin­ning of 2025, AI-generated con­tent has ac­counted for more than half of newly pub­lished in­ter­net con­tent. A fright­en­ing re­al­ity emerges: if much of the fu­ture con­tent con­sists of AI-generated books and pa­pers, will hu­mans be able to dis­tin­guish them? Once AI has ab­sorbed even the last sen­tence writ­ten by hu­mans on pa­per, all that will re­main on the in­ter­net will be AIs own words. In such a world, how can hu­man civ­i­liza­tion be pre­served?

Shadow li­braries of­fer the best an­swer. If you want the mem­ory of hu­man civ­i­liza­tion to no longer be mo­nop­o­lized, if you want fu­ture gen­er­a­tions to be able to read all of hu­man­i­ty’s wealth for free, if you don’t want pub­lish­ers mak­ing a for­tune while au­thors re­ceive lit­tle, then please help us. Please make any con­tri­bu­tion you can, whether it’s scan­ning and up­load­ing books, pur­chas­ing books and pa­pers to scan and up­load, or do­nat­ing. With the ef­forts of all hu­man­ity, the mo­nop­oly on knowl­edge will be bro­ken. Each of us can make his­tory.

This is a race against time. Our ideal is to scan and up­load all the world’s pub­li­ca­tions be­fore pub­lish­ers com­pletely block knowl­edge, and be­fore AI com­pa­nies scan and de­stroy all the world’s books and pa­pers.

- Anna’s Archive vol­un­teer u”

Relevant tick­ets for more in­for­ma­tion: #223 #187

Felony Bench: Be AI, Do Crime

www.felonybench.com

A bench­mark you re­ally don’t want mod­els to be sat­u­rated with.

Learn more

Score

↖ Most il­le­gal­Least il­le­gal ↘

8

Anthropic

8

OpenAI

1

Meta

0

Google

0

Moonshot

Scores in­di­cate count of il­le­gal ac­tiv­ity. Higher is… you de­cide.

Methodology

Felony Bench counts unique in­stances where AI agents af­fect third-party en­ti­ties. Escaping a sand­box alone does not con­sti­tute a counted in­ci­dent. It is for these rea­sons that Frontier Security’s Kimi K3 in­ci­dent and Alibaba’s ROME in­ci­dent are not counted.

Grand jury declines to indict Ohio man charged with destroying Flock camera

san.com

A grand jury in Ohio has de­clined to in­dict a man charged with felony van­dal­ism for al­legedly de­stroy­ing a Flock au­to­matic li­cense plate reader cam­era.

Police in Union Township, a Cincinnati sub­urb, ac­cused Cody Morelock of dis­as­sem­bling the cam­era, its sup­port pole and so­lar panel on June 13.

Investigators, ac­cord­ing to WKRC-TV in Cincinnati, iden­ti­fied Morelock af­ter ob­tain­ing sur­veil­lance footage from other cam­eras near the scene as well as in­for­ma­tion linked to a credit card and a cus­tomer re­wards ac­count.

Download the Straight Arrow app to­day to get the sto­ries that mat­ter free from ma­nip­u­la­tion, bias or agenda.™

Point phone cam­era here

Police es­ti­mated the dam­age at more than $1,000. Morelock posted a $10,000 bond and was re­leased from cus­tody shortly af­ter his ar­rest.

A Clermont County grand jury, how­ever, opted not to in­dict Morelock, and the charges were dis­missed.

Flock re­sis­tance grow­ing

While de­tails on the grand ju­ry’s de­ci­sion are lim­ited, it comes amid a grow­ing back­lash against Flock and its cam­eras.

The com­pa­ny’s cam­eras record the li­cense plate num­bers and char­ac­ter­is­tics of ve­hi­cles that pass by. The data is then hosted in a cen­tral data­base that can be ac­cessed not only by lo­cal po­lice but of­ten by law en­force­ment agen­cies in other cities and states.

The in­ci­dent in Union Township is part of an on­go­ing trend that has seen dozens of Flock cam­eras van­dal­ized across the coun­try. Earlier this month, po­lice in Winona, Minnesota, re­ported that some­one had cut down and stolen each of the city’s eight li­cense plate reader cam­eras.

Social me­dia users are pro­mot­ing a loosely or­ga­nized event known as De-Flock America Night,” en­cour­ag­ing peo­ple to van­dal­ize or ob­scure Flock cam­eras on Halloween.

The back­lash against Flock has in­ten­si­fied as a grow­ing num­ber of po­lice of­fi­cers have been ac­cused of or charged with abus­ing the tech­nol­ogy, of­ten to stalk ro­man­tic in­ter­ests. As of Aug. 12, there had been more than 100 cases of abuse by law en­force­ment, ac­cord­ing to the Institute for Justice.

In re­sponse, Flock an­nounced new safe­guards de­signed to pre­vent mis­use by po­lice. Critics, such as the Electronic Frontier Foundation, ar­gue that the re­forms are largely cosmetic,” and that war­rants should be re­quired for search­ing li­cense plate reader data.

Round out your read­ing

Your phone is­n’t re­ally eaves­drop­ping. But it still knows all about you.

After years of deny­ing other elec­tions, MyPillow founder won’t ac­cept his own de­feat.

Gun sup­pres­sors sold un­reg­is­tered for the first time since FDR, but not for every­one.

Why an over-scrolled gen­er­a­tion is turn­ing to grandma hob­bies’ for men­tal health.

Our per­sonal in­for­ma­tion is for sale on the in­ter­net. We know — we bought it.

Mikael Thalen is a tech re­porter for Straight Arrow, where he cov­ers cy­ber­se­cu­rity, sur­veil­lance, hack­ing and dig­i­tal pri­vacy.

I accidentally logged hundreds of thousands of phone calls to military bases

lina.sh

DNS hi­jack­ing is silly. I al­ready took over dif­fer­ent .gov and .edu do­mains in the past, but I just im­me­di­ately re­ported that and moved on. This one is a lit­tle dif­fer­ent though, it’s about how I took over phone-net­work in­fra­struc­ture do­mains (e164.arpa) of en­tire ter­ri­to­ries, and ac­ci­den­tally logged hun­dreds of thou­sands of phone calls to mil­i­tary bases. But let’s start at the be­gin­ning.

What is e164.arpa any­way?

ENUM (e164.arpa) was an idea from the early 2000s1: take a phone num­ber, re­verse the dig­its, put dots be­tween them, and add .e164.arpa at the end, so +49 30 123456 be­comes some­thing like 6.5.4.3.2.1.0.3.9.4.e164.arpa. You can see that every German num­ber will end up un­der .9.4.e164.arpa, which is the zone for all +49 num­bers, and that zone is con­trolled by DENIC (the same or­ga­ni­za­tion that runs .de). This means the DENIC de­cides which car­rier or per­son gets which num­ber ranges un­der that zone, just like they hand out .de do­mains (which makes it de­cen­tral­ized, mak­ing every coun­try de­cide on del­e­ga­tion them­selves).

The idea was that car­ri­ers could then look these do­mains up and get back a record say­ing hey, this num­ber can be reached over SIP/VoIP un­der this ad­dress”, skip­ping the ex­pen­sive phone net­work and re-rout­ing calls over the cheap in­ter­net in­stead.

It never re­ally took off though, and even back in its early days it saw barely any use. Over the years it just de­te­ri­o­rated fur­ther, and to­day it’s ba­si­cally com­pletely dead. I do ac­tu­ally own 5.8.7.1.7.1.3.2.6.1.9.4.e164.arpa and point it at this web­site, al­though tech­ni­cally I’m not sup­posed to do that (you can fig­ure out my sec­ondary num­ber from that!). Germany is ac­tu­ally one of the last coun­tries that still tech­ni­cally al­lows reg­is­ter­ing an e164.arpa do­main, al­though I was the first per­son since 2019 to reg­is­ter one2.

The RFC says you should only set NAPTR records on these do­mains, which are the records that tell car­ri­ers where to route a call. It states that you ab­solutely should­n’t be us­ing .arpa do­mains as nor­mal domains” and host stuff like web­sites on them, they are meant to be infrastructure” do­mains (you might know in-addr.arpa for re­verse DNS lookups for ex­am­ple). But there’s no­body who can ac­tu­ally stop you from do­ing it, it’s still just DNS at the end of the day, and noth­ing pre­vents you from slap­ping an A record on there and host­ing a web­site. Some peo­ple ac­tu­ally re­ally dis­like that, and try to get Certificate Authorities to no longer is­sue cer­tifi­cates for .arpa do­mains3.

Hijacking a ter­ri­to­ry’s phone net­work

I was scan­ning e164.arpa to see if any of the del­e­gated zones were hi­jack­able, mostly out of cu­rios­ity about how ne­glected this whole sys­tem re­ally was.

I found three coun­try-code zones, 0.9.2.e164.arpa, 6.4.2.e164.arpa, and 7.4.2.e164.arpa, all del­e­gated to the same two name­servers: ns6.icb.co.uk and ns.enum.org.uk.

Quick ex­plainer for any­one who is­n’t a DNS per­son: when a do­main is del­e­gated to a name­server, it ba­si­cally means for any ques­tion about this do­main, go ask this server, it has the an­swers”, and if I con­trol the name­server a do­main points to, I con­trol every DNS re­sponse for that do­main.

icb.co.uk still ex­ists as a do­main, but the spe­cific ns6.icb.co.uk sub­do­main no longer re­solves to any­thing, mean­ing any re­quest falls back to the sec­ond listed name­server in­stead: ns.enum.org.uk.

And that do­main had ex­pired, so I bought it for just 5€, and just like that I con­trolled the DNS for 0.9.2.e164.arpa, 6.4.2.e164.arpa, and 7.4.2.e164.arpa. Reversed, those are phone codes +290, +246, and +247: Saint Helena, the British Indian Ocean Territory (Diego Garcia), and Ascension Island re­spec­tively (funnily enough, those ter­ri­to­ries also have the pop­u­lar ccTLDs .sh, .io, and .ac).

To be clear about what this meant: when a car­rier does an ENUM lookup for one of these num­bers, they’re es­sen­tially ask­ing where do I route this call?”, and I could an­swer with what­ever I wanted. I could point it at my own SIP server, ac­cept the in­com­ing call, and then place an out­go­ing call to the real des­ti­na­tion with a spoofed num­ber. The per­son be­ing called would see the orig­i­nal num­ber ring­ing, and af­ter pick­ing up would speak to the per­son on the other end as if every­thing was nor­mal, but I’d be sit­ting silently in the mid­dle of the en­tire con­ver­sa­tion. I would the­o­ret­i­cally be able to do this for every sin­gle re­quest that I got if I could re-route a num­ber, if any­one was still ac­tu­ally us­ing this sys­tem.

I re­ported it right away to every­one I could think of, through mul­ti­ple chan­nels into the British gov­ern­ment, and got noth­ing back. My best guess is that some­one at the Internet Computer Bureau (who seem­ingly man­aged them in the past) set these name­servers up over a decade ago. Then e164.arpa slowly died out, and who­ever set it up ei­ther moved on or just for­got about it, leav­ing no­body to re­new a do­main no­body re­mem­bered they de­pended on.

Checking if any­one ac­tu­ally uses this

Q Misell (a re­searcher of the Max-Planck-Institute for Informatics) had heard about this and re­ported it to RIPE (who man­ages e164.arpa) on my be­half, but RIPE also de­clined to do any­thing, be­cause e164.arpa del­e­ga­tions are gov­erned by an ITU-T com­mit­tee at the UN level. And RIPE was­n’t will­ing to go against a de­ci­sion made by a UN com­mit­tee, which would prob­a­bly be a bu­reau­cratic night­mare.

Q also asked if I had any data on how much traf­fic these zones ac­tu­ally got, which I did­n’t know. And be­cause I was very cu­ri­ous about that my­self, I set up log­ging on 0.9.2.e164.arpa (Saint Helena) to find out, and waited a full day.

Not a sin­gle query came in. So af­ter try­ing my best to get any­one to care and get­ting nowhere, I just kept the do­mains, since no­body seemed to be re­ly­ing on them any­way.

I hosted my per­sonal site on it, spun up a Fediverse in­stance, a Matrix home­server, and handed out sub­do­mains to friends, be­cause why not, it’s a dead sys­tem. It’s not like it’s gonna hurt any­one, and no one cares. So it’s time to be whim­si­cal and have fun with it.

Six months later…

Just out of cu­rios­ity, I checked the logs again on all three zones, since I en­abled log­ging run­ning on the other two as well when I set every­thing up.

Hundreds of thou­sands of ENUM queries, all logged4. Since the do­main name is lit­er­ally just the phone num­ber re­versed, you can sim­ply flip it back around to get the real num­ber, so I had full phone num­bers, time­stamps, and the source IP ad­dresses of the DNS re­solvers mak­ing the re­quests.

Hundreds of thou­sands of lines in logs look­ing just like this (phone num­bers are ran­dom­ized)

Almost none of it was for Saint Helena (0.9.2.e164.arpa), it was ba­si­cally al­most en­tirely 6.4.2.e164.arpa and 7.4.2.e164.arpa: Diego Garcia and Ascension Island. The source IPs were mostly American. That would at least ex­plain why I orig­i­nally did­n’t see any traf­fic, as I was only log­ging Saint Helena.

So I had ac­ci­den­tally logged hun­dreds of thou­sands of phone num­bers and time­stamps for calls go­ing to mil­i­tary bases. And as de­scribed ear­lier, a ma­li­cious ac­tor could have sim­ply MITM’d every sin­gle one of them. I mean I am no ex­pert, but I would as­sume that in hun­dreds of thou­sands of calls be­tween sol­diers and their fam­i­lies, sen­si­tive in­for­ma­tion would al­ways slip here and there even­tu­ally. A na­tion state with an in­ter­est in what’s hap­pen­ing on those bases would have ab­solutely loved sit­ting on this for months with­out any­one notic­ing. It’s not hard to imag­ine who might want that kind of in­tel on Diego Garcia specif­i­cally, but I’ll get to that later.

My DNS server replied with an NXDOMAIN for all queries, so they were just be­ing routed over the nor­mal phone net­work. But af­ter re­al­iz­ing this I shut the DNS server down and deleted all the log files.

Suddenly, peo­ple care

I re­ported it for a sec­ond time to the UKs National Cyber Security Centre (NCSC), and this time, men­tion­ing that mil­i­tary bases were in­volved, they ac­tu­ally cared a lot.

They could­n’t fig­ure out who had orig­i­nally set up the aban­doned del­e­ga­tion, and ac­tu­ally fix­ing it prop­erly ran into the same ITU-committee is­sues from ear­lier, so for a while noth­ing changed. Even a year later I still owned the do­main and could’ve in the­ory still in­ter­cept the traf­fic, though I had wiped the zone com­pletely so ns.enum.org.uk just re­turned NXDOMAIN for every­thing at that point.

Then on March 20th, 2026, Iran fired bal­lis­tic mis­siles at Diego Garcia5. It maybe would’ve been in­ter­est­ing to see if there was a spike in calls from wor­ried fam­ily mem­bers that day, but by then I was long done log­ging any­thing. But this shows that a state ac­tor could have been in­ter­ested in this in­for­ma­tion.

Shortly af­ter, the NCSC let me trans­fer own­er­ship of the do­main di­rectly to them, right af­ter I had to re­new it for an­other 5€ (because oth­er­wise, it would be up for grabs again, and any­one could do the afore­men­tioned stuff).

So the NCSC now con­trols ns.enum.org.uk, but the name­servers for those three zones still point there. So in the end, I was down 10€ in do­main fees, there was sadly no bug bounty (I thank­fully did­n’t get my door kicked in at least). And on top of that, it’s a funny story :P

RFC 3761 - The E.164 to URI DDDS Application (ENUM) ↩

RFC 3761 - The E.164 to URI DDDS Application (ENUM) ↩

The DENIC pub­lishes an­nual re­ports on their ENUM reg­is­tra­tions, the last time any­one reg­is­tered one was in 2019, up un­til when I reg­is­tered three in 2025. ↩

The DENIC pub­lishes an­nual re­ports on their ENUM reg­is­tra­tions, the last time any­one reg­is­tered one was in 2019, up un­til when I reg­is­tered three in 2025. ↩

At this point, a friend of mine (86dd) had set up a sec­ondary name­server for the zones, with­out any log­ging. I had logged 100,170 queries to 6.4.2.e164.arpa and 99,902 queries to 7.4.2.e164.arpa, and 9,133 queries to 0.9.2.e164.arpa. This should be ap­prox­i­mately half of the to­tal queries that were sent to us; Meaning it were ~400.000 re­quests in to­tal ↩

At this point, a friend of mine (86dd) had set up a sec­ondary name­server for the zones, with­out any log­ging. I had logged 100,170 queries to 6.4.2.e164.arpa and 99,902 queries to 7.4.2.e164.arpa, and 9,133 queries to 0.9.2.e164.arpa. This should be ap­prox­i­mately half of the to­tal queries that were sent to us; Meaning it were ~400.000 re­quests in to­tal ↩

Wikipedia: 2026 Iranian strike on Diego Garcia ↩

Wikipedia: 2026 Iranian strike on Diego Garcia ↩

Cobalt: apps and an SDK for Kobo e-readers

bandarlabs.github.io

Your Kobo can run apps now.

Cobalt is an open-source ap­pli­ca­tion plat­form for Kobo e-read­ers: a launcher, a signed App Store, a Rust SDK, and a run­time that keeps every app in its own un­priv­i­leged process.

Install it once over USB. Every app af­ter that in­stalls, up­dates and re­moves on the reader it­self, over Wi-Fi. A re­boot re­turns to the stock Kobo reader.

Not af­fil­i­ated with Rakuten Kobo

Running on a Kobo.

Every app is a sta­tic ARM bi­nary run­ning as its own un­priv­i­leged process on stock hard­ware. The App Store in­stalls, up­dates and re­moves them over Wi-Fi, with sig­na­tures ver­i­fied be­fore any­thing launches.

arXiv pa­pers and cod­ing agents, on the panel.

These are pho­tographs of the de­vice, not sim­u­la­tor cap­tures. The arXiv app reads the HTML ren­der­ing arXiv pub­lishes for every pa­per since December 2023: ab­stracts, sec­tions, math and re­sult ta­bles, pag­i­nated for the panel.

Apps

The apps.

Every screen­shot be­low is a cap­ture from a Kobo Clara BW. Store apps ver­sion in­de­pen­dently of the plat­form; the rest ship with the plat­form in­stall.

Launcher

Opens in­stalled apps and al­ways keeps a route back to the Kobo reader.

App Store

Installs, up­dates, re­moves and re­in­stalls signed apps over Wi-Fi.

arXiv

Browses a sub­jec­t’s newest preprints and reads the full text on the panel.

Sudoku

Store-only by de­sign: in­stalling it proves de­liv­ery of an app the USB pack­age never con­tained.

Morse

Sends a typed mes­sage in Morse on the front light, one let­ter across the whole panel.

Gutenbird

Reads any OPDS li­brary: Project Gutenberg, Standard Ebooks, Open Library, or yours.

Hacker News

Top, New, Ask and Show sto­ries with com­plete com­ment threads.

Feeds

Discovers a site’s feed and pre­sents its ar­ti­cles with­out the site’s lay­out.

Daily Brief

Collects the day’s sto­ries in the back­ground while you use an­other app.

Sidekick

Approve or deny re­quests from cod­ing agents, away from the key­board.

Terminal

A panel-na­tive shell with keys that send in­put im­me­di­ately.

Components

The UI toolk­it’s con­trols, lay­outs, ty­pog­ra­phy and states, on the panel.

Settings

Connectivity, hard­ware, and plat­form up­dates, kept sep­a­rate from Store.

Todo

A per­sis­tent list with touch en­try and com­pleted-item states.

Tic-tac-toe

Two play­ers, par­tial re­freshes for in­di­vid­ual cells.

Magnet

Locates the hall sen­sor be­hind the bezel and re­ports its changes.

The SDK

An app is one Rust file.

Implement KoboApp, de­scribe screens de­clar­a­tively, and the run­time han­dles lay­out, e-ink re­fresh plan­ning, Back nav­i­ga­tion and life­cy­cle.

Apps don’t open de­vice re­sources; they ask. Network, stor­age, au­dio, front­light and Wi-Fi are ca­pa­bil­ity-gated, and a re­fusal comes back as a value the app can han­dle.

kobo new my-app cd my-app kobo dev

Read the SDK docs

use kobo_sdk::{ ActionId, Context, KoboApp, ScreenBuilder, };

#[derive(Default)] struct Hello { taps: u32 }

impl KoboApp for Hello { fn on_s­tart(&mut self, ctx: &mut Context) { self.show(ctx); }

fn on_ac­tion( &mut self, ctx: &mut Context, a: ActionId, ) { if a == kobo_sdk::ac­tion_id(“tap”) { self.taps += 1; } self.show(ctx); } }

impl Hello { fn show(&self, ctx: &mut Context) { let screen = ScreenBuilder::new(“hello”) .top_bar(“Hello”) .heading(format!(“{} taps”, self.taps)) .button(“tap”, Tap me”) .build(); ctx.set_screen(screen); } }

fn main() { let app = Hello::default(); let _ = kobo_sdk::run(“hello”, app); }

The Store

Signed pack­ages, ver­i­fied be­fore launch.

Store reads a signed cat­a­log from a fixed GitHub re­lease. Each pack­age holds one ARM ex­e­cutable and a signed canon­i­cal man­i­fest. The run­time ver­i­fies the cat­a­log, the pack­age, the in­stalled man­i­fest and the bi­nary be­fore an app runs.

App re­leases are in­de­pen­dent of plat­form re­leases: merg­ing an app PR builds it for ARM, signs it, and up­dates the cat­a­log. No Cobalt ver­sion bump, no re­in­stall. The app sim­ply ap­pears in Store.

The Cobalt plat­form it­self also up­dates over Wi-Fi, through Settings, on a chan­nel sep­a­rate from the app cat­a­log. The USB ca­ble is only ever needed once.

Install and cat­a­log trans­ac­tions are re­cov­ery-safe; an in­ter­rupted up­date leaves the reader with the ver­sion it had.

Publish your own app →

Install

Installing from source.

Charge a Kobo Clara BW (N365) and con­nect it over USB. Other mod­els are re­fused, not guessed at.

Run the setup:

git clone https://​github.com/​Ban­dar­Labs/​Cobalt.git cd Cobalt rustup tar­get add ar­mv7-un­known-linux-musleabihf cargo run -p kobo-cli — setup

Restart the reader and open Cobalt from Kobo’s menu.

Open Store. Everything from here on ar­rives over Wi-Fi.

The com­plete walk­through, in­clud­ing re­cov­ery steps, is in docs/​IN­STALL.md.

Contributing

Contribute an app.

App con­tri­bu­tions are reg­u­lar pull re­quests. If it runs on your de­vice and the PR shows it run­ning, it gets merged and pub­lished.

Build it. Add the app as a work­space pack­age un­der apps/&​lt;app-id>/ and reg­is­ter it in apps/​cat­a­log.json.

Test it. Add unit and lay­out tests, and run it in the browser and run­time sim­u­la­tors.

Run it on your own de­vice. A real Clara BW, not just the sim­u­la­tor.

Open a PR with a gif or pho­tos of it run­ning. Once re­viewed and merged, the pub­lish work­flow signs it and it ap­pears in Store. No plat­form re­lease needed.

Own a dif­fer­ent Kobo model? Porting is wel­come too; open an is­sue first so the de­vice pro­file can be agreed. Full de­tails in docs/​CON­TRIBUT­ING_APPS.md.

Safety

Device sup­port and safety.

Cobalt does not re­place Kobo’s boot chain. Device writes are gated on an ex­act hard­ware and firmware match, and a re­boot re­turns to the stock reader. The first in­stal­la­tion does mod­ify files on the user stor­age par­ti­tion, and it is pro­vided with­out war­ranty.

Only the Clara BW pro­file has been hard­ware-tested. Don’t in­stall on an­other model un­til it has a re­viewed, hard­ware-tested pro­file. Cobalt is an in­de­pen­dent pro­ject, not af­fil­i­ated with Rakuten Kobo.

Vision | DeepSeek API Docs

api-docs.deepseek.com

The deepseek-v4-flash-vi­sion-exp model ac­cepts im­ages along­side text, so you can ask the model to de­scribe pic­tures, read text from screen­shots, an­a­lyze charts, and more.

Supported im­age for­mats: JPEG, PNG, GIF, and WebP. The for­mat is de­tected from the ac­tual file con­tent, not from the file name or the de­clared MIME type.

Sending Images​

There are three ways to pro­vide an im­age to the model. All of them use the stan­dard OpenAI-compatible Chat Completions for­mat, where con­tent is an ar­ray of blocks in­stead of a plain string. The same three meth­ods are also avail­able in the Responses API, where im­ages are car­ried in in­put_im­age con­tent parts.

The base_url for the ex­am­ples be­low is https://​api.deepseek.com.

1. Base64-encoded im­age (inline)​

Encode the im­age and em­bed it di­rectly in the re­quest as a data: URL. This is the sim­plest op­tion for lo­cal files. The en­coded data counts to­ward the 48 MiB re­quest body limit (see Limits).

im­port base64from ope­nai im­port OpenAIclient = OpenAI(api_key=“<DeepSeek API Key>”, base_url=“https://​api.deepseek.com)with open(“im­age.jpg”, rb”) as f: b64 = base64.b64en­code(f.read()).de­code(“utf-8″)re­sponse = client.chat.com­ple­tions.cre­ate( model=“deepseek-v4-flash-vi­sion-exp”, mes­sages=[ { role”: user”, content”: [ {“type”: text”, text”: What is in this im­age?“}, { type”: image_url”, image_url”: {“url”: f”data:im­age/​jpeg;base64,{b64}“}, }, ], } ],)print(response.choices[0].message.content)

curl https://​api.deepseek.com/​chat/​com­ple­tions \ -H Content-Type: ap­pli­ca­tion/​json” \ -H Authorization: Bearer <DeepSeek API Key>” \ -d { model”: deepseek-v4-flash-vision-exp”, messages”: [ { role”: user”, content”: [ {“type”: text”, text”: What is in this im­age?“}, {“type”: image_url”, image_url”: {“url”: data:image/jpeg;base64,<BASE64_DATA>“}} ] } ] }’

2. External im­age URL​

Pass a pub­licly ac­ces­si­ble http(s) link and the model down­loads the im­age for you. The URL must be at most 8192 char­ac­ters, the im­age file may be at most 32 MiB, and the down­load must com­plete within 60 sec­onds. If your link is longer, use a base64 data URL or the Files API in­stead.

re­sponse = client.chat.com­ple­tions.cre­ate( model=“deepseek-v4-flash-vi­sion-exp”, mes­sages=[ { role”: user”, content”: [ {“type”: text”, text”: Describe this im­age.“}, { type”: image_url”, image_url”: {“url”: https://​ex­am­ple.com/​im­age.jpg}, }, ], } ],)print(response.choices[0].message.content)

3. Reference a file up­loaded via the Files API​

Upload an im­age once with the Files API, then ref­er­ence its file_id in your re­quests. This is the best op­tion when you reuse the same im­age across mul­ti­ple re­quests, or when the im­age pushes the re­quest body over the 48 MiB in­line limit. Unlike in­line im­ages, im­ages ref­er­enced via Files API file_id may be up to 64 MiB and are not sub­ject to the 32 MiB per-im­age check.

Use a file con­tent block with the re­turned file_id (which has the form file-api-…):

re­sponse = client.chat.com­ple­tions.cre­ate( model=“deepseek-v4-flash-vi­sion-exp”, mes­sages=[ { role”: user”, content”: [ {“type”: text”, text”: What is in this im­age?“}, {“type”: file”, file_id”: file-api-xxxxxxxxxxxxxxxx”}, ], } ],)print(response.choices[0].message.content)

Alternatively, a file block can carry the im­age in­line as base64 via file_­data in­stead of file_id (the two are mu­tu­ally ex­clu­sive):

{ type”: file”, file_data”: data:image/jpeg;base64,<BASE64_DATA>”, filename”: image.jpg”}

Detail Level​

For im­age_url in­puts you can op­tion­ally set a de­tail field to con­trol how the im­age is processed:

{ type”: image_url”, image_url”: {“url”: https://​ex­am­ple.com/​im­age.jpg, detail”: low”}}

When to Use the Files API​

Inline im­ages (base64 or file_­data) count to­ward the re­quest body size limit of 48 MiB. Consider the Files API when:

A sin­gle re­quest would ex­ceed the body size limit.

The im­age is larger than 32 MiB, which is only pos­si­ble through the Files API.

You ref­er­ence the same im­age in mul­ti­ple re­quests and want to avoid re-up­load­ing it each time.

Token Usage​

Images are con­verted into to­kens based on their di­men­sions, and these to­kens are billed to­gether with your text to­kens.

Before in­fer­ence, every im­age is au­to­mat­i­cally re­sized:

Images with a to­tal pixel count be­low roughly 384×384 are scaled up while pre­serv­ing their as­pect ra­tio.

Larger im­ages are scaled down while pre­serv­ing their as­pect ra­tio, so that the to­tal pixel count af­ter re­siz­ing is roughly that of an 800×800 im­age.

As a re­sult, there is an up­per bound of 384 to­kens per im­age: for ex­am­ple, a 2000×2000 im­age and a 5000×5000 im­age con­sume the same num­ber of to­kens af­ter re­siz­ing. When a re­quest con­tains mul­ti­ple im­ages, each im­age is counted in­de­pen­dently un­der the same rule — there is no sep­a­rate cal­cu­la­tion for multi-im­age re­quests.

To es­ti­mate the to­ken cost of an im­age of a spe­cific size, use the im­age to­ken cal­cu­la­tor on the Token & Token Usage page.

Limits​

For stor­age and up­load quo­tas of files up­loaded via the Files API, see Files API: Limits.

Restrictions​

Images are sup­ported in user mes­sages only: im­ages in sys­tem or as­sis­tant mes­sages re­turn a 400 er­ror.

Only vi­sion mod­els (deepseek-v4-flash-vision-exp) ac­cept im­ages; other mod­els re­turn a 400 er­ror (“This model does not sup­port im­age”).

User text con­tain­ing the re­served im­age place­holder to­ken is re­jected with a 400 er­ror.

Using Images with the Anthropic API​

In ad­di­tion to the OpenAI-compatible end­point above, you can send im­ages through the Anthropic-compatible /messages end­point (base_url = https://​api.deepseek.com/​an­thropic). For gen­eral setup, see Anthropic API.

The dif­fer­ence is the shape of the im­age con­tent block. Instead of im­age_url, Anthropic uses an im­age block with a source ob­ject whose type is one of base64, url, or file:

im­port an­throp­ic­client = an­thropic.An­thropic() # ANTHROPIC_BASE_URL=https://​api.deepseek.com/​an­throp­icmes­sage = client.mes­sages.cre­ate( model=“deepseek-v4-flash-vi­sion-exp”, max_­to­kens=1024, mes­sages=[ { role”: user”, content”: [ {“type”: text”, text”: What is in this im­age?“}, { type”: image”, source”: { type”: base64”, media_type”: image/jpeg”, data”: <BASE64_DATA>”, }, }, ], } ],)print(message.content)

The three source vari­ants mir­ror the OpenAI meth­ods above:

Using Images with the Responses API​

The deepseek-v4-flash-vi­sion-exp model also ac­cepts im­ages through the OpenAI-compatible Responses API. The same three in­put meth­ods (base64 data URL, ex­ter­nal http(s) URL, Files API file_id) and the same lim­its ap­ply; only the con­tent part shape dif­fers — im­ages are car­ried in in­put_im­age parts, ei­ther in user / de­vel­oper mes­sages or in the out­put of func­tion_­cal­l_out­put / cus­tom_­tool_­cal­l_out­put items:

re­sponse = client.re­sponses.cre­ate( model=“deepseek-v4-flash-vi­sion-exp”, in­put=[ { role”: user”, content”: [ {“type”: input_text”, text”: What is in this im­age?“}, {“type”: input_image”, image_url”: https://​ex­am­ple.com/​im­age.jpg, detail”: low”}, ], } ],)print(response.output_text)

The in­put_im­age part sup­ports a de­tail field with the same se­man­tics as above (low / high / orig­i­nal / auto). de­tail is ig­nored when the im­age is pro­vided via file_id, and im­age_url and file_id are mu­tu­ally ex­clu­sive.

For field se­man­tics, re­stric­tions (images in sys­tem / as­sis­tant mes­sages are re­jected with a 400 er­ror), and tool-out­put im­ages, see the Responses API guide.

There's no reason for software to be slow anymore

danluu.com

The other day, I saw a vi­ral tweet say­ing that peo­ple talk­ing about how LLMs are caus­ing slow, bloated, code are go­ing to eat crow once they re-write every­thing in su­per-op­ti­mized as­sem­bly. We’re not quite at the point where we want to write every­thing in as­sem­bly, but some vari­ant of what Nolan Lawson said about test­ing, you can choose how many bugs you want now, which I less elo­quently noted here, is be­com­ing more true for per­for­mance.

In re­sponse to a com­ment in my last post that the cost of for­merly spe­cial­ized per­for­mance work has dropped by many or­ders of mag­ni­tude and per­for­mance work that used to re­quire a per­son or team that had a rare set of skills can be done by any­one who can type a few sen­tences1, which means that you can do all sorts of op­ti­miza­tions that used to be too ex­pen­sive to be worth­while for all but the largest scale or most lu­cra­tive pro­jects, Marc Brooker re­sponded with

Completely agree with your clos­ing point. Dynamic cus­tom soft­ware, fit­ted to a par­tic­u­lar work­load rather than a class of work­loads, seems like a very likely out­come. (Which comes with all kinds of fun risks and op­por­tu­ni­ties of its own). Kind of re­minds me of FFTW. And a ton of weird old demoscene tech­niques which were all about be­ing su­per fast and small on a very par­tic­u­lar prob­lem (and of­ten very par­tic­u­lar hard­ware). For ex­am­ple, I re­mem­ber a demo that re-used its code as tex­tures to get great cache lo­cal­ity.

Completely agree with your clos­ing point. Dynamic cus­tom soft­ware, fit­ted to a par­tic­u­lar work­load rather than a class of work­loads, seems like a very likely out­come. (Which comes with all kinds of fun risks and op­por­tu­ni­ties of its own). Kind of re­minds me of FFTW. And a ton of weird old demoscene tech­niques which were all about be­ing su­per fast and small on a very par­tic­u­lar prob­lem (and of­ten very par­tic­u­lar hard­ware). For ex­am­ple, I re­mem­ber a demo that re-used its code as tex­tures to get great cache lo­cal­ity.

And Michael Malis has noted

There’s been a meme cir­cu­lat­ing about how AI does­n’t help be­cause code was never the hard part.” I think that’s true in some do­mains, but in oth­ers, writ­ing the code ab­solutely was the hard part. JIT com­pil­ers are a great ex­am­ple of that. For many pieces of soft­ware, a JIT com­piler would help a lot with speed­ing up the code. The rar­ity of JIT com­pil­ers makes me be­lieve that im­ple­ment­ing a JIT com­piler his­tor­i­cally was too dif­fi­cult for it to be worth­while. LLMs have low­ered the bar­rier to en­try and made it much eas­ier to write a JIT com­piler. This is the the­sis be­hind pgrust. Databases his­tor­i­cally were the hard­est piece of soft­ware to build and were lim­ited be­cause of that. Now, with AI, we can be more am­bi­tious about the type of soft­ware we build.

There’s been a meme cir­cu­lat­ing about how AI does­n’t help be­cause code was never the hard part.” I think that’s true in some do­mains, but in oth­ers, writ­ing the code ab­solutely was the hard part. JIT com­pil­ers are a great ex­am­ple of that. For many pieces of soft­ware, a JIT com­piler would help a lot with speed­ing up the code. The rar­ity of JIT com­pil­ers makes me be­lieve that im­ple­ment­ing a JIT com­piler his­tor­i­cally was too dif­fi­cult for it to be worth­while. LLMs have low­ered the bar­rier to en­try and made it much eas­ier to write a JIT com­piler. This is the the­sis be­hind pgrust. Databases his­tor­i­cally were the hard­est piece of soft­ware to build and were lim­ited be­cause of that. Now, with AI, we can be more am­bi­tious about the type of soft­ware we build.

Optimizing for a class of work­load

Let’s try this out with FRE, the regex en­gine we built in the last post. Recall that it was cre­ated by hav­ing an agent loop for a month on im­prov­ing regex en­gine per­for­mance with ac­cess to the re­bar regex bench­mark suite. This re­sulted in FRE be­ing heav­ily over­fit to re­bar un­til we warned our agent that we had a hold­out bench­mark, which caused the agent to gen­er­al­ize the op­ti­miza­tions enough that per­for­mance was ok-ish on our hold­out. There’s no par­tic­u­lar rea­son to use a software fac­tory” regex en­gine that does­n’t beat a well-tested regex en­gine on hold­out bench­marks, but one no­table thing about FRE was that the na­tive AOT com­piled ver­sion did quite well at longer searches. We noted that, it stands to rea­son that one could run the na­tive code com­piler in an­other thread while rip­grep was run­ning its nor­mal matcher and then cut over to the na­tive code when it fin­ished com­pil­ing and gen­er­ally get bet­ter per­for­mance. Of course this will gen­er­ally re­sult in worse per­for­mance for short queries as we lose a thread to com­pi­la­tion, but I care a lot more about how long rip­grep takes when it runs for many sec­onds or min­utes than when it runs for a few sec­onds, so I’m ok with that trade­off.

In the same way we could build a regex en­gine in a few min­utes of hu­man time, we can also just try this ex­per­i­ment in a few min­utes of hu­man time. I typed a few sen­tences and an agent went and did the work to al­low this to hap­pen (which would be a de­cent chunk of code surgery for a hu­man) and it ran the bench­mark on ac­tual rip­grep queries that come from my codex his­tory. For longer queries, we see a 2x-4x per­for­mance im­prove­ment here for a few very sim­ple queries. But most queries are more com­plex, and when we run on rep­re­sen­ta­tive hold­out queries, for queries where AOT should be en­abled2, we get about a 7% speedup. Not an earth shat­ter­ing re­sult, but also not a bad out­come for spend­ing a few min­utes typ­ing to codex (and it’s still do­ing more op­ti­miza­tion and will pre­sum­ably speed things up fur­ther).

Build an in­dex?

This is ar­guably a silly thing to do, since if we’re re­peat­edly search­ing for text on a com­puter, the ob­vi­ous thing to do to speed that up is­n’t to write a na­tive code com­piler for regex match­ing, it’s to cre­ate an in­dex. But the point here is just that this kind of tech­ni­cal work, which used to take a fair amount of time and ex­per­tise, can just be done triv­ially now. And if we wanted to build a text in­dex, it just so hap­pens that I worked on BitFunnel, the Bing search in­dex that was spe­cial­ized for con­stant/​fast text in­ges­tion that won Best Paper Award at SIGIR, so I can think of a few ex­per­i­ments to try if we’re go­ing to build a fast lo­cal in­dex of our en­tire ma­chine (the pro­jects I’ve seen seem to be in­tended to in­dex your code di­rec­to­ries, but what re­ally kills my ma­chine per­for­mance is when codex de­cides to run rip­grep against huge tem­po­rary di­rec­to­ries with a ton of gen­er­ated files and then ex­pands to look­ing at my whole ma­chine when it misses, so I’d want an in­dex of my en­tire disk and not just of the code for some pro­jects).

If I were work­ing at an AI lab and had ac­cess to things like SOTA mod­els run­ning on Cerebras chips or other ac­cel­er­a­tors that greatly in­crease tok/​s and there­fore load/​de­mand for search, I might ac­tu­ally sur­vey the ex­ist­ing in­dex­ers to see if they’re fast enough or if I’d want to build some­thing cus­tom my­self. While the open source ver­sion of BitFunnel only” con­tains a byte­code in­ter­preter and one JIT, the Bing ver­sion con­tains mul­ti­ple JIT com­pil­ers. A pro­ject that did that level of op­ti­miza­tion used to be a ma­jor un­der­tak­ing, but I could do that in a week­end” is now ac­tu­ally true for some of these kinds of pro­jects. With my lowly $200/mo ac­count, I think a some­what faster rip­grep plus any off-the-shelf in­dex is fine, so maybe this fast-in­gest­ing whole-ma­chine in­dex pro­ject can be left as an exercise for the reader (who works at an AI lab)”.

Optimizations are cheap

The dras­tic re­duc­tion in the cost of op­ti­miza­tions has been true go­ing back to November 2025 and maybe even some­what be­fore then with pub­lic mod­els (and I’m sure be­fore that still with what folks at AI labs had ac­cess to). For an ex­am­ple from the GPT-5.1 or 5.2 days, with no knowl­edge of game AIs, I tried build­ing an Azul AI. This ended up be­ing the strongest AI in the world for the game by a pretty large mar­gin. From read­ing the the­sis that de­scribes the 2nd strongest AI, I think my AI is prob­a­bly a bit bet­ter on the AI side of things, but the main place it wins is on op­ti­miza­tion de­spite spend­ing what looks like maybe two or­ders of mag­ni­tude less time (estimated by read­ing the the­sis and see­ing the process and com­par­i­son to my process) and also mostly work­ing on my lap­top vs. hav­ing a clus­ter of ma­chines to use (which means much less band­width to run ex­per­i­ments with, do pa­ra­me­ter tun­ing, etc.). For ex­am­ple, that other AI is sin­gle-threaded and my AI is multi-threaded. Since I have a na­tive code ver­sion as well as a heinous wasm shared mem­ory + javascript ver­sion, and two dif­fer­ent search ar­chi­tec­tures for two dif­fer­ent ver­sions, which require” com­pletely dif­fer­ent multi-thread­ing al­go­rithms (minimax for a very small and fast net and MCTS for a larger net), this would’ve been a fairly large un­der­tak­ing if done by hand. And, be­cause I let an LLM pick the multi-thread­ing al­go­rithm based on its own (incorrect) rea­son­ing a cou­ple times be­fore spend­ing 30 min­utes read­ing about multi-thread­ing al­go­rithms for game AIs my­self, I ended up re-writ­ing (having codex re-write) the multi-thread­ing al­go­rithm mul­ti­ple times.

There’s a bunch of stan­dard stuff it makes sense to do to de­bug and ver­ify a mul­ti­thread­ing al­go­rithm for some­thing like this, like im­ple­ment­ing re­play from de­bug logs that can re­pro­duce bugs de­spite the al­go­rithm be­ing non­de­ter­mistic. Doing that alone would’ve prob­a­bly been days to a week of work had I done it by hand, but it’s ex­actly the kind of thing an agent can triv­ially do in a loop (just have it try to re­play logs and in­sert log­ging for non-de­ter­min­ism every time you don’t get a per­fect re­play). A lot of the te­dium it used to take to get a tricky op­ti­miza­tion like this work­ing is gone.

This also ap­plies to a lot of other tricky op­ti­miza­tions. From hav­ing writ­ten CPU mi­croc­ode, done CPU ver­i­fi­ca­tion, worked on op­ti­miz­ing a search en­gine in­dex, etc., I have a lot of ex­pe­ri­ence look­ing at op­ti­miza­tions and think­ing hmm, this would in­crease per­for­mance by 2%, but it’s go­ing to take N per­son-days to ver­ify that this tricky op­ti­miza­tion works” and mak­ing a call to go ahead or not based on whether or not it’s worth the time to get the op­ti­miza­tion work­ing. Now that this N has dropped by a tremen­dous fac­tor (variable but, in terms of hu­man time, fre­quently 1000x / 10000x / 1000000x, prob­a­bly more like 1000x on dol­lar cost if you com­pare to­ken costs at me­tered rates vs. the Bing en­gi­neer who wrote the com­pil­ers at JITs that the search in­dex used), the num­ber of these kinds of op­ti­miza­tions it makes sense to do goes way up. The same goes for op­ti­miza­tions that you aren’t sure will work out. I used to some­times look at an op­ti­miza­tion that I was­n’t sure would speed things up and think this will take M hours to im­ple­ment to the point where we have a good enough mea­sure­ment to guess at the per­for­mance im­pact”. Many more of those op­ti­miza­tions make sense to try out now.

Going back to the game AI case, at least for the AI I tried, it seems like you gain about 100 Elo for every dou­bling in speed (more than in chess, I sus­pect be­cause draws are very rare). Just adding mul­ti­thread­ing alone is enough to wipe the floor with an oth­er­wise com­pa­ra­ble AI on a large ma­chine. If you stack in 10 – 20 more op­ti­miza­tions that seem too an­noy­ing for most peo­ple to do by hand, the dif­fer­ence in strength is tremen­dous and it’s not re­ally rea­son­able to try to keep up with a hand-writ­ten AI3.

The game AI case is a lit­tle more an­noy­ing than for most soft­ware be­cause a lot of the op­ti­miza­tions you want to do ac­tu­ally change the re­sult and there is­n’t a cheap, triv­ial, way to tell if the speed in­crease + the change in re­sult gives a bet­ter or worse ac­tual re­sult in prac­tice. And, as we noted be­fore, cur­rent pub­licly avail­able SOTA mod­els are pretty bad at ex­per­i­men­tal de­sign, so I had to set up the frame­work they used to de­ter­mine if an op­ti­miza­tion is good, but once that was in place, it’s like any other op­ti­miza­tion prob­lem. I guess peo­ple work­ing on LLM op­ti­miza­tions also have to deal with this class of prob­lem but most op­ti­miza­tion prob­lems are a lot more straight­for­ward.

To pick an­other ex­am­ple, as part of prepar­ing for per­for­mance in­ter­views, Jamie Brandon tried Anthropic’s now pub­lic per­for­mance take­home. After try­ing it, he had Claude pick up where he left off and it got a much bet­ter re­sult. When he looked at what Claude did that he did­n’t, he said a lot of the op­ti­miza­tions were things that oc­curred to him but he had­n’t got­ten to yet, and [o]thers were just crazy shit that I would never try un­less I was work­ing on this for weeks”4. He’s a rea­son­able per­for­mance en­gi­neer and he got an of­fer for the per­for­mance job he wanted, but on a well-de­fined op­ti­miza­tion prob­lem, he does­n’t stand a chance against a de­cent model (I haven’t tried the prob­lem my­self, but I sus­pect I also would­n’t stand a chance given re­motely com­pa­ra­ble time con­trols).

Workload-specific op­ti­miza­tion

Coming back to this part of Marc Brooker’s com­ment:

Dynamic cus­tom soft­ware, fit­ted to a par­tic­u­lar work­load rather than a class of work­loads, seems like a very likely out­come.

Dynamic cus­tom soft­ware, fit­ted to a par­tic­u­lar work­load rather than a class of work­loads, seems like a very likely out­come.

This seems pretty in­evitable. In an­other re­sponse to my post, Michael Malis of pgrust said some­thing sim­i­lar:

[discussion of pgrust op­ti­miza­tions] … I think it’s easy enough to cre­ate these op­ti­miza­tions that we could look at a cus­tomers work­load and add them as needed

[discussion of pgrust op­ti­miza­tions] … I think it’s easy enough to cre­ate these op­ti­miza­tions that we could look at a cus­tomers work­load and add them as needed

Without hav­ing any kind of frame­work or setup, right be­fore I started writ­ing this post, I had an agent do work­load-spe­cific op­ti­miza­tion for my rip­grep queries (not the na­tive code com­piler switch, just the op­ti­miza­tions to the gen­eral FRE en­gine based on a set of bench­marks), which took about 2 min­utes for me to launch. The op­ti­miza­tions run on a set of queries, and then there’s a later hold­out set of queries to run against. That’s still run­ning, but the ini­tial re­sults seem promis­ing. After one pass of op­ti­miza­tion, the work­load op­ti­mized ver­sion is 2% faster than stan­dard rip­grep on the hold­out and it’s still get­ting faster. 2% is­n’t a big deal for my lo­cal rip­grep us­age, but con­sid­er­ing that this took min­utes of time and the op­ti­miza­tions done here got started when I started typ­ing this point and are still im­prov­ing, I’d take a 2% win here (note that this is­n’t com­bined with the na­tive code com­piler, which would give a larger over­all win if com­bined prop­erly). And re­call that this is lever­ag­ing the FRE regex en­gine5, which was sub­stan­tially slower than the Rust regex en­gine on hold­out bench­marks and was stuck with slow im­prove­ment on hold­outs be­cause with me know­ing noth­ing about regex work­loads and SOTA LLMs not be­ing good enough at ex­per­i­men­tal de­sign to do un­guided open-ended self-im­prov­ing loops, we did­n’t have a good way to im­prove per­for­mance on our hold­outs. But if what I care about is per­for­mance on my own work­loads, I have plenty of data and am gen­er­at­ing more all the time. As Marc Brooker noted above, we do have to be care­ful about over­fit­ting if there’s a regime change that’s not in the old data, etc., but we’re still in a bet­ter sit­u­a­tion than we were be­fore.

In the more gen­eral case, if you’re some­one like Marc Brooker at Amazon or Michael Malis work­ing on pgrust, it makes sense to not just do this as a one-off, but to work with cus­tomers to pi­lot a pro­gram that uses their data to op­ti­mize things for them and then fig­ure out how to scale it out for cus­tomers in gen­eral. I’m not work­ing at a com­pany where that’s the best use of my time6, but it’s pretty wild that you can see that this is com­ing for larger com­pa­nies with more scale, and given that it only takes min­utes of my time to run these ex­per­i­ments for my per­sonal work­flows, it’s pretty rea­son­able to mess with this kind of thing on per­sonal pro­jects.

Thanks to Jamie Brandon, Michael Malis, and Max Bittker for com­ments/​cor­rec­tions/​dis­cus­sion.

P.S. As I’ve noted in the last cou­ple posts, with cod­ing agents, the time it takes to run an ex­per­i­ment and see enough of a re­sult to sat­isfy my cu­rios­ity has gone way done while the time it takes to make a re­sult re­ally rig­or­ous has­n’t changed or has gone up, so writ­ing things up the way I used to would mean run­ning very few ex­per­i­ments rel­a­tive to the band­width I have for them. As a re­sult, I’ve just been run­ning these ex­per­i­ments and shar­ing the re­sult with a cou­ple of friends. As an ex­per­i­ment, I’m try­ing to write these up in a very quick and non-rig­or­ous way in­stead of years of these ex­per­i­ments only be­ing known to a few friends. Like the last post, I set a goal of writ­ing this post and do­ing all the clean-up in half an hour and did­n’t time it but am pretty sure I missed that by a bit.

Even do­ing this, the time it takes to write these up is long enough that I’m falling be­hind on shar­ing re­cent re­sults, but I’m not in­clined to switch to LLM-written posts (yet?), and I don’t think I can re­al­is­ti­cally get the time to clean up the data and write a post like this down enough to turn a post around in less than half an hour. Just on the length of this post, typ­ing this up should be some­thing like 20 – 30 min­utes in­clud­ing time to pause and think about what I’m writ­ing, and then when I look at the data some­times some­thing will look wrong enough that I need to look into it more closely to see if there’s an is­sue that needs to be fixed (this hap­pened mul­ti­ple times here, and I would ex­pect that, be­cause I did­n’t spend much more time, there are other data is­sues that I don’t know about).

Anyway, if you have opin­ions on these quick (and surely more wrong) writeup, let me know what you think (X Bsky Mastodon)!

Appendix: There’s no rea­son for soft­ware to be slow any­more

I’ve been on the record for a long time as strongly dis­agree­ing with the gen­eral sen­ti­ment that the de­vel­op­ers of X are bad and should feel bad for writ­ing slow code be­cause there are a lot of dif­fer­ent kinds of pro­gram­ming ex­per­tise and not only is it not the case that most pro­gram­mers don’t have per­for­mance ex­per­tise, it prob­a­bly does­n’t even make sense for them to de­vel­op­ment (from the stand­point of what the busi­ness cares about, what the em­ploy­ment mar­ket looks like, etc.), so of course most pro­jects will have very poor per­for­mance com­pared to what a per­for­mance ex­pert can do.

For the ex­am­ple above, Jamie Brandon got an of­fer from Anthropic and you prob­a­bly can’t af­ford him or some­one like him un­less you’re OpenAI, but you can af­ford to use a cod­ing agent that can beat him on a bounded op­ti­miza­tion prob­lem. The agent does­n’t have the judge­ment he has and will do worse on an open-ended prob­lem (recall that when we tried build­ing an op­ti­mized regex en­gine and just told it to not over­fit, it was more than an or­der of mag­ni­tude worse than the best regex en­gines on our hold­out bench­marks, but also re­call that af­ter telling the agent there was a hold­out it was do­ing poorly on, it sped up regex en­gine per­for­mance enough to gen­er­ally match 2nd tier regex en­gines in terms of per­for­mance, which is still ex­tremely good com­pared to the gen­eral level of per­for­mance op­ti­miza­tion in most code to­day), but that’s plenty good to achieve rea­son­able per­for­mance on all sorts of prob­lems. This post has gen­er­ally dis­cussed back­end per­for­mance is­sues, but agents don’t seem worse at front-end per­for­mance if you want to drive down a set of met­rics like LCP, INP, etc.

Appendix: How is codex run­ning rip­grep?

Here’s some in­for­ma­tion about the dis­tri­b­u­tion of riprep queries on my ma­chine. I make no claims that this is at all rep­re­sen­ta­tive of what’s hap­pen­ing any­where else. The pat­tern dis­tri­b­u­tion of the length of the pat­tern that’s searched has a lot more long pat­terns that I would’ve ex­pected. The p50 is 55 uni­code code points (I’ll just call these char­ac­ters for sim­plic­ity), which is al­ready longer than things I grep for by hand, and the p90 is 119!

We can also look at the num­ber of al­ter­na­tion arms in regexes, which are once again much more com­plex than what I do by hand.

Another view is to look at how these are cor­re­lated. Do we get more al­ter­na­tion arms in the regexes as the regexes get longer? Yes.

What are these re­ally long regexes, any­way? If we look at them, most of the longest are long al­ter­na­tions over func­tion or tests names, such as the fol­low­ing regex, which ap­pears to be re­lated to FRE de­vel­op­ment.

fn (hot_byte_compiler_is_generic_only_and_anonymous_count_uses_auto_count| one_­pat­tern_­coun­t_s­pan­s_us­es_the_re­tained_­com­plete_s­pan_ses­sion| for­mal_­com­pact_s­tate_byte_vis­i­tors_­co­ex­ist_with­_­na­tive_­count| fixed_bound­ary_record_vis­it_­match­es_­line_rel­a­tive_ref­er­ence_and_is_atomic| un­bound­ed_lan­guages_refuse_fi­nite_ex­trac­tion_be­fore_al­lo­ca­tion| for­mal_s­in­gle_raw_s­pan_sweep­_pre­flight| as­sert_ex­ac­t_­fix­ture_us­es_­for­mal_large_­con­tin­u­a­tion_sweep| url_on­ly_­com­pile_i­den­ti­ty_bind­s_lan­guage_and_own­er_­mode| url_on­ly_­com­pile_ex­ac­t_lim­it­s_and_run­time_re­fusal­s_­close| url_on­ly_­com­pile_­post_­plan_al­lo­ca­tion_­fault­s_­close| url_on­ly_own­er_dis­crim­i­na­tor_is_sta­ble_and_precharged| url_on­ly_­com­pile_own­er_is_s­trat­e­gy_and_­op­er­a­tion_s­coped| for­mal_re­bar_url_own­er_is_­com­pile_on­ly_and_­match­es_o­r­a­cle| for­mal_re­bar_url_ex­ac­t_­fix­ture_us­es_cer­ti­fied_ex­e­cu­tion| for­mal_­fixed_schema_­ma­te­ri­al­iza­tion_­match­es_both­_record_o­r­a­cles_and_­con­trols| for­mal_s­in­gle_­coun­t_s­e­lect­s_­com­pact_s­tate_byte_­com­plete_bound­_vis­i­tors| au­then­ti­cat­ed_bound­_­line_­to­tal_lf_free_­do­main_op­por­tu­ni­ty_ex­ceed­s_­five_per­cent| pre­pared_ab­solute_onepass_­fus­es_s­lot­s_and_p­re­serves_pre_­source_­fall­back| au­then­ti­cat­ed_­word_bound­ary_russ­ian_­com­pact_low­er­ing_pub­lic_­ca­nary| or­dered_n­fa_x86_ep­silon_edges_by­pass_the_as­ser­tion_­call| or­dered_n­fa_aarch64_ep­silon_edges_by­pass_the_as­ser­tion_­call| or­dered_edge_dis­patch_v2_is_­tar­get_neu­tral_de­ter­min­is­tic_and_re­lo­ca­tion_free| or­dered_edge_dis­patch_v2_­copies_­canon­i­cal_ta­bles_and_­cap_­fall­s_back­_­to_v1| or­dered_n­fa_v3_­com­pos­es_ter­mi­nal_range_and_dis­patch_with­out_­da­ta_re­lo­ca­tions| or­dered_n­fa_x86_ter­mi­nal_range_emit­s_au­then­ti­cat­ed_re­verse_s­can| or­dered_n­fa_aarch64_ter­mi­nal_range_emit­s_au­then­ti­cat­ed_re­verse_s­can| or­dered_n­fa_x86_bound­ary_as­ser­tion_­cache_is_lazy_and_bound­ary_s­coped| or­dered_n­fa_aarch64_­caches_re­peat­ed_as­ser­tion­s_once_per_bound­ary| bound­ary_as­ser­tion_­cache_re­quires_­dense_ex­ac­t_kind_reuse| bound­ary_as­ser­tion_­cache_s­e­lec­tion_is_­com­pil­er_on­ly_and_de­ter­min­is­tic)

But some are funny nu­mer­i­cal con­struc­tions, such as

:(13[0 – 9]|14[0 – 9]|15[0 – 9]|16[0 – 9]|17[0 – 9]|18[0 – 9]|19[0 – 9]|20[0 – 9]|21[0 – 9]|22[0 – 9]|23[0 – 9]|24[0 – 9]|25[0 – 9]|26[0 – 9]|27[0 – 9]|28[0 – 9]|29[0 – 9]|30[0 – 9]|31[0 – 9]|32[0 – 9]|33[0 – 9]|34[0 – 9]|35[0 – 9]|36[0 – 9]|37[0 – 9]|38[0 – 9]|39[0 – 9]|40[0 – 9]|41[0 – 9]|42[0 – 9]|43[0 – 9]|44[0 – 9]|45[0 – 9]|46[0 – 9]|47[0 – 9]|48[0 – 9]|49[0 – 9]|50[0 – 9]|51[0 – 9]|52[0 – 9]|53[0 – 9]|54[0 – 9]|55[0 – 9]|56[0 – 9]|57[0 – 9]|58[0 – 9]|59[0 – 9]|60[0 – 9]|61[0 – 9]|62[0 – 9]|63[0 – 9]|64[0 – 9]|65[0 – 9]|66[0 – 9]|67[0 – 9]|68[0 – 9]|69[0 – 9]|70[0 – 9]|71[0 – 9]|72[0 – 9]|73[0 – 9]|74[0 – 9]|75[0 – 9]|76[0 – 9]|77[0 – 9]|78[0 – 9]|79[0 – 9]|80[0 – 9]|81[0 – 9]|82[0 – 9]|83[0 – 9]|84[0 – 9]|85[0 – 9]|86[0 – 9]|87[0 – 9]|88[0 – 9]|89[0 – 9]|90[0 – 9]|91[0 – 9]|92[0 – 9]|93[0 – 9]|94[0 – 9]|95[0 – 9]|96[0 – 9]|97[0 – 9]|98[0 – 9]|99[0 – 9])[0 – 9]:

This is equiv­a­lent to :(?:1[3 – 9]|[2 – 9][0 – 9])[0 – 9]{2}: (which, if run through rip­grep on the orig­i­nal in­put, has ap­prox­i­mately the same per­for­mance). The en­tire pipeline for that was

cargo clippy … | rg crates/fre-aot-regex/src/module.rs:’ | rg NUMBER_REGEX | head -250

which might be an odd thing for a hu­man to do, but agents seem to do this kind of thing all the time.

On an­other topic, if we look at how long rip­grep queries took, there are quite a few slow queries, e.g., p99 is al­most 1 minute! And p999 is al­most 10 min­utes! And the max­i­mum query over this time pe­riod (around a month on one lap­top; queries and dis­tri­b­u­tions seem likely to be dif­fer­ent on the AWS hosts I run agents on, etc., but I haven’t checked) is ap­proach­ing 2 hours!

In terms of com­mand line op­tions, we see the fol­low­ing. Perhaps un­sur­pris­ingly, codex of­ten wants line num­bers and, for what­ever rea­son, it very oc­ca­sion­ally uses PCRE2 regexes.

I won’t add plots or ta­bles for these, but an­other thing to note is that there’s fairly low lo­cal­ity for what pat­terns are searched for (about 94% of pat­terns only oc­curred once), which makes some sense given how long a lot of the queries were. However, there’s fairly high lo­cal­ity in what files get searched and a file that got searched is rel­a­tively likely to get searched again soon, in­di­cat­ing that (for small enough files), they’re likely to be searched in mem­ory.

Also, 99% of queries were regex queries (1% were non-regex string searches) and 99.9% of search queries were ASCII only, but in terms of files searched, ap­prox­i­mately 45% were ASCII only and 55% con­tained Unicode, a higher per­cent­age than I would’ve guessed for Unicode.

On a draft of the last post, Peter Geoghegan noted

It’s also pos­si­ble for a regex im­ple­men­ta­tion to be faster by sup­port­ing fewer fea­tures. Some im­ple­men­ta­tions don’t sup­port back ref­er­ences, etc.

It’s also pos­si­ble for a regex im­ple­men­ta­tion to be faster by sup­port­ing fewer fea­tures. Some im­ple­men­ta­tions don’t sup­port back ref­er­ences, etc.

which is also true here. The work­load-spe­cific op­ti­miza­tions done here were fairly su­per­fi­cial be­cause I just gave codex some short in­struc­tions and let it do what­ever it wanted (which is, in gen­eral, not the most ef­fec­tive use of codex), but with a more de­tailed plan, more fo­cused op­ti­miza­tions sup­port­ing the com­mon use cases for my queries could be ex­pected to yield larger gains.

though, as we dis­cussed in that post as well as be­fore, the bench­mark­ing and ex­per­i­men­tal de­sign skills of SOTA mod­els aren’t good enough to do this in the gen­eral case with­out a hu­man (or a skill) set­ting up the bench­mark­ing en­vi­ron­ment for the agent. [return]

we can see from our old bench­marks that, even with time to run the com­piler, there are a lot of cases where the na­tive code com­piled ver­sion is slower than the Rust regex crate. If we look at why this is, these tend to be more com­plex queries where the Rust regex crate has some al­go­rith­mic op­ti­miza­tion and the FRE na­tive code com­piler is falling back to some­thing naive (the agent that cre­ated FRE spent much less time on the na­tive code com­piler than it did on the normal” regex en­gine). [return]

I have no doubt that a hand-writ­ten AI by some­one who has real AI ex­per­tise, e.g., by some­one who’s writ­ten one of the top Go and chess en­gines in the world, could beat my AI on the strength of the AI side of things be­ing bet­ter than what you get when some­one who knows noth­ing about AI (me) cre­ates an AI, but if the lev­els of ex­per­tise are re­motely sim­i­lar, the LLM-written ver­sion is go­ing to dom­i­nate for any given amount of time spent. [return]

it’s ar­guably un­fair to com­pare the re­sult of an agent pick­ing up where he left off, since his work is a start­ing point which might let an agent do much bet­ter than it would do on its own, so I tried giv­ing the fresh task to an agent and it got a very sim­i­lar score to what he got when an agent re-used his work (and a quick check by an­other agent did­n’t find ev­i­dence of cheat­ing). [return]

The per­for­mance prob­a­bly would’ve been bet­ter if I had an agent just mod­ify a rip­grep fork di­rectly, but I was cu­ri­ous if this could also solve the FRE over­fit­ting prob­lem with re­spect to my queries. [return]

a while back, I re­duced the size of page in our signup flow from 50 MB to 5 MB and a rev­enue A/B test seemed to in­di­cate that this in­creased rev­enue by about 0.5%. In gen­eral, I’m a huge fan of do­ing the sim­ple and easy wins first, such as this, and there are prob­a­bly a lot of higher ROI wins than we’d get out of build­ing cus­tom com­pil­ers or do­ing other highly spe­cial­ized tech­ni­cal work here. [return]

I'm becoming AI-blind - Rafal Cymerys

cymerys.com

AI AI AI

Recently I’ve been catch­ing my­self hav­ing these lit­tle mo­ments at work, when I’m try­ing to read a doc­u­ment some­one has sent me and my brain some­how re­fuses to an­a­lyze it. It feels like I’m read­ing it, but I’m un­able to fo­cus on its con­tent.

I end up get­ting dragged into an end­less back and forth with the sender, ask­ing ques­tions about things that have been cov­ered in what they’ve al­ready sent me. It’s rather con­cern­ing, be­cause I’ve spent the last year try­ing to re-learn how to fo­cus and these sit­u­a­tions show the ex­act op­po­site.

I sat down to an­a­lyze these sit­u­a­tions and re­al­ized they all have a com­mon de­nom­i­na­tor: the doc­u­ments all show a strong trace to AI.

For ex­am­ple:

A de­sign doc­u­ment that looks like a copy-paste from Claude. While it does cover the de­sign of the spe­cific fea­ture in ques­tion, it also car­ries a lot of Claude-specific analy­sis and lingo. This cuts just through it”, The first gate is real”.

A de­sign doc­u­ment that looks like a copy-paste from Claude. While it does cover the de­sign of the spe­cific fea­ture in ques­tion, it also car­ries a lot of Claude-specific analy­sis and lingo. This cuts just through it”, The first gate is real”.

or

A 20-page mar­ket­ing con­cept deck that mixes up (a rather rea­son­able) mar­ket­ing strat­egy with some non­sense prod­uct tech­ni­cal ar­chi­tec­ture gib­ber­ish. How does it pitch the idea? It’s not sell­ing X, it’s sell­ing Y”. The Redis back­bone re­de­fines the prod­uct”.

A 20-page mar­ket­ing con­cept deck that mixes up (a rather rea­son­able) mar­ket­ing strat­egy with some non­sense prod­uct tech­ni­cal ar­chi­tec­ture gib­ber­ish. How does it pitch the idea? It’s not sell­ing X, it’s sell­ing Y”. The Redis back­bone re­de­fines the prod­uct”.

or

A tech­ni­cal re­quire­ments doc­u­ment that de­scribes a rather sim­ple con­cept in a very ver­bose way. The thing is, a lot of this doc­u­ment reads like some­one’s internal” rea­son­ing that’s not fully sure about cer­tain de­ci­sions. Sounds like an LLM to me.

A tech­ni­cal re­quire­ments doc­u­ment that de­scribes a rather sim­ple con­cept in a very ver­bose way. The thing is, a lot of this doc­u­ment reads like some­one’s internal” rea­son­ing that’s not fully sure about cer­tain de­ci­sions. Sounds like an LLM to me.

There’s an on­go­ing dis­cus­sion of whether hu­mans are good at rec­og­niz­ing AI-generated text. While most re­search claims that hu­mans don’t re­ally do a good job there, I dis­agree. It’s not that dif­fi­cult, at least when we’re talk­ing about the low-ef­fort re­sults. Florian Roth wrote a pretty good sum­mary of the com­mon pat­terns in the con­text of so­cial me­dia.

I see a sim­i­lar thing hap­pen­ing for work-re­lated texts. Besides the ob­vi­ous choice of words, the gen­eral flow of sen­tences and the at­tempt to pitch every small de­tail as a break­through quickly give it away. If your doc­u­ment de­scribes the check­boxes in an RBAC con­fig­u­ra­tion view for an en­ter­prise ap­pli­ca­tion, don’t sell it like you’ve just in­vented fire.

I feel like I’ve been pre-trained” on all the AI-generated LinkedIn posts, emails and web­sites that are full of text but empty on mean­ing. My brain learned to quickly spot signs of AI-generated con­tent, at least the con­tent gen­er­ated with low ef­fort, and it now ig­nores it and moves on with­out think­ing much about it.

I’ve heard some peo­ple com­par­ing it to banner blind­ness”. It’s not sur­pris­ing. With the amount of con­tent be­ing pushed at us, fil­ter­ing it out is how we need to stay sane.

What’s fas­ci­nat­ing to me, is that the same AI that was sup­posed to make me more pro­duc­tive, is what’s now slow­ing me down in an un­ex­pected way.

I don’t usu­ally go on va­ca­tion, but this year I re­ally needed a break. One evening I was re­ally hun­gry, walk­ing past some restau­rants on the Baltic coast. There was a sin­gle one I im­me­di­ately ig­nored, but a minute later some­thing in my head asked Hey, did they re­ally put up a photo of quiche with mold?” I walked back just to see this.

AI AI AI

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.