10 interesting stories served every morning and every evening.

Firefox is now the last major browser that still supports uBlock Origin

www.pcworld.com

When you pur­chase through links in our ar­ti­cles, we may earn a small com­mis­sion. This does­n’t af­fect our ed­i­to­r­ial in­de­pen­dence.

News

Aug 13, 2026

Firefox re­cently an­nounced via Bluesky post: Our sup­port for uBlock Origin is­n’t go­ing any­where.” The mo­ment comes in re­sponse to news that Microsoft Edge is soon go­ing to lock out uBlock Origin and other ad-block­ing ex­ten­sions that run on Manifest V2 ar­chi­tec­ture.

Once Microsoft Edge moves to Manifest V3, ad-block­ing ex­ten­sions won’t have ac­cess to the func­tions needed to prop­erly iden­tify and block ads that oc­cur while brows­ing web­sites and watch­ing videos.

Microsoft’s move is­n’t sur­pris­ing, as Edge is based on Chromium, the open-source browser en­gine that pow­ers most web browsers to­day, in­clud­ing Opera, Brave, Vivaldi, and Samsung Browser. Google ini­ti­ated the mi­gra­tion from Manifest V2 to V3 in Chrome/Chromium, and Microsoft Edge is now fol­low­ing Google’s lead.

But Firefox is one of the few web browsers re­main­ing that is­n’t based on Chromium, and it’s now the only ma­jor browser to still sup­port uBlock Origin. Neither Safari nor DuckDuckGo—the two other ma­jor non-Chromium browsers out there—sup­port uBlock Origin.

For die-hard uBlock Origin fans, Firefox ap­pears to be the only browser left with­out com­pro­mises. With any other browser, you’ll need to set­tle for uBlock Origin Lite (with fewer fea­tures and less ad-block­ing suc­cess) or what­ever built-in ad-block­ing fea­ture comes with the browser.

This ar­ti­cle orig­i­nally ap­peared on our sis­ter pub­li­ca­tion PC för Alla and was trans­lated and lo­cal­ized from Swedish.

Everything is about to “go dark”

blog.cryptographyengineering.com

I’m com­ing down from spend­ing a few days at Usenix Security, right here in my home­town of Baltimore. This means that my days have been taken up with two kinds of con­ver­sa­tion: first, ex­plain­ing to col­leagues why Baltimore is­n’t ac­tu­ally like The Wire. And sec­ond, try­ing not to talk about AI.

Here I’m go­ing to break that sec­ond rule.

I have many wor­ries about what AI means for our field, for var­i­ous de­f­i­n­i­tions of field”. But in this post I want to fo­cus on just one thing I’ve started wor­ry­ing about, and it’s a per­verse thing: specif­i­cally, I’m con­cerned that AI is go­ing to make soft­ware much too se­cure.

While that does­n’t sound so bad on the sur­face, there’s a con­se­quence to this. I mean some­thing very spe­cific: I’m con­cerned that U.S. in­tel­li­gence and law en­force­ment agen­cies are about to go dark, mean­ing: that they’re go­ing to sud­denly lose a huge por­tion of their ca­pa­bil­ity. And that this is­n’t go­ing to be sim­ply a prob­lem for those agen­cies, but also for those of us who value com­puter se­cu­rity and pri­vacy in gen­eral.

Going Dark, and the era of law en­force­ment hack­ing

To ex­plain how we got here, we need to talk about re­cent his­tory. This ac­tu­ally gives me a real ex­cuse to ref­er­ence The Wire, just be­cause it’s a per­fect snap­shot of what elec­tronic sur­veil­lance looked like way back in 2002. If you’ve seen the first sea­son, you’ll re­call that it’s about cops wire­tap­ping drug deal­ers who use pay­phones and burn­ers. The mo­bile phones in the show are rel­a­tively new tech­nol­ogy for the time, but from a tech­no­log­i­cal per­spec­tive noth­ing in this sce­nario would have shocked a cop who jumped for­ward from, say, 1989.

In less than a decade from the pre­mier, every­thing in those episodes be­came to­tally quaint.

The change be­gan in the late 2000s, thanks to the rise of smart­phones and tex­ting. Because smart­phones can ac­tu­ally store data as well as con­vey­ing it, the con­tents of those phones quickly be­came a use­ful new source of law-en­force­ment ca­pa­bil­ity. Or they were un­til 2010, when Apple be­gan en­crypt­ing iPhone stor­age us­ing a key de­rived from the user’s pass­code (Android phones fol­lowed shortly there­after.) The next year, Apple de­ployed end-to-end en­cryp­tion in iPhone text mes­sages. By 2014, a tiny tex­ting startup named WhatsApp had gath­ered 600 mil­lion users world­wide. By 2016 those users, now nearly a bil­lion strong, were all us­ing de­fault end-to-end en­crypted mes­sag­ing and calls. These two trends — the move from calls to texts, and texts to en­crypted data — hap­pened very rapidly. The chart be­low gives one view of the tran­si­tion:

The FBI and law en­force­ment agen­cies were not in­sen­si­tive to what was hap­pen­ing. In 2014, Director Comey an­nounced an ini­tia­tive called Going Dark, which would launch a national con­ver­sa­tion” about what providers could do — or be com­pelled to do — to make these new com­mu­ni­ca­tions me­dia leg­i­ble to law en­force­ment and coun­ter­in­tel­li­gence.

In 2016, the agency quit talk­ing and took their the­ory to court. When a ter­ror­ist at­tack left the FBI hold­ing a shooter’s locked iPhone, the agency or­dered Apple to give them ac­cess. The com­pany re­fused. What broke the stale­mate — and, to some ex­tent, ended Going Dark” it­self — was some­thing that nei­ther the FBI nor Apple ex­pected. An out­side com­pany an­nounced that there was no need for Apple’s as­sis­tance: they could sim­ply hack the phone.

The Apple v. FBI case has turned out to be a mi­cro­cosm of the whole Going Dark de­bate. For the next decade, law en­force­ment and in­tel­li­gence agen­cies con­tin­ued to ask for exceptional ac­cess” back­doors, but the ur­gency was gone. Agencies and man­u­fac­tur­ers both knew that law en­force­ment could and would pur­chase tar­geted hack­ing tools like GrayKey (for phone un­lock­ing), or even re­mote ex­ploita­tion tools like NSO Group’s Pegasus, if they needed them badly enough. Vendors like Apple and Google con­tin­ued to play a vig­or­ous de­fense, clos­ing vul­ner­a­bil­i­ties as soon as they learned about them. But of­fen­sive vul­ner­a­bil­ity hunters con­sis­tently man­aged to keep the edge.

And now there’s a very good chance that all this is about to be his­tory.

The era of AI bug hunt­ing is here

This April (just four months ago!) Anthropic an­nounced a new model called Mythos that hap­pened to be un­usu­ally skilled at soft­ware vul­ner­a­bil­ity find­ing. The U.S. gov­ern­ment tem­porar­ily blocked its ex­port, re­strict­ing ac­cess to U.S. agen­cies and trusted ven­dors. While the ban was dra­matic and made for good PR, it turned out to be mostly point­less. OpenAI, along with Chinese open-weight model labs like Z.ai and Moonshot, have since demon­strated that vul­ner­a­bil­ity find­ing is­n’t some­thing that a sin­gle lab is likely to hold a mo­nop­oly on. The list of se­ri­ous vul­ner­a­bil­i­ties that these mod­els have found is get­ting scarier (or more im­pres­sive) by the day.

At first glance, this might seems like good news for the of­fen­sive team, and for hack­ers in gen­eral. But I doubt that’s how this will play out in the long term. Defenders are now in the process of patch­ing every bug they can find — of­ten decades worth of bugs — and the back­log feels huge. But they’re mak­ing progress. Entire CI tool­chains are be­ing re­built to in­cor­po­rate AI-based vul­ner­a­bil­ity scan­ning be­fore soft­ware ever reaches the point where a hu­man will touch it. While I doubt this means that every bug will be found (even cal­cu­lat­ing the num­ber of bugs in a piece of code is prob­a­bly un­com­putable), in the real world, it does feel likely that we’re go­ing to hit some sort of a ceil­ing on the num­ber of use­ful bugs, and prob­a­bly we’ll hit it soon.

Thus: over the next two years, ma­jor pieces of soft­ware are likely to run out of re­motely-ex­ploitable bugs.

Obviously I think this is great. But for law en­force­ment and of­fen­sive in­tel­li­gence agen­cies, it’s go­ing to be a night­mare. For the first time since 2010, law en­force­ment might ex­pe­ri­ence what it looks like to re­ally go dark”, across a huge cat­e­gory of ad­vanced (well-maintained) de­vices and pieces of soft­ware.

So how is this a prob­lem?

The de­bate over exceptional ac­cess” mech­a­nisms never re­ally went away. In some places, like the UK, it even metas­ta­sized into some­thing worse. Here in the US it mostly went into hi­ber­na­tion. Some of the slow­down can le­git­i­mately be at­trib­uted to ex­pert push­back — aca­d­e­mics and in­dus­try en­gi­neers point­ing out the risk that back­doors might be abused by the very ad­ver­saries that Agencies are sup­posed to be pro­tect­ing us against. But I fear that this was less of a prin­ci­pled pause, and more of a mar­ket that was just pric­ing sup­ply.

The de­struc­tion of the low-hang­ing vul­ner­a­bil­ity fruit will make law en­force­ment (and in­tel­li­gence) agen­cies’ need much more acute. The de­mand for con­structed, in­ten­tional back­doors will re-start in earnest. The re­sult will be enor­mous pres­sure on in­dus­try to re-ar­chi­tect their sys­tems to make their sys­tems amenable to ex­cep­tional ac­cess. In some cases, gov­ern­ments will ask for these ca­pa­bil­i­ties in the ex­pec­ta­tion that they’ll be use­ful for spy­ing on other gov­ern­ments — a strat­egy that might have been un­de­tectable in the pre-AI era, but that prob­a­bly will be less pro­duc­tive now. The re­sults are un­pre­dictable. One re­sult might be that non-US gov­ern­ments en­tirely re­move their de­pen­dence on US soft­ware.

The worst part about this dy­namic is that these po­ten­tial new back­doors will prob­a­bly only af­fect the coun­tries that de­mand them, mean­ing that they will be pri­mar­ily use­ful for al­low­ing the US to weaken its own sys­tems. This will in turn al­low for­eign ad­ver­saries to find new ways to at­tack our com­mu­ni­ca­tions. This de­lib­er­ate self-sab­o­tage will hap­pen just at a mo­ment when we’re fi­nally get­ting a han­dle on se­cur­ing our own in­fra­struc­ture.

So what do we do about it?

I hon­estly have no idea. This is not a call to ac­tion for ex­perts to rally be­hind a so­phis­ti­cated plan. Like so many things about the AI rev­o­lu­tion, it’s just oc­cur­ring to me that we’re on a long greasy slide to a place that will look dif­fer­ent than where we are to­day. Just re­al­iz­ing this does­n’t mean that I have any strat­egy in mind to avoid it. In this case, we’re just go­ing to have to hope that this time we make the right choices, for no other rea­son than that they’re right.

The other Sean Byrne doesn't exist

conic.al

Earlier this year Apple de­nied me ac­cess to App Store Connect af­ter de­cid­ing that I matched some­one on a U.S. gov­ern­ment re­stricted-party list.

Their ex­pla­na­tion was fairly de­fin­i­tive:

The in­for­ma­tion you pro­vided fully matches one or more re­stricted par­ties on the U.S. gov­ern­ment con­sol­i­dated screen­ing list or an­other gov­ern­men­t’s sanc­tions list.”

The in­for­ma­tion you pro­vided fully matches one or more re­stricted par­ties on the U.S. gov­ern­ment con­sol­i­dated screen­ing list or an­other gov­ern­men­t’s sanc­tions list.”

They al­ready had my pass­port.

I replied with my full le­gal name, Sean Joseph Byrne, up­loaded my dri­ver’s li­cense, and pointed out the ad­dress on the gov­ern­ment record they ap­peared to be match­ing me against. I’ve never lived at that ad­dress, never lived in County Sligo, and have no con­nec­tion to the com­pany in­volved. I asked them to es­ca­late it to their sanc­tions com­pli­ance folks and make a proper non-match de­ter­mi­na­tion.

Apple still has­n’t replied.

Apple’s re­sponse af­ter re­view­ing my iden­tity in­for­ma­tion.

I knew what had hap­pened be­cause this was­n’t the first time.

Cloonmull House

Search the U.S. gov­ern­men­t’s Consolidated Screening List for Sean Byrne and you get ex­actly one re­sult:

Sean Byrne Cloonmull House Drumcliffe, County Sligo Ireland

Source: Entity List, Bureau of Industry and Security Added: July 21, 2009 License re­quire­ment: All items sub­ject to the EAR License pol­icy: Presumption of de­nial

Sean Byrne Cloonmull House Drumcliffe, County Sligo Ireland

Source: Entity List, Bureau of Industry and Security Added: July 21, 2009 License re­quire­ment: All items sub­ject to the EAR License pol­icy: Presumption of de­nial

The Consolidated Screening List is­n’t it­self a sanc­tions list. It’s a U.S. gov­ern­ment screen­ing tool that com­bines a num­ber of ex­port-con­trol, sanc­tions and other re­stricted-party lists main­tained by the Departments of Commerce, State and Treasury.

The re­sult comes from the Commerce Department’s Bureau of Industry and Security Entity List. All items sub­ject to the EAR means the Export Administration Regulations, the rules gov­ern­ing what U.S. com­pa­nies can ship abroad. Presumption of de­nial” is a li­cens­ing pos­ture: if some­one ap­plies for a li­cence to ex­port some­thing to this per­son, the de­fault an­swer is no. That’s the en­tire pur­pose. It’s an ex­port-con­trol in­stru­ment but it says noth­ing about who can be em­ployed, or who can sell shares.

The per­son in the search re­sult is­n’t me. More in­ter­est­ingly, it does­n’t ap­pear to be any­one.

The en­try came out of the pros­e­cu­tion of an Irish air­craft-parts busi­ness called Mac Aviation. In 2009, the Department of Justice de­scribed Sean Byrne as Mac Aviation’s com­mer­cial man­ager and charged him along­side Thomas and Sean McGuinn over the il­le­gal ex­port of U.S. air­craft equip­ment to Iran.

Except Mac Aviation had ap­par­ently in­vented em­ploy­ees to make the com­pany look big­ger than it was.

Mac Aviation was a fa­ther and son work­ing out of a cot­tage on the edge of Drumcliffe vil­lage, Ben Bulben be­hind it. A Rolls-Royce of­fi­cial who came to visit was re­port­edly speech­less. The com­pany he had been sell­ing he­li­copter en­gines to, and had taken for a global op­er­a­tion em­ploy­ing hun­dreds of pro­fes­sion­als, was a house in Sligo.

Drumcliffe, County Sligo, with Ben Bulben in the back­ground.

To keep up the im­pres­sion of a much larger firm, the McGuinns signed doc­u­ments with false names. Sean Byrne was one of them. John Mooney re­ported in the Sunday Times that the name ap­peared on so much com­pany pa­per­work that the American au­thor­i­ties became con­vinced Byrne ex­isted and tried to in­dict him.”

When DOJ filed a su­per­sed­ing in­dict­ment in 2010, re­plac­ing the orig­i­nal, Sean Byrne was no longer a de­fen­dant. The de­fen­dants were Mac Aviation and Thomas and Sean McGuinn.

More im­por­tantly, the su­per­sed­ing in­dict­ment re­peat­edly de­scribes Sean Byrne as an alias used by one or more co-con­spir­a­tors. The phrase ap­pears more than fif­teen times, at­tached to spe­cific in­voices, emails and an own­er­ship state­ment. Mac Aviation staff used the name with sup­pli­ers in the U.S. and cus­tomers in Iran.

At some stage the U.S. gov­ern­ment ap­pears to have worked out that Sean Byrne was­n’t ac­tu­ally a sep­a­rate per­son. And yet the en­try in the Entity List sur­vived. Sixteen years later it still has no date of birth, pass­port num­ber, mid­dle name or other use­ful per­sonal iden­ti­fier. It’s ba­si­cally a com­mon Irish name, an ad­dress in Sligo and Ireland.

Cloonmull House in Sligo was Thomas McGuinn’s home, and the in­dict­ment gives it as Mac Aviation’s reg­is­tered mail­ing ad­dress. So the en­try is­n’t a record of a man in Sligo. It’s a name at­tached to some­body else’s house. And I’ve never lived in Sligo.

This has hap­pened be­fore

Years be­fore I moved back to Ireland, I was sell­ing stock through a ten­der of­fer when Nasdaq stopped my or­der af­ter a back­ground check re­turned a match on my name.

Their Head of Account Management emailed me say­ing that the check had found a match as­so­ci­ated with a pre­vi­ous in­ci­dent and that he was con­fi­dent it was a false pos­i­tive, but com­pli­ance wanted ad­di­tional proof of my California ad­dress. On the phone he gave me more de­tail and specif­i­cally asked me about Mac Aviation. I ex­plained that I’d never had any­thing to do with the com­pany, pro­vided the ex­tra doc­u­men­ta­tion they wanted and the sale went through with an en­ter­tain­ing story to tell peo­ple.

Nasdaq’s re­sponse af­ter a back­ground check matched my name.

More re­cently I or­dered a Starlink mount­ing pole from California. DHL, ship­ping on be­half of SpaceX, told me the prob­lem was a re­stricted-party match and asked for my pass­port. I sent it, they cleared it and the pole ar­rived. DHL also would­n’t send a hat I’d or­dered in the U.S. on to me in Ireland with­out a copy of my pass­port.

The fake Sean Byrne was as­so­ci­ated with at­tempts to pro­cure he­li­copter en­gines, fighter-air­craft parts and other U.S. equip­ment for Iran. The real Sean Byrne oc­ca­sion­ally needs to pro­duce a pass­port be­fore some­one will send him a hat.

Apple is the odd one out. Nasdaq and the ship­pers both gen­er­ated false pos­i­tives, asked for enough in­for­ma­tion to re­solve them, and then re­solved them. Apple al­ready had my pass­port, re­ceived my dri­ver’s li­cense and a fairly de­tailed ex­pla­na­tion of ex­actly why the match was wrong, and still told me that I fully” matched a re­stricted party.

Twelve men named Robert Johnson

None of this is novel. In October 2006, 60 Minutes found twelve American men named Robert Johnson who all had trou­ble board­ing flights, and brought them to New York to­gether. A politi­cian, a soc­cer coach, busi­ness­men and a serv­ing mem­ber of the mil­i­tary.

The Robert Johnson they kept be­ing con­fused with was­n’t a man named Robert Johnson. It was a known alias of some­one con­victed of plot­ting to bomb a Hindu tem­ple and a cin­ema in Toronto, who by then had served his sen­tence and been de­ported to Trinidad. The air­line agents check­ing the twelve real ones against it had a name and noth­ing else. Not even a date of birth.

Asked about it, the head of the FBIs Terrorist Screening Center said Robert Johnson would never get off the list, and that any­one with the name would be in­con­ve­nienced every time they tried to check in.

I’m not on the Entity List. I’m be­ing misiden­ti­fied as an en­try on it. Apple did­n’t make that dis­tinc­tion.

The con­sumer ver­sion of this has been lit­i­gated. In 2005 Sandra Cortez was held up buy­ing a car in Colorado be­cause TransUnion matched her against a woman on the Treasury Department’s sanc­tions list who was born 27 years af­ter her. The credit bu­reau had com­pared first and last names only, not dates of birth. A jury awarded her dam­ages and the Third Circuit up­held it, de­scrib­ing the fail­ure to com­pare birth dates as rep­re­hen­si­ble. Sergio Ramirez had the same ex­pe­ri­ence at a car deal­er­ship six years later, and his case reached the Supreme Court in 2021.

In both cases the courts called for bet­ter match­ing. Compare the date of birth. Compare the mid­dle name.

There is no ver­sion of that avail­able to me. The list­ing has no date of birth to com­pare. No mid­dle name and no pass­port num­ber, be­cause the per­son does­n’t ex­ist. A screen­ing sys­tem that does its job per­fectly will still flag me, for­ever, on the only two facts the record con­tains: a com­mon Irish name and a coun­try.

Which is why ar­gu­ing with com­pa­nies one at a time is the wrong ap­proach.

The real fake Sean Byrne

Remote hir­ing has de­vel­oped a se­ri­ous iden­tity-fraud prob­lem. It has also de­vel­oped, some­what un­be­liev­ably, a North Korea prob­lem.

North Korean IT work­ers have been get­ting re­mote jobs at U.S. com­pa­nies us­ing stolen or fab­ri­cated iden­ti­ties, proxy in­ter­view­ers and U.S.-based laptop farms” that make work­ers over­seas ap­pear to be con­nect­ing from in­side the United States. The FBI has been warn­ing com­pa­nies about it, and the DOJ has pros­e­cuted schemes that suc­cess­fully placed work­ers at more than 100 U.S. com­pa­nies.

So if you’re hir­ing re­mote en­gi­neers, is this per­son ac­tu­ally who they claim to be?” is now a le­git­i­mate se­cu­rity prob­lem.

A new class of re­cruit­ing prod­ucts is be­ing built around that prob­lem. They sit in­side the soft­ware com­pa­nies use to man­age job ap­pli­ca­tions, the ap­pli­cant track­ing sys­tem or ATS.

Tofu is one of them. Their pitch is that they screen every ap­pli­cant across more than forty sig­nals be­fore a re­cruiter opens a ré­sumé, and that when a screened ap­pli­cant trig­gers a sanc­tions match the sig­nal routes straight to com­pli­ance re­view. They are ex­plicit that this has to hap­pen early: screen­ing at the back­ground-check stage is, on their ac­count, al­ready too late, so it should run at ap­pli­ca­tion sub­mis­sion be­fore any re­cruiter makes con­tact. They also say a can­di­date flagged by one of their cus­tomers is flagged across their whole cus­tomer net­work through their API. Brainner makes a sim­i­lar case, check­ing ap­pli­cants against 3.5 bil­lion data points and flag­ging high-risk pro­files be­fore a re­cruiter re­views them.

Tofu says its data­base is built from analysed ap­pli­cant pro­files. Their home­page says more than 18 mil­lion. Most of their other pages say more than 5 mil­lion.

The sanc­tions screen­ing these ven­dors de­scribe is OFAC and the Specially Designated Nationals list, which is the right list for the risk they’re sell­ing against: pay­ing wages to a sanc­tioned per­son. I’ve no ev­i­dence that ei­ther com­pany queries the BIS Entity List, and no idea whether any com­pany I’ve ap­plied to uses ei­ther prod­uct.

But screen­ing my name against U.S. re­stricted-party data pro­duces a false pos­i­tive. It did at Nasdaq, at SpaceX, at DHL and at Apple. Four times in six years. And the in­dus­try’s an­swer to re­mote-hir­ing fraud is to run that class of check ear­lier, be­fore a hu­man is in­volved, and prop­a­gate the re­sult across a net­work of em­ploy­ers.

Mac Aviation fab­ri­cated an em­ployee to make it­self look like a big­ger com­pany. That fab­ri­cated em­ployee ended up in an au­thor­i­ta­tive U.S. gov­ern­ment data­base. Sixteen years later, com­pa­nies are build­ing sys­tems us­ing that data­base to de­tect fake ap­pli­cants, fraud­sters, crim­i­nals and state-backed ac­tors be­fore they get through the hir­ing process.

Has this cost me a job? I don’t know. I’ve spent my ca­reer in in­for­ma­tion se­cu­rity, much of it in the U.S., and I’m now ap­ply­ing for roles from Ireland. There have been jobs where I’ve a back­ground that should at least get a con­ver­sa­tion and I’ve heard noth­ing. That’s hardly re­mark­able on its own. Hiring is messy, roles get frozen, re­cruiters dis­ap­pear and com­pa­nies re­ject per­fectly good can­di­dates for rea­sons the can­di­date will never know. There is an en­tire web­site called Did They Ghost You?, so we’re not deal­ing with an un­ex­plained phe­nom­e­non.

But Nasdaq told me it was Mac Aviation. DHL ship­ping on be­half of SpaceX asked for a pass­port and told me it was due to a hit against the re­stricted par­ties on the U.S. gov­ern­ment con­sol­i­dated screen­ing list. Apple at least told me I’d matched some­thing, even if it then stopped talk­ing. An ap­pli­cant track­ing sys­tem will tell me noth­ing at all.

Upstream

I asked the Bureau of Industry and Security’s End-User Review Committee to re­view the orig­i­nal Entity List en­try. I’m not sure any­thing will come of it. The nor­mal process is de­signed for a listed per­son ask­ing to be re­moved, which cre­ates an in­ter­est­ing prob­lem here. I’m not the listed Sean Byrne, and the avail­able ev­i­dence sug­gests that per­son may never have ex­isted.

For now the U.S. gov­ern­men­t’s screen­ing data still says Sean Byrne, Cloonmull House, Drumcliffe, County Sligo.

I’ve still never lived in Sligo.

Auto-research with codex: How I achieved a 232x Faster Kernel over baseline with Codex in GPU Mode's qr_v2 problem

sankalp.bearblog.dev

08 Jul, 2026

Table of Contents

Intro Contest in short Problem in­tro

Contest in short

Problem in­tro

Why this prob­lem is auto-re­search-able

Learning Enough to Ask Better Questions

(Optional) Math for QR de­com­po­si­tion: Householder re­flec­tions

Make se­r­ial work small with the help of the blocked Householder al­go­rithm

Other chal­lenges

Codex-maxxing Kernel progress break­throughs Breakthrough ideas

Kernel progress break­throughs

Breakthrough ideas

Introducing idea di­ver­sity to es­cape the lo­cal max­ima

Implementation Hints

What I could have done bet­ter

Conclusion

References

Acknowledgements

Intro

Contest in short

GPU Mode, in col­lab with Core Automation, re­cently hosted an auto-re­search themed con­test. The prob­lem state­ment was to im­ple­ment batched square com­pact-House­holder QR fac­tor­iza­tion aka QR de­com­po­si­tion. I placed 12th out of 183 par­tic­i­pants, end­ing up with a 232x speedup over the base­line so­lu­tion. This post is about how I got there. I will go through my ap­proach, learn­ings, and bot­tle­necks I ran into dur­ing the con­test. It was my first se­ri­ous at­tempt at auto-re­search. Some peo­ple will call this loop en­gi­neer­ing”, and hon­estly that is fine too.

Note that you don’t need to go through the math­e­mat­ics or the prob­lem it­self in de­tail to fol­low most of this blog post. I have fo­cused on my ap­proach while keep­ing the math and the prob­lem it­self sec­ondary as most peo­ple who will read this won’t have par­tic­i­pated in the con­test.

You can check out the full con­test page here: Problem Link and Leaderboard

This con­test was part of GPU Mode’s Linear Algebra Kernels in the Age of Research se­ries.

Problem in­tro

We were given a batch of square FP32 CUDA ma­tri­ces A with shape batch x n x n, and had to re­turn the same com­pact Householder QR rep­re­sen­ta­tion as torch.geqrf(A): an H ma­trix whose up­per tri­an­gle is R and whose lower tri­an­gle stores Householder vec­tors, plus a tau vec­tor of re­flec­tor co­ef­fi­cients. The checker re­built Q with torch.linalg.house­hold­er_prod­uct(H, tau), took R = triu(H), and ver­i­fied:

A≈QR,Q⊤Q≈I,Q⊤A≈R

Among cor­rect sub­mis­sions, the leader­board ranked run­time by geo­met­ric mean across shapes and con­di­tion­ing cases. The im­por­tant sizes were batched square ma­tri­ces like 512 x 512, with larger 1024, 2048, and 4096 cases too. Low-bit FP16, FP8, or NVFP4 was al­lowed in­ter­nally, but re­turned fac­tors still had to sat­isfy FP32-style QR checks.

A tiny 3 x 3 ex­am­ple is:

A=[12−5146167−68−424−41]=[6/7−69/175−58/1753/7158/1756/175−2/76/35−33/35]⏟Q[1421−140175−700035]⏟R

Here Q is or­thog­o­nal, which means its columns are unit-length and per­pen­dic­u­lar to each other, and R is up­per tri­an­gu­lar, which means every­thing be­low the di­ag­o­nal is zero. The con­test was not ask­ing us to print dense Q and R di­rectly; it asked for the com­pact Householder ver­sion that lets the checker re­con­struct Q and read R from the up­per tri­an­gle.

For the 3×3 ex­am­ple above, the very first re­flec­tor maps the first col­umn (12, 6, −4) straight onto (−14, 0, 0) in one shot. The −14 be­comes R11. How that works is in the math sec­tion.

Why this prob­lem is auto-re­search-able

GPU Mode pro­vides par­tic­i­pants with the pop­corn CLI mak­ing it agent-friendly. Agents can use this to test, bench­mark, and sub­mit to the leader­board di­rectly. The checker also pro­vided shape-wise feed­back along with the over­all geo­met­ric mean tim­ing.

Astute ob­servers will no­tice this is an apt setup for writ­ing a loop. Agents yearn for tight feed­back loops. They al­low them to hill-climb to their heart’s con­tent.

GPU Mode con­tests usu­ally give you some way to it­er­ate on ker­nels. Either you sub­mit di­rectly, or a spon­sor like Modal chips in cred­its. Here the or­ga­niz­ers ba­si­cally al­lowed un­lim­ited sub­mis­sions as long as you spaced them out. If you did­n’t, the queues got long and every­body’s runs timed out. At one point the work­space even ran out of Modal cred­its be­cause every­one had been ham­mer­ing sub­mis­sions. It’s a nice way to make learn­ing ac­ces­si­ble.

Over the course of 14 days, I made over 1500 sub­mis­sions.

Learning Enough to Ask Better Questions

I have known the ba­sics of GPU ker­nel op­ti­miza­tion (mainly in Triton with some un­der­stand­ing of CUDA) for a year, but haven’t worked in this do­main pro­fes­sion­ally. What I am try­ing to tell you is that I was an un­der­dog among the peo­ple around me on the leader­board. The per­son just above me on the leader­board (CUDA Colonel) is a prin­ci­pal en­gi­neer at NVIDIA.

Anyway aura farm­ing aside, since I know the ba­sics and had re­cently read about GatedDeltaNet, I was fresh on the gen­eral GPU ker­nel lingo.

The bet­ter you know some­thing, the bet­ter you can prompt the LLMs, be­cause you con­vert un­known un­knowns into known un­knowns.

At the same time, it’s worth not­ing that this con­test was doable with­out do­main knowl­edge - like you prob­a­bly won’t make it to the top 10, but you can get a re­spectable speedup over base­line by just re­ly­ing on your har­ness/​agent loop or what­ever.

My first steps in the con­test were to learn what QR de­com­po­si­tion is and how it can be done. There are a bunch of ways to do it - like Gram-Schmidt and Householder re­flec­tions. The con­test man­dated Householder re­flec­tions. I went back and forth with Claude and watched a few YouTube videos to build in­tu­ition. After my dis­cus­sions with Claude, it was clear that we needed to use the blocked Householder al­go­rithm as the main ar­chi­tec­ture with the trail­ing WY-update. As it turns out, GPT-5.5 also had a good idea about this. QR de­com­po­si­tion is a fairly well known prob­lem.

I found the con­cept in­ter­est­ing as ma­trix de­com­po­si­tions show up in sev­eral mod­ern op­ti­mizer vari­ants for LLM train­ing, es­pe­cially in meth­ods that use ma­trix pre­con­di­tion­ing, such as Shampoo-style op­ti­miz­ers and re­lated ap­proaches. Muon (used by Kimi) is an­other good ex­am­ple: in­stead of treat­ing a weight up­date as one gi­ant flat­tened vec­tor, it keeps the ma­trix struc­ture around and or­thog­o­nal­izes the mo­men­tum up­date, usu­ally through a few Newton-Schulz it­er­a­tions that ap­prox­i­mate the po­lar fac­tor.

(Optional) Math for QR de­com­po­si­tion: Householder re­flec­tions

I rec­om­mend skim­ming through this sec­tion if you are cu­ri­ous about the math oth­er­wise feel free to skip. The only thing to note is that there is a se­quen­tial de­pen­dency in Householder QR which makes it prob­lem­atic to do GEMM. We use blocked Householder to make it more ma­trix-mul­ti­pli­ca­tion shaped.

The con­tract

Quickly re­view­ing the con­tract: in­put is a batch of square FP32 ma­tri­ces A; out­put is the com­pact (H, tau) for­mat that torch.geqrf re­turns. The up­per tri­an­gle of H is R. Below the di­ag­o­nal, H stores the Householder vec­tors, and tau stores one scalar per col­umn. The checker re­builds Q from (H, tau) and ver­i­fies A ≈ QR.

Mirrors

Forget ma­tri­ces for a sec­ond. In a bath­room mir­ror, your re­flec­tion is ex­actly as far be­hind the glass as you are in front of it, straight through.

If x⟂ is the part of x stick­ing out per­pen­dic­u­lar to the glass, re­flec­tion just sub­tracts that part twice:

xre­flected=x−2x⟂

So a Householder re­flec­tion is about find­ing the per­pen­dic­u­lar part and sub­tract­ing it twice.

Storing the mir­ror

A Householder vec­tor is the mir­ror, stored com­pactly. In code, we don’t carry around the whole mir­ror plane. We store one vec­tor v stick­ing straight out of it. The mir­ror is every­thing per­pen­dic­u­lar to v, and the re­flec­tion moves along v.

The per­pen­dic­u­lar part is just the shadow of x along v, which is v⊤xv⊤v copies of v. Plug that into the sub­trac­tion above:

ℋx=x−τv(v⊤x),τ=2v⊤v

So tau is just 2v⊤v: the fac­tor of 2 and the length of v bun­dled into one pre­com­puted num­ber. v picks the mir­ror, tau scales the up­date. (I’ll write the math­e­mat­i­cal re­flec­tor as ℋj and re­serve H for the com­pact out­put ma­trix.)

Householder re­flec­tion in 2D tau = 0.00

Why QR cares about mir­rors

QR wants to turn A into an up­per-tri­an­gu­lar ma­trix R. Column 1 should be­come some­thing like (*, 0, 0), col­umn 2 should have ze­ros be­low row 2, and so on.

A Householder mir­ror is use­ful be­cause it can do that to a col­umn in one shot. Take the first col­umn of the 3×3 ex­am­ple above: (12, 6, -4). We want to send it to the x-axis so the lower en­tries be­come zero. A re­flec­tion can only change di­rec­tion, not length, so the tar­get must also have length 14. One valid tar­get is (-14, 0, 0). After that re­flec­tion, the 6 and -4 en­tries are gone, which is ex­actly what we wanted.

How do we find the mir­ror? It sits halfway be­tween the col­umn and its tar­get, so v, the vec­tor pok­ing through the mir­ror, is just the col­umn mi­nus its tar­get:

v = (12, 6, -4) - (-14, 0, 0) = (26, 6, -4)

The Householder up­date

Then com­pute tau = 2 / (vᵀv). The re­flec­tor it­self is:

ℋ=I−τvv⊤,τ=2v⊤v

ℋx=(I−τvv⊤)x

ℋx=x−τv(v⊤x)

ℋAactive=(I−τvv⊤)Aactive

ℋAactive=Aactive−τv(v⊤Aactive)

This turns the cur­rent col­umn into (-14, 0, 0) and rewrites the other columns con­sis­tently, so the next re­flec­tor is built from the up­dated ma­trix.

We keep do­ing this once per col­umn. In rough no­ta­tion, the re­peated up­dates look like:

A(1)=A(0)−τ1v1(v1⊤A(0))

A(2)=A(1)−τ2v2(v2⊤A(1))

A(3)=A(2)−τ3v3(v3⊤A(2))

R=A(n)

Each line uses the ma­trix pro­duced by the pre­vi­ous line. Each re­flec­tor ze­roes out every­thing be­low the di­ag­o­nal of its col­umn with­out dis­turb­ing the columns al­ready fin­ished. After the last one, A has walked down to an up­per-tri­an­gu­lar R:

A=ℋ1ℋ2⋯ℋn⏟QR

Mirrors don’t change lengths or an­gles, so each ℋj is or­thog­o­nal, and so is their prod­uct Q. That’s where the or­thog­o­nal­ity the checker ver­i­fies comes from for free.

What’s the com­pact for­mat

Once col­umn j is processed, every­thing be­low its di­ag­o­nal is dead space. geqrf reuses those slots to stash the tail of v_j (the lead­ing 1 is im­plicit). On and above the di­ag­o­nal you’re look­ing at R; be­low it, the re­flec­tors; and tau rides along as a sep­a­rate vec­tor. That’s why the checker needs both H and tau to re­build Q.

Compact geqrf stor­age (n = 6) hover a col­umn

H  (upper tri­an­gle = R, be­low di­ag­o­nal = re­flec­tor tails)

tau  (one scalar per re­flec­tor)

r en­tries of R tail of vj τj

Make se­r­ial work small with the help of the blocked Householder al­go­rithm

Householder QR ze­roes out A be­low the di­ag­o­nal one col­umn at a time. Each step builds a re­flec­tor from the cur­rent col­umn and ap­plies it to every­thing on the right. The prob­lem is re­flec­tor j+1 is built from the ma­trix af­ter re­flec­tor j has al­ready hit it. So you can’t re­order the steps and you can’t fuse them. It’s se­r­ial, and the se­r­ial ma­trix-vec­tor work runs in the slow vec­tor lanes of the SM while the ten­sor cores just sit there idle.

Householder QR, one re­flec­tor at a time (n = 5) math view: ze­ros ap­pear, trail­ing block gets rewrit­ten

orig­i­nal a fi­nal en­try of R rewrit­ten by this re­flec­tor 0 zeroed (vj gets stashed here)

The clas­sic fix is the blocked al­go­rithm. You pick a nar­row panel of b columns (say 32 or 64) and do all the se­r­ial work in­side it. That’s fine, be­cause the panel is only b columns wide, so it stays cheap. Then, in­stead of ap­ply­ing the pan­el’s b re­flec­tors to the rest of the ma­trix one at a time, you com­press them into a sin­gle rank-b up­date (the WY rep­re­sen­ta­tion”) and hit the en­tire trail­ing block in one shot with three back-to-back ma­trix mul­ti­plies. The se­r­ial work stays con­fined to the panel, and every­thing else turns into GEMMs, which is ex­actly the shape the ten­sor cores want.

Concretely, the WY rep­re­sen­ta­tion col­lapses a pan­el’s b re­flec­tors into a sin­gle rank-b up­date. Stack the pan­el’s Householder vec­tors as columns of V=[v1,v2,…,vb], build a small b×b up­per-tri­an­gu­lar T, then:

ℋ1ℋ2⋯ℋb=I−VTV⊤

and the trail­ing-block up­date be­comes three GEMM-shaped steps:

W=V⊤Atrail

Z=T⊤W

Atrail←Atrail−VZ

Blocked Householder (n = 6, panel width b = 2) con­fine the se­r­ial work, GEMM the rest

Eigendrum - draw a shape and hear it as a real drum

eigendrum.com

how it works

A drum­head clamped at its rim can only vi­brate in cer­tain shapes, at cer­tain fre­quen­cies. Those shapes and fre­quen­cies are the so­lu­tions of

−∇²u = λu  inside the shape,  u = 0 on the edge

Each so­lu­tion u is a mode, a stand­ing wave, and each λ gives a fre­quency pro­por­tional to √λ. This is an eigen­value prob­lem, and for al­most every shape it has no for­mula. So Eigendrum solves it nu­mer­i­cally: it cov­ers your shape with a mesh of tri­an­gles, builds the fi­nite el­e­ment stiff­ness and mass ma­tri­ces, and finds the small­est eigen­val­ues of Kφ = λMφ.

why you can trust the num­bers

A few shapes have spec­tra that can be writ­ten down ex­actly, and the solver is tested against them on every change. A cir­cle’s fre­quen­cies are the ze­ros of Bessel func­tions; a rec­tan­gle’s are π²(m²/​a² + n²/​b²). The solver re­pro­duces both to bet­ter than a tenth of a per­cent, and be­cause a con­form­ing fi­nite el­e­ment method min­imises en­ergy over a re­stricted space, its an­swers are guar­an­teed slight over­es­ti­mates, never un­der. The mea­sured er­ror is in the num­bers”.

where you strike it mat­ters

Striking a spot dri­ves each mode in pro­por­tion to how much that mode moves there. Hit a line where a mode stands still and you can­not ex­cite it at all. That was not pro­grammed in; it falls out of pro­ject­ing the mal­let onto the modes.

So a strike is never one mode: it is every mode at once, in a mix­ture set by where your mal­let landed. The rules along the mode list are that mix­ture, and the modes marked with a square were the ones your mal­let could not reach. Pressing a row in­stead plays that sin­gle mode alone - some­thing no mal­let can do, and the only way to hear what one fre­quency of a shape ac­tu­ally sounds like.

drums from equa­tions

Besides trac­ing an out­line you can write one. r(t) gives the ra­dius as t sweeps one full turn, so 1 + 0.3cos(5t) is a five-lobed flower; a para­met­ric x(t), y(t) pair reaches the closed curves po­lar can­not, like a nephroid or an egg. This is not a short­cut for draw­ing. It reaches shapes no hand traces ac­cu­rately - eleven even lobes, a super­el­lipse part­way be­tween a cir­cle and a square - and it makes a shape some­thing you vary: change one num­ber and hear what moved.

A writ­ten shape trav­els as its own text. The link for a for­mula holds the for­mula, so it is some­thing you can read and re­type rather than a few hun­dred char­ac­ters of en­coded out­line, and edit­ing it in the ad­dress bar works. Anything too thin to mesh hon­estly is re­fused rather than an­swered, be­cause a sliver would still re­turn num­bers and they would be wrong.

can one hear the shape of a drum?

Mark Kac asked ex­actly that in 1966. In 1992 Carolyn Gordon, David Webb and Scott Wolpert an­swered no, by build­ing two dif­fer­ent shapes with iden­ti­cal spec­tra. Both are in the form list as Kac drum I and II. Each is made from the same seven tri­an­gles, re­arranged. They en­close the same area and the same perime­ter, and every fre­quency matches. Switch be­tween them and lis­ten: the out­lines are plainly dif­fer­ent and the sound is not.

what is a mod­el­ling choice

The fre­quency ra­tios, the mode shapes and the pitch of the fun­da­men­tal are physics, fixed en­tirely by the out­line. What is not in the out­line is the wave speed, which is ten­sion and den­sity: the pitch slider sets that by nam­ing the note a cir­cle of this area would sound, and each shape then lands above the ref­er­ence by its own amount. Every shape is scaled to the same area be­fore solv­ing, so that off­set is shape and not size - about six semi­tones across the built-in shapes, with the cir­cle low­est, which is Faber-Krahn rather than a choice. How fast each over­tone fades is ma­te­r­ial and air, so that stays a slider rather than a silent as­sump­tion.

The mal­let is mod­elled too. Its width is a slider; its con­tact time is fixed at a few mil­lisec­onds, be­cause no real beater is in­stan­ta­neous and one that was would drive every mode equally hard. Both de­cide how much of a mode a strike can reach, and nei­ther can move a mod­e’s fre­quency. Damping is Rayleigh damp­ing, so loss rises with the square of fre­quency: the high over­tones die away first, which is why a drum dark­ens as it rings.

where it lives, and how to reach me

Eigendrum is hosted at eigen­drum.com. That is the ad­dress to link to and to cite; the older base­lashraf81.github.io/​eigen­drum is a mir­ror that now redi­rects there.

For ad­ver­tis­ing or part­ner­ship en­quiries, write to u2679054@uel.ac.uk. For any­thing wrong with the maths or the in­ter­face, an is­sue on the repos­i­tory is bet­ter, be­cause then the fix is pub­lic.

colophon

No build step and no ap­pli­ca­tion back­end: the mesh, the solve and the au­dio all run on your own ma­chine. The de­ployed site uses Vercel Analytics and Google Analytics; it car­ries no ad­ver­tis­ing net­work and no con­sent ban­ner. Support to­ward the do­main and host­ing is vol­un­tary, via the link above. The shape you draw lives in the ad­dress bar af­ter the #, which browsers never send to a server, and an­a­lyt­ics is con­fig­ured not to record it. Details in the pri­vacy no­tice. Set in Jost* by in­de­struc­tible type*. After Kac, Can One Hear the Shape of a Drum? (1966); Gordon, Webb and Wolpert (1992); and Driscoll, Eigenmodes of Isospectral Drums (1997), whose co­or­di­nates the two Kac drums use.

Source, in­clud­ing the solver and the tests that check it against the closed-form spec­tra: github.com/​Base­lAshraf81/​eigen­drum

Free to use, with no ac­count and noth­ing to in­stall. If you would like to put some­thing to­wards it, or would rather it were not ad-sup­ported: ko-fi.com/​base­lashraf

earthquake.usgs.gov

The Earthquake Event Page ap­pli­ca­tion sup­ports most re­cent browsers, view sup­ported browsers. Or, try our Real-time Notifications, Feeds, and Web Services.

Just a moment...

alz-journals.onlinelibrary.wiley.com

Working With AI Feels More Like Leadership Than Coding

allen.bargi.org

allen@bargi:~/​notes$ cat work­ing-with-ai.md

For most of my ca­reer, code gave me cer­tainty. A pro­gram did what its in­struc­tions told it to do. If the same in­put pro­duced a dif­fer­ent re­sult, we called it a bug.

People were never like that. As a leader, I can ex­plain a task and get ex­actly what I asked for. I can also get some­thing bet­ter be­cause a col­league un­der­stood the in­tent be­hind the re­quest. Sometimes the re­sult shows that I was not as clear as I thought.

Working with AI feels closer to the sec­ond ex­pe­ri­ence.

AI runs on soft­ware, but work­ing with it is not fully pre­dictable. The same re­quest can pro­duce a dif­fer­ent an­swer. It can make a use­ful con­nec­tion, miss an ob­vi­ous point, or sur­prise me with an ap­proach I had not con­sid­ered.

This is frus­trat­ing when I treat AI like a com­piler. It be­comes more use­ful when I treat the in­ter­ac­tion as a form of col­lab­o­ra­tion.

That does not make AI a per­son. It has no lived ex­pe­ri­ence, ac­count­abil­ity, or hu­man judg­ment. The com­par­i­son is about how we work. Good lead­ers do more than is­sue in­struc­tions. They share con­text, ex­plain the de­sired out­come, set bound­aries, and re­spond to what comes back.

The same habits im­prove my work with AI. A good prompt helps, but a shared work­ing con­text helps more. Examples, cor­rec­tions, and reusable in­struc­tions re­duce mis­un­der­stand­ings. Over time, the sys­tem be­comes bet­ter aligned with how I think and what I need from it.

The in­vest­ment is not in pre­tend­ing that AI is hu­man. It is in be­com­ing bet­ter at ex­press­ing in­tent.

We spent years learn­ing how to tell com­put­ers ex­actly what to do. Now we also need to ex­plain why the work mat­ters, what a good re­sult looks like, and where judg­ment is needed.

For me, that is the shift. AI is mak­ing soft­ware work less like is­su­ing com­mands to a ma­chine and more like lead­ing through a con­ver­sa­tion. The tech­nol­ogy is new. The lead­er­ship skills are not.

AI Isn’t Outthinking Mathematicians. It’s Out-Remembering Them.

davidepiffer.com

When an AI sys­tem solves a dif­fi­cult math­e­mat­i­cal prob­lem, the usual ex­pla­na­tion is that it has be­come more in­tel­li­gent.

Perhaps it has ab­sorbed mil­lions of math­e­mat­i­cal ex­am­ples. Perhaps re­in­force­ment learn­ing has taught it bet­ter rea­son­ing strate­gies. Perhaps it is be­gin­ning to de­velop some­thing re­sem­bling gen­uine math­e­mat­i­cal in­tu­ition.

All of these ex­pla­na­tions may con­tain some truth. But they over­look a sim­pler pos­si­bil­ity:

AI has ac­cess to a vastly larger work­ing mem­ory than the hu­man brain.

Or, more pre­cisely, it has ac­cess to an enor­mous ex­ter­nal sym­bolic work­space that per­forms many of the func­tions that work­ing mem­ory per­forms in hu­mans.

This dif­fer­ence may be es­pe­cially im­por­tant in math­e­mat­ics.

A hu­man math­e­mati­cian can hold only a small num­ber of un­fa­mil­iar el­e­ments in mind si­mul­ta­ne­ously. An AI model can keep the en­tire prob­lem state­ment, hun­dreds of in­ter­me­di­ate equa­tions, sev­eral aban­doned ap­proaches, de­f­i­n­i­tions, con­straints and ear­lier con­clu­sions in­side its con­text win­dow.

We nor­mally in­ter­pret the re­sult­ing per­for­mance as ev­i­dence of su­pe­rior rea­son­ing. But some of it may in­stead re­flect the re­moval of one of the most im­por­tant bi­o­log­i­cal lim­its on hu­man rea­son­ing: our ex­tremely re­stricted work­ing-mem­ory ca­pac­ity.

Working mem­ory is the men­tal sys­tem that al­lows us to hold and ma­nip­u­late in­for­ma­tion over short pe­ri­ods.

When solv­ing an equa­tion, you must re­mem­ber what each vari­able rep­re­sents, which op­er­a­tions have al­ready been per­formed and what the cur­rent goal is. During a proof, you may need to keep track of as­sump­tions, in­ter­me­di­ate lem­mas, ex­cep­tions and mul­ti­ple pos­si­ble cases.

Human work­ing mem­ory is re­mark­ably lim­ited.

Its ex­act ca­pac­ity de­pends on the task and on how in­for­ma­tion is or­ga­nized, but the gen­eral lim­i­ta­tion is ob­vi­ous from every­day ex­pe­ri­ence. Try mul­ti­ply­ing two three-digit num­bers in your head. The un­der­ly­ing op­er­a­tions are sim­ple. The dif­fi­culty comes largely from hav­ing to pre­serve par­tial re­sults while per­form­ing ad­di­tional cal­cu­la­tions.

Writing the num­bers down trans­forms the prob­lem.

Paper does not make you more in­tel­li­gent. It ex­pands your ef­fec­tive work­ing mem­ory.

The same prin­ci­ple ap­plies at higher lev­els of math­e­mat­ics. A math­e­mati­cian uses no­ta­tion, scratch pa­per, di­a­grams and pre­vi­ously writ­ten lem­mas not merely to com­mu­ni­cate the so­lu­tion, but to make the rea­son­ing cog­ni­tively pos­si­ble.

Experts com­pen­sate through chunking.” A novice sees a long se­quence of sym­bols. An ex­pert rec­og­nizes a fa­mil­iar struc­ture and treats it as a sin­gle con­cep­tual ob­ject. This al­lows far more in­for­ma­tion to fit in­side the same bi­o­log­i­cal work­ing-mem­ory limit.

But chunk­ing does not elim­i­nate the limit. It merely com­presses the in­for­ma­tion.

An AI model faces a very dif­fer­ent con­straint.

The im­por­tance of work­ing mem­ory for math­e­mat­ics is not merely the­o­ret­i­cal. It is vis­i­ble in the dif­fer­ences be­tween hu­man be­ings.

Working mem­ory is strongly re­lated to gen­eral in­tel­li­gence, which raises an ob­vi­ous ques­tion: does it in­de­pen­dently pre­dict math­e­mat­i­cal per­for­mance, or is it merely an­other im­per­fect mea­sure of IQ?

Several stud­ies sug­gest that it con­tributes some­thing be­yond con­ven­tional in­tel­li­gence mea­sures. Alloway and Passolunghi (2011), for ex­am­ple, ex­am­ined work­ing mem­ory, ver­bal abil­ity and math­e­mat­i­cal skills in chil­dren. They found that work­ing-mem­ory mea­sures made a dis­tinct con­tri­bu­tion to math­e­mat­i­cal per­for­mance rather than sim­ply re­pro­duc­ing the as­so­ci­a­tion be­tween math­e­mat­ics and gen­eral ver­bal abil­ity.

In a sep­a­rate six-year lon­gi­tu­di­nal study, Alloway and Alloway (2010) mea­sured chil­dren at age five and then ex­am­ined their aca­d­e­mic achieve­ment six years later. Early work­ing-mem­ory per­for­mance pre­dicted later lit­er­acy and nu­mer­acy even af­ter IQ was in­cluded in the analy­sis. Indeed, work­ing mem­ory was a stronger pre­dic­tor of the later aca­d­e­mic out­comes than the IQ mea­sure used in the study.

Blankenship and col­leagues (2015) sim­i­larly re­ported that work­ing mem­ory ex­plained unique vari­a­tion in math­e­mat­i­cal flu­ency and cal­cu­la­tion af­ter sta­tis­ti­cally con­trol­ling for IQ and age. A large meta-analy­sis by Friso-van den Bos and col­leagues (2013) also found a con­sis­tent re­la­tion­ship be­tween work­ing mem­ory and math­e­mat­ics across pri­mary-school stud­ies, al­though the strength of the re­la­tion­ship var­ied ac­cord­ing to the type of work­ing-mem­ory and math­e­mat­i­cal task be­ing mea­sured.

These find­ings should not be ex­ag­ger­ated. Working mem­ory and in­tel­li­gence over­lap sub­stan­tially, and sta­tis­ti­cal con­trol can­not per­fectly iso­late them as in­de­pen­dent psy­cho­log­i­cal mech­a­nisms. Nor does the ev­i­dence im­ply that com­mer­cially train­ing work­ing mem­ory will nec­es­sar­ily pro­duce large im­prove­ments in in­tel­li­gence or math­e­mat­ics.

The nar­rower con­clu­sion is nev­er­the­less im­por­tant: among chil­dren with sim­i­lar mea­sured in­tel­li­gence, dif­fer­ences in the abil­ity to hold, up­date and ma­nip­u­late in­for­ma­tion still pre­dict dif­fer­ences in math­e­mat­i­cal per­for­mance.

This pro­vides a cru­cial clue for un­der­stand­ing AI. If hu­man math­e­mat­i­cal per­for­mance is partly capped by a work­ing-mem­ory bot­tle­neck, then giv­ing a ma­chine an enor­mous sym­bolic work­space changes the na­ture of the con­test. The ma­chine may ap­pear more math­e­mat­i­cally in­tel­li­gent partly be­cause it is much less con­strained by a cog­ni­tive lim­i­ta­tion that sup­presses hu­man per­for­mance.

A mod­ern lan­guage model can process an enor­mous se­quence of to­kens at once. This se­quence may in­clude the orig­i­nal ques­tion, de­f­i­n­i­tions, ex­am­ples, in­ter­me­di­ate cal­cu­la­tions and the mod­el’s own ear­lier rea­son­ing.

The con­text win­dow is not iden­ti­cal to hu­man work­ing mem­ory. It is bet­ter un­der­stood as a gi­gan­tic ex­ter­nal note­book com­bined with an im­per­fect sys­tem for search­ing and us­ing what has been writ­ten in it.

This dis­tinc­tion mat­ters.

Humans pos­sess a form of ac­tive in­ter­nal mem­ory. We can silently choose a num­ber, hold it in mind, trans­form it and re­place it with a new value with­out say­ing or writ­ing any­thing.

Standard lan­guage mod­els are much weaker at main­tain­ing this kind of pri­vate, con­tin­u­ously up­dated men­tal state. Their most sta­ble form of mem­ory is usu­ally the se­quence of to­kens that has al­ready been gen­er­ated.

If the model writes:

x=6

and later writes:

x+3=9,

those state­ments re­main in­side the con­text. The model can at­tend to them again when gen­er­at­ing the next step.

Its rea­son­ing is there­fore of­ten ex­ter­nal­ized. The text is not merely a re­port of a com­pleted thought process. The text is part of the mech­a­nism by which the rea­son­ing oc­curs.

Humans do some­thing sim­i­lar when us­ing scratch pa­per. The ma­jor dif­fer­ence is scale.

An un­aided hu­man may strug­gle to keep five un­fa­mil­iar con­di­tions ac­tive si­mul­ta­ne­ously. An AI can pre­serve dozens or hun­dreds of them in ex­plicit form.

This does not mean that every item in a long con­text is re­trieved per­fectly. Models can over­look rel­e­vant in­for­ma­tion, be­come dis­tracted or lose track of de­tails. Advertised con­text length is not the same as per­fectly us­able mem­ory.

Nevertheless, the dif­fer­ence in po­ten­tial ca­pac­ity is enor­mous.

The con­text-win­dow ad­van­tage is not equally use­ful in every kind of rea­son­ing.

It mat­ters es­pe­cially for math­e­mat­ics be­cause math­e­mat­i­cal rea­son­ing can be trans­lated un­usu­ally well into ex­plicit sym­bols.

Almost every rel­e­vant el­e­ment of a math­e­mat­i­cal prob­lem can be writ­ten down:

the as­sump­tions;

the as­sump­tions;

the de­f­i­n­i­tions;

the de­f­i­n­i­tions;

the known equa­tions;

the known equa­tions;

the cur­rent ob­jec­tive;

the cur­rent ob­jec­tive;

the re­sults al­ready proved;

the re­sults al­ready proved;

the cases that have been elim­i­nated;

the cases that have been elim­i­nated;

the con­di­tions un­der which each step re­mains valid.

the con­di­tions un­der which each step re­mains valid.

Once writ­ten, this in­for­ma­tion re­mains sta­ble.

If (x) is de­fined as an in­te­ger at the be­gin­ning of a proof, it re­mains an in­te­ger un­less the proof ex­plic­itly changes the de­f­i­n­i­tion. A strict in­equal­ity does not grad­u­ally be­come a non-strict in­equal­ity be­cause of changes in mood, con­text or in­ter­pre­ta­tion.

Mathematical sym­bols are de­signed to re­duce am­bi­gu­ity.

This makes math­e­mat­ics al­most per­fectly suited to an in­tel­li­gence that op­er­ates through a large tex­tual work­space.

Consider a prob­lem re­quir­ing the solver to re­mem­ber that:

n is odd;

n is odd;

p is prime;

p is prime;

x≠0

x≠0

and that one branch of the ar­gu­ment has al­ready pro­duced a con­tra­dic­tion.

A hu­man may un­der­stand the ba­sic strat­egy but di­vide by x be­fore es­tab­lish­ing that x≠0. The er­ror is not nec­es­sar­ily caused by a lack of in­tel­li­gence. It may be a fail­ure of book­keep­ing.

An AI can re­state the ac­tive con­straints at each stage:

We are work­ing un­der the as­sump­tions that (n) is odd, (p) is prime and (x≠0).

We are work­ing un­der the as­sump­tions that (n) is odd, (p) is prime and (x≠0).

The con­text be­comes a ledger of the rea­son­ing state.

Many dif­fi­cult math­e­mat­i­cal prob­lems con­tain a pro­found in­sight some­where near the be­gin­ning, but they also con­tain a large amount of less glam­orous work af­ter­ward: ex­pand­ing ex­pres­sions, check­ing cases, car­ry­ing con­di­tions through trans­for­ma­tions and mak­ing sure that the fi­nal con­clu­sion is com­pat­i­ble with every ear­lier as­sump­tion.

A ma­chine does not need to pos­sess deeper in­sight than a hu­man to gain an ad­van­tage here. It may sim­ply be bet­ter equipped to pre­serve the en­tire state of the prob­lem while com­plet­ing a long se­quence of op­er­a­tions.

Mathematics is highly com­po­si­tional.

A proof can of­ten be rep­re­sented as:

A → B → C → D

If each step is valid and the chain is pre­served ac­cu­rately, the con­clu­sion fol­lows.

A large work­ing space al­lows the model to con­struct much longer chains be­fore los­ing the thread.

This is im­por­tant be­cause the dif­fi­culty of a prob­lem does not de­pend only on the dif­fi­culty of each in­di­vid­ual step. It also de­pends on how many steps must be co­or­di­nated.

A per­son may be per­fectly ca­pa­ble of un­der­stand­ing every lo­cal in­fer­ence in a 100-step ar­gu­ment while still be­ing un­able to gen­er­ate the en­tire ar­gu­ment un­aided. The prob­lem ex­ceeds the per­son’s abil­ity to main­tain the global struc­ture.

AI can po­ten­tially com­pen­sate by writ­ing down nearly every­thing.

This may ex­plain why ad­di­tional thinking time” of­ten im­proves model per­for­mance. More com­pu­ta­tion al­lows the sys­tem to pro­duce more in­ter­me­di­ate states, ex­am­ine al­ter­na­tive branches and pre­serve par­tial con­clu­sions.

What looks like deeper thought may some­times be broader search con­ducted in­side a much larger note­book.

Now com­pare a math­e­mat­i­cal prob­lem with a so­cial ques­tion:

Why has Maria sud­denly stopped re­ply­ing to my mes­sages?

Why has Maria sud­denly stopped re­ply­ing to my mes­sages?

A larger con­text win­dow might al­low an AI to ex­am­ine years of cor­re­spon­dence. It could iden­tify changes in tone, tim­ing and vo­cab­u­lary.

But the de­ci­sive in­for­ma­tion may still be miss­ing.

Perhaps Maria is an­gry. Perhaps she is busy. Perhaps she is ill. Perhaps she has lost her phone. Perhaps she is avoid­ing an un­re­lated prob­lem.

No amount of mem­ory can re­trieve facts that were never ob­served.

The chal­lenge is not sim­ply to pre­serve a known set of premises and de­rive their con­se­quences. It is to rea­son un­der un­cer­tainty about hid­den causes.

Informal rea­son­ing also de­pends heav­ily on con­cepts whose mean­ings are un­sta­ble.

Words such as fair,” successful,” responsible,” harmful” or intelligent” do not pos­sess the ex­act­ness of math­e­mat­i­cal vari­ables. Their mean­ing de­pends on cul­ture, goals and con­text.

A model can re­mem­ber every sen­tence in a dis­cus­sion while still mis­un­der­stand­ing what the par­tic­i­pants mean.

The same ap­plies to po­lit­i­cal analy­sis, his­tor­i­cal in­ter­pre­ta­tion, busi­ness strat­egy and psy­cho­log­i­cal judg­ment. In these do­mains, the cen­tral prob­lem is of­ten not work­ing-mem­ory ca­pac­ity. It is iden­ti­fy­ing the cor­rect causal model when the ev­i­dence is in­com­plete and am­bigu­ous.

A larger note­book helps, but it does not solve the fun­da­men­tal prob­lem.

Eigendrum - draw a shape and hear it as a real drum

baselashraf81.github.io

how it works

A drum­head clamped at its rim can only vi­brate in cer­tain shapes, at cer­tain fre­quen­cies. Those shapes and fre­quen­cies are the so­lu­tions of

−∇²u = λu  inside the shape,  u = 0 on the edge

Each so­lu­tion u is a mode, a stand­ing wave, and each λ gives a fre­quency pro­por­tional to √λ. This is an eigen­value prob­lem, and for al­most every shape it has no for­mula. So Eigendrum solves it nu­mer­i­cally: it cov­ers your shape with a mesh of tri­an­gles, builds the fi­nite el­e­ment stiff­ness and mass ma­tri­ces, and finds the small­est eigen­val­ues of Kφ = λMφ.

why you can trust the num­bers

A few shapes have spec­tra that can be writ­ten down ex­actly, and the solver is tested against them on every change. A cir­cle’s fre­quen­cies are the ze­ros of Bessel func­tions; a rec­tan­gle’s are π²(m²/​a² + n²/​b²). The solver re­pro­duces both to bet­ter than a tenth of a per­cent, and be­cause a con­form­ing fi­nite el­e­ment method min­imises en­ergy over a re­stricted space, its an­swers are guar­an­teed slight over­es­ti­mates, never un­der. The mea­sured er­ror is in the num­bers”.

where you strike it mat­ters

Striking a spot dri­ves each mode in pro­por­tion to how much that mode moves there. Hit a line where a mode stands still and you can­not ex­cite it at all. That was not pro­grammed in; it falls out of pro­ject­ing the mal­let onto the modes.

So a strike is never one mode: it is every mode at once, in a mix­ture set by where your mal­let landed. The rules along the mode list are that mix­ture, and the modes marked with a square were the ones your mal­let could not reach. Pressing a row in­stead plays that sin­gle mode alone - some­thing no mal­let can do, and the only way to hear what one fre­quency of a shape ac­tu­ally sounds like.

drums from equa­tions

Besides trac­ing an out­line you can write one. r(t) gives the ra­dius as t sweeps one full turn, so 1 + 0.3cos(5t) is a five-lobed flower; a para­met­ric x(t), y(t) pair reaches the closed curves po­lar can­not, like a nephroid or an egg. This is not a short­cut for draw­ing. It reaches shapes no hand traces ac­cu­rately - eleven even lobes, a super­el­lipse part­way be­tween a cir­cle and a square - and it makes a shape some­thing you vary: change one num­ber and hear what moved.

A writ­ten shape trav­els as its own text. The link for a for­mula holds the for­mula, so it is some­thing you can read and re­type rather than a few hun­dred char­ac­ters of en­coded out­line, and edit­ing it in the ad­dress bar works. Anything too thin to mesh hon­estly is re­fused rather than an­swered, be­cause a sliver would still re­turn num­bers and they would be wrong.

can one hear the shape of a drum?

Mark Kac asked ex­actly that in 1966. In 1992 Carolyn Gordon, David Webb and Scott Wolpert an­swered no, by build­ing two dif­fer­ent shapes with iden­ti­cal spec­tra. Both are in the form list as Kac drum I and II. Each is made from the same seven tri­an­gles, re­arranged. They en­close the same area and the same perime­ter, and every fre­quency matches. Switch be­tween them and lis­ten: the out­lines are plainly dif­fer­ent and the sound is not.

what is a mod­el­ling choice

The fre­quency ra­tios, the mode shapes and the pitch of the fun­da­men­tal are physics, fixed en­tirely by the out­line. What is not in the out­line is the wave speed, which is ten­sion and den­sity: the pitch slider sets that by nam­ing the note a cir­cle of this area would sound, and each shape then lands above the ref­er­ence by its own amount. Every shape is scaled to the same area be­fore solv­ing, so that off­set is shape and not size - about six semi­tones across the built-in shapes, with the cir­cle low­est, which is Faber-Krahn rather than a choice. How fast each over­tone fades is ma­te­r­ial and air, so that stays a slider rather than a silent as­sump­tion.

The mal­let is mod­elled too. Its width is a slider; its con­tact time is fixed at a few mil­lisec­onds, be­cause no real beater is in­stan­ta­neous and one that was would drive every mode equally hard. Both de­cide how much of a mode a strike can reach, and nei­ther can move a mod­e’s fre­quency. Damping is Rayleigh damp­ing, so loss rises with the square of fre­quency: the high over­tones die away first, which is why a drum dark­ens as it rings.

where it lives, and how to reach me

Eigendrum is hosted at eigen­drum.com. That is the ad­dress to link to and to cite; the older base­lashraf81.github.io/​eigen­drum is a mir­ror that now redi­rects there.

For ad­ver­tis­ing or part­ner­ship en­quiries, write to u2679054@uel.ac.uk. For any­thing wrong with the maths or the in­ter­face, an is­sue on the repos­i­tory is bet­ter, be­cause then the fix is pub­lic.

colophon

No build step and no ap­pli­ca­tion back­end: the mesh, the solve and the au­dio all run on your own ma­chine. The de­ployed site uses Vercel Analytics and Google Analytics; it car­ries no ad­ver­tis­ing net­work and no con­sent ban­ner. Support to­ward the do­main and host­ing is vol­un­tary, via the link above. The shape you draw lives in the ad­dress bar af­ter the #, which browsers never send to a server, and an­a­lyt­ics is con­fig­ured not to record it. Details in the pri­vacy no­tice. Set in Jost* by in­de­struc­tible type*. After Kac, Can One Hear the Shape of a Drum? (1966); Gordon, Webb and Wolpert (1992); and Driscoll, Eigenmodes of Isospectral Drums (1997), whose co­or­di­nates the two Kac drums use.

Source, in­clud­ing the solver and the tests that check it against the closed-form spec­tra: github.com/​Base­lAshraf81/​eigen­drum

Free to use, with no ac­count and noth­ing to in­stall. If you would like to put some­thing to­wards it, or would rather it were not ad-sup­ported: ko-fi.com/​base­lashraf

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.