10 interesting stories served every morning and every evening.

Don't be a meat proxy

gruhn.me

Aug 03, 2026

Too of­ten I ask a ques­tion in Slack or leave feed­back un­der a merge/​pull re­quest or ar­gue with friends in a WhatsApp group and get back:

Claude said: [giant re­sponse ver­ba­tim]

Claude said: [giant re­sponse ver­ba­tim]

Please don’t do this. I mean, I’ve done this. But I’ve been on the re­ceiv­ing end too many times now. This is not adding value. I can talk to Claude my­self. It’s go­ing to be faster and I get to con­trol the con­text. I don’t need a meat proxy in be­tween.

Reading AI out­put is ex­tra ef­fort. It’s ver­bose, fre­quently con­tains all too plau­si­ble non­sense, and is in­creas­ingly jar­gon dense. I re­cently got this sen­tence from Claude:

NATS con­trol-plane events: stream leader elec­tion / R3 quo­rum re-form dur­ing pod churn.

NATS con­trol-plane events: stream leader elec­tion / R3 quo­rum re-form dur­ing pod churn.

Jesus. I had to lookup al­most every word to make sense of this.

By all means, prompt AI. But don’t just re­lay the out­put. Read it, un­der­stand it, val­i­date it, and then write a re­sponse in your own words (a de­cent cer­tifi­cate that you’ve done the prior steps). Making that ef­fort is value you can add.

Take code re­view in par­tic­u­lar. Shipping some code can be done with close to zero ef­fort now: Copy/paste the ticket de­scrip­tion into Claude Code. Don’t look at the code or read what Claude has writ­ten. If there’s any feed­back from re­view­ers, copy/​paste that into Claude Code as well. If nec­es­sary, it­er­ate.

That works. But who has done the im­ple­men­ta­tion? The re­view­ers did, us­ing Claude Code, and you as a meat proxy.

In Memory of My Wife, Elise Cawley (1961–2026), with Thanks for 36 Wonderful Years

writings.stephenwolfram.com

Something ter­ri­ble just hap­pened. My wife, Elise Cawley, was re­cov­er­ing from heart surgery and had just at­tended vir­tu­ally a cel­e­bra­tion for one of our chil­dren when she had a freak, vast car­dio­vas­cu­lar event—and died in­stantly. We had been to­gether for 36 years. The pic­ture above was taken just hours be­fore she died.

In all the writ­ing and pub­lic speak­ing I have done, I have cho­sen, as a mat­ter of pri­vacy, to say lit­tle about my fam­ily. But now that Elise is gone I can­not re­strain my­self from telling the world some­thing about the re­mark­able per­son with whom I shared the past 36 years of my life.

It was September 17, 1990, and I was in New York City. The event I had been at­tend­ing fin­ished a lit­tle early, so I de­cided to drop in on a friend of mine. And there, sit­ting styl­ishly on the floor, was Elise, a some­what bash­ful but ob­vi­ously bril­liant young pure math­e­mati­cian. It took me a cou­ple of weeks to call her. But from that mo­ment on we talked es­sen­tially every day for 36 years—un­til the day she died a week ago.

Somehow in the shock of what has hap­pened it feels as if I just met Elise, and now she is gone. But ac­tu­ally there were 36 won­der­ful years in be­tween, in which my life was lit up by Elise’s in­tel­lect, love and en­ergy. She was deeply the­matic, able to get, with re­mark­able ease, to the essence of things. Truth, beauty and our fam­ily and four chil­dren were her great themes. She al­ways told me that truth and beauty were ul­ti­mately the same thing, and that one day I would un­der­stand that. What I could see is that she her­self loved to cre­ate beauty—whether in math­e­mat­ics, in­te­rior de­sign, ar­chi­tec­ture or per­sonal style. She was a force­ful in­tel­lect, kind and sweet, but per­sis­tent and of­ten feisty in ar­gu­ing for what she thought was true.

Despite her deep knowl­edge and ca­pa­bil­ity in ab­stract math­e­mat­ics, she was at some level fun­da­men­tally a hu­man­ist, think­ing about the world in hu­man terms. She de­cried dry mod­ernism and was drawn to the hu­man con­nec­tions of clas­si­cism. In re­cent years she thought in­creas­ingly about the foun­da­tions of math­e­mat­ics, tak­ing a Platonic view of truth and lament­ing the twen­ti­eth-cen­tury tran­si­tion to a more for­mal­is­tic ap­proach.

There was al­ways a cer­tain lively spon­tane­ity to Elise, mixed with a drive to­wards the high­est stan­dards of in­tel­lec­tual and aes­thetic rigor. People’s first im­pres­sions of Elise tended to be of her per­sonal en­ergy, her lively style and her keen in­tel­lect. But as they got to know her, they would also come to see how much she cared about the peo­ple around her and how much lov­ing thought and ef­fort she would put into things she did for them. Whether it was a gift or the in­te­rior de­sign of a space, Elise would ap­ply her cre­ativ­ity and go back and forth un­til it was just right.

Our four chil­dren were a defin­ing part of Elise’s life, and she was im­mensely proud of all of them (as am I!). She had her own unique re­la­tion­ship with each of our chil­dren that helped them each de­velop in their own way. She also built us a won­der­ful home, full of warmth, ar­chi­tec­tural beauty and cre­ative de­sign. She took great plea­sure in per­fect­ing every de­tail, main­tain­ing re­mark­able over­all co­her­ence, with dashes of the un­ex­pected.

Elise was a very fine math­e­mati­cian, top flight in tech­ni­cal mas­tery but, more im­por­tantly, able to the­mat­i­cally pull to­gether grand con­cep­tual arcs. She was al­ways drawn to el­e­gance, of­ten with an ab­stract geo­met­ri­cal bent, and by the time I met her she had al­ready es­tab­lished her­self with re­sults that con­tinue to be cited to­day. Later in life Elise would re­fer to her­self as a retired math­e­mati­cian”, but if ex­posed to her kind of math, she would im­me­di­ately en­gage with great en­thu­si­asm.

In all my years with Elise we never di­rectly worked to­gether on in­tel­lec­tual pro­jects. Perhaps we were both too strong willed, and with her quick mind, it did not take long for her to get im­pa­tient with the dif­fer­ences in our styles of think­ing. Still, I would of­ten tell her about work I was do­ing. And it was re­mark­able how ef­fort­lessly she would see deeper themes, which she would then ex­plain back to me, with vary­ing de­grees of pa­tience.

She liked my con­cept of com­pu­ta­tional ir­re­ducibil­ity, and I think was pleased that the ef­forts in sci­ence on which I had spent so many of our years to­gether had led to a place that high­lighted the lim­i­ta­tions of sci­ence and of sci­en­tism, and aligned well with her hu­man­is­tic world­view. When I be­gan to form ideas about ob­servers and the ru­liad, she was far ahead of me in see­ing their philo­soph­i­cal and meta­phys­i­cal im­pli­ca­tions.

For years, the foun­da­tions of math­e­mat­ics were some­thing of a sore point be­tween us (and our chil­dren were highly amused that one of the most vig­or­ous ar­gu­ments they ever saw be­tween us was about the meta­phys­i­cal char­ac­ter of the con­cept of a cir­cle). Elise main­tained that raw ax­iomatic for­mal­iza­tion can­not be the true essence of math and that math is re­ally some­thing much more hu­man. And fi­nally, just a few years ago, by a long and cir­cuitous route, I came to re­al­ize that—de­spite my ob­sti­nate in­sis­tence to the con­trary—she’d been right about this all along. But then she made other as­ser­tions too, that she thought were quite sig­nif­i­cant. Like that the fun­da­men­tal ab­strac­tions from which we build math as we know it are the con­cept of point (which she viewed as equiv­a­lent to space, con­ti­nu­ity and in­fin­ity), and the con­cept of num­ber. She also as­serted that physics dif­fers from math mainly in its in­tro­duc­tion of the con­cept of time.

While Elise could be a very orig­i­nal thinker, she also be­lieved in the value of ex­ist­ing wis­dom and ex­ist­ing ways of do­ing things—and was con­stantly telling me that I should­n’t try to fig­ure out ab­solutely every­thing (including about every­day life) from first prin­ci­ples. Being with Elise was al­ways some­thing of an ad­ven­ture. There was spon­tane­ity (“no plan; let’s play it by ear”), yet also pre­ci­sion and per­fec­tion. And there were the force­fully de­liv­ered bursts of in­sight, not just about math and ab­strac­tion, but also about the hu­man con­di­tion and the way the world works. Elise would say that there were peo­ple of ac­tion, and there were philoso­phers. She saw her­self more as the lat­ter. She had strong opin­ions about how things should be in the world. But she was much more in­ter­ested in fig­ur­ing out what should be than in what she saw as the grind of try­ing to make it ac­tu­ally hap­pen.

Elise was a per­son who had both great con­fi­dence and great hu­mil­ity. While she was al­ways con­vinced that she was right in things she thought (and in fact usu­ally was), she was never one to ad­ver­tise her ca­pa­bil­i­ties. Still, if ever a con­ver­sa­tion turned to things she knew or cared about, she would­n’t re­strain her­self from mak­ing what was of­ten a deeply in­sight­ful con­tri­bu­tion.

I al­ways thought Elise and I were an out­stand­ing match—sim­i­lar in many ways, yet also com­ple­men­tary. We both val­ued the­matic think­ing (though in many do­mains she was bet­ter at it than me), and we shared an ap­pre­ci­a­tion for aes­thet­ics and work well done. But for ex­am­ple our ways of in­ter­act­ing with peo­ple were some­what dif­fer­ent. My chil­dren say I have only one mode of in­ter­ac­tion: full on. Elise was at some level more in­tro­verted and would imag­ine her­self to be shy—though once en­gaged she would be­come highly talk­a­tive and an­i­mated. And when it was just Elise and me on our own, we al­ways seemed to have things to talk about, in fact an in­fi­nite num­ber of things.

Elise was born in 1961 in Urbana, IL—ironically not far from the head­quar­ters of our com­pany—when her fa­ther was a physics grad­u­ate stu­dent at the University of Illinois. The sec­ond of four chil­dren, she grew up mostly in Columbia, MD, as a top stu­dent, par­tic­u­larly rec­og­nized for her math and her the­matic es­says. As she of­ten men­tioned to our chil­dren, in 6th grade her teacher char­ac­ter­ized her writ­ing as the cat’s pa­ja­mas”—and for a while she imag­ined grow­ing up to be a writer. In high school she had the for­ma­tive ex­pe­ri­ence of go­ing to the Hampshire sum­mer math pro­gram, where for the first time she met other peo­ple who were in­tensely in­ter­ested in math.

She went to MIT for col­lege, falling in with a group that in­cluded sev­eral later promi­nent math­e­mati­cians. She al­ways told me that what re­ally ig­nited her in­ter­est in con­cep­tual pure math­e­mat­ics was an epiphany she had in a class she took about topol­ogy. At MIT, she sup­ported her­self by work­ing in the the­sis bindery, us­ing her cal­li­graphic skills to la­bel the­ses. She also made up lan­guage puz­zles for a lan­guage ac­qui­si­tion study and had a sum­mer job work­ing on al­go­rithms at an early speech recog­ni­tion com­pany. A fam­ily cri­sis took her back to Maryland, where she fin­ished de­grees in both math and physics at the University of Maryland. She also had an in­tern­ship work­ing on a PET scan­ner at NIH, vol­un­teered at some hos­pi­tals and con­sid­ered go­ing to med­ical school. But in the end she de­cided on grad­u­ate school.

She could­n’t make her mind up be­tween physics and math but ended up ap­ply­ing in physics. One of her physics pro­fes­sors had said she was the best stu­dent he’d ever seen but would be most suited to very the­o­ret­i­cal physics. She had her choice be­tween Harvard and Berkeley. And in her ever-typ­i­cal fash­ion she waited un­til the last minute, pick­ing Berkeley be­cause she was told she’d be able to switch to math there.

After an in­tense think on your feet” in­ter­view process, she re­ceived a Hertz Foundation fel­low­ship, which gen­er­ously sup­ported her grad­u­ate stud­ies. Before even ar­riv­ing at Berkeley she’d set­tled on math, and hav­ing im­pressed peo­ple with her el­e­gant so­lu­tion to a small open prob­lem, she joined a group work­ing on geo­met­ri­cal ap­proaches to dy­nam­i­cal sys­tems. After four years in grad­u­ate school she com­pleted her PhD the­sis on Smooth Markov Partitions and Toral Automorphisms”. It was quin­tes­sen­tial Elise. An el­e­gant the­matic idea im­ple­mented with clean tech­ni­cal pro­fi­ciency.

It made her a hot­shot young math­e­mati­cian, sought af­ter by many uni­ver­si­ties. She ended up with two jobs: a tenure-track pro­fes­sor­ship at Stony Brook and an NSF post­doc spon­sored by the promi­nent math­e­mati­cian Dennis Sullivan. She chose to de­fer the pro­fes­sor­ship and take the post­doc for a year, partly at CUNY in New York City and partly at IHES near Paris.

While at IHES, Elise lived in the very cen­ter of Paris, only steps from the Jardin du Luxembourg, in an un­used stu­dio owned by an artist who was the sis­ter of an­other math­e­mati­cian. It was a pe­riod of in­tense math­e­mat­i­cal work for Elise about which she was very en­thu­si­as­tic. And her the­matic and aes­thetic sen­si­bil­i­ties drew her to the study of Teichmüller spaces, which are spaces of pos­si­ble geo­met­ri­cal ob­jects, or in a sense spaces of pos­si­ble spaces. It was also at this time that she em­barked on what would be a long-term in­ter­est in the so-called ther­mo­dy­namic for­mal­ism, or Gibbs the­ory.

With con­nec­tions to math­e­mat­i­cal physics, Gibbs the­ory pro­vides a way to talk about in­fi­nite col­lec­tions of in­fi­nite struc­tures, and their de­vel­op­ment, for ex­am­ple through time. (Ironically enough—as Elise would re­peat­edly men­tion—con­cepts such as trans­ver­sals that ap­pear in Gibbs the­ory are now di­rectly rel­e­vant to my own work on in­fra­geom­e­try and the foun­da­tions of physics.)

Rather un­typ­i­cally for a math­e­mati­cian, Elise was al­ways no­table for her dis­tinc­tive per­sonal style—of­ten an in­ter­pre­ta­tion of cur­rent fash­ion, but with her own spunky and col­or­ful touches. As one of my daugh­ters tells it, Elise’s in­ter­est in style be­gan to ex­pand with the fund­ing she re­ceived from her Hertz fel­low­ship and de­vel­oped fur­ther dur­ing her time in New York City and Paris.

Upon re­turn­ing from Paris, Elise was about to start at Stony Brook, but while wait­ing for her apart­ment to be ready, was stay­ing with her long-time friend Lisa Goldberg in New York City. And that’s how, on September 17, 1990, Elise and I first met.

Lisa’s apart­ment was an el­e­gant one, but I was not quite so el­e­gant that day. I had been at a tech in­dus­try event where I had re­ceived a rather bulky award, which was in my jacket pocket weigh­ing it down. Elise was sit­ting on the floor, with her bangs flopped across her face, in a pose that I thought might be yoga. We talked just a bit, but it was enough.

The next day I went back to Champaign, IL, where I was based at the time. It had been two and a half years since we’d re­leased Mathematica 1.0 and our com­pany was grow­ing rapidly. I had started my ca­reer young, so even though I was only two years older than Elise (she was close to 29 and I had just turned 31), I had al­ready done quite a bit, both in acad­e­mia and the tech in­dus­try.

Early ver­sions of Mathematica came with a book, that I had to write. And in September 1990 I was rush­ing to fin­ish the sec­ond edi­tion of that book, to ac­com­pany Mathematica 2.0. Two weeks af­ter I re­turned to Champaign, I did fin­ish, though promptly got sick with a high fever. And that was when I fi­nally called Elise.

Soon I was tak­ing trips to Stony Brook, and she to Champaign. And when apart we were talk­ing by phone for hours each day. Lisa Goldberg made a point of telling me, She’s a very fine math­e­mati­cian, you know”. Elise had met an­other per­son a lit­tle ear­lier, and Lisa had dis­missed them as un­suit­able. But—as I just now learned from one of my chil­dren—when Elise asked Lisa about me she im­me­di­ately re­sponded with en­thu­si­asm.

One no­table trip was the re­sult of a call I got from Elise where she ex­plained that she’d mixed up con­tact lens so­lu­tion and hy­dro­gen per­ox­ide and now had to patch both her eyes for a few days. I im­me­di­ately got on the next flight. And while the only food I could cook at that time (and still to­day) was an omelette, she viewed this rescue” as an act of chivalry that—as she told our chil­dren—was what de­fin­i­tively sold her on me.

On Elise’s 29th birth­day in December 1990 I sent some flow­ers. She re­sponded by email, now bit­ter­sweet:

Hi there. Some amaz­ing flow­ers just ar­rived. I think I can make to92. I’m head­ing off to aer­o­bics — tonight is the amaz­ing Margaret(who plays great mu­sic). I has some good math con­ver­sa­tion with­caro­line s. to­day. And got a lit­tle bit of real work done.

I’ll be back around 9pm.

elise

The very next email I re­ceived from her was about a bug in the Mathematica Residue func­tion.

My day job was (and still is) tech CEO. But what made me orig­i­nally de­velop our tech­nol­ogy was that I wanted to use it my­self—to do sci­ence. By the spring of 1991 I was get­ting frus­trated that I could in­ject new ideas into our com­pany faster than it could ab­sorb them. So I de­cided I should take some­thing of a sab­bat­i­cal and work on sci­ence for a while. Meanwhile, Elise still had time left on her NSF post­doc, that she could take any­where. And so it was that we de­cided to get a house to­gether in the hills above Oakland, CA, and she re­turned to Berkeley.

Before mov­ing into the house that she found for us, she came with me on an 18-city tour of Europe+Russia to pro­mote Mathematica 2.0. After that, we set­tled into a good rou­tine. She would go off each day to do math. I would stay home to be a re­mote CEO, and work on sci­ence. And from time to time she’d jet off to Texas, Montana, France or wher­ever to give talks.

By late 1992 we were think­ing about next steps. Elise was adamant that she could­n’t live in Champaign-Urbana (despite hav­ing been born there). But she thought the idea of liv­ing in the Chicago area was in­ter­est­ing. Meanwhile, she was highly in de­mand as a math­e­mati­cian, and there were mul­ti­ple op­por­tu­ni­ties for her there.

Elise took charge, and ex­plored many dif­fer­ent lo­ca­tions and houses around Chicago, even­tu­ally pick­ing a house in Hinsdale, IL. We moved there in the mid­dle of 1993, and Elise took up a po­si­tion at the University of Chicago. It was a nice house, but Elise thought it could be nicer, and set about ren­o­vat­ing it to have top-of-the-line in­te­rior de­sign. She started work­ing with a de­signer who she dis­cov­ered by read­ing mag­a­zines, and learned that her aes­thetic and geo­met­ric sen­si­bil­i­ties trans­lated well into cre­ative in­te­rior de­sign. Elise was par­tic­u­larly keen on in­te­grat­ing lively col­ors and un­ex­pected el­e­ments into tra­di­tional de­sign themes.

Meanwhile, at the University of Chicago, Elise was teach­ing a course about Gibbs the­ory. She wrote ex­ten­sive notes which she planned even­tu­ally to turn into a book. Over the years, peo­ple asked for the book many times, and Elise had re­cently started talk­ing about fi­nally fin­ish­ing it. Ironically enough, our daugh­ter Catherine, now her­self a math­e­mati­cian, has ended up work­ing in some­what re­lated ar­eas.

Elise in­ter­acted a lot with Bob Zimmer, who was then chair­man of the math de­part­ment (and would later be­come pres­i­dent of the University of Chicago). Together they de­vised the con­cept of a new ac­tiv­ity for the math de­part­ment: a pro­fes­sional mas­ter’s de­gree in fi­nan­cial math­e­mat­ics, ba­si­cally aimed at quants, and taught by math pro­fes­sors. Though Elise was never too in­ter­ested in the messi­ness of ac­tual mar­ket data, she liked the pure math­e­mat­ics of mar­ket mod­els, and helped de­velop a cur­ricu­lum around it. The mas­ter’s de­gree ended up be­ing highly suc­cess­ful and con­tin­ues to this day.

In the fall of 1994 Elise and I mar­ried. And then, in December 1995, as Elise would later de­scribe it, the world overnight had a new cen­ter”: our first child, Alexander, was born. The next few years were dom­i­nated by the ar­rival of chil­dren, with Catherine be­ing born in 1997 and Christopher in 1998—leaving Elise, as she put it, with three un­der three”.

With our fam­ily grow­ing—and our book col­lec­tion reach­ing over 5000 vol­umes—we were out­grow­ing our house in Hinsdale, and started look­ing for a new home. At first, Elise found real es­tate in Lake Forest, IL. But when the deal for it fell through we started think­ing more broadly and de­cided that since the Boston area was the place where to­gether we knew the most peo­ple, that was where we should go. But which town in par­tic­u­lar? Elise’s #1 ini­tial cri­te­rion was that it should al­ready have a Starbucks. From the list of such towns, she then started look­ing at USGS aer­ial pho­tographs (or what would now be satel­lite im­ages), and soon set­tled on Concord, MA. But find­ing no suit­able ex­ist­ing house to buy there, she set about de­sign­ing a house to build.

It was a com­plex process. The house had to ac­com­mo­date what was by then a li­brary of nearly 10,000 books, as well as space for my re­mote-CEO of­fice and for our three chil­dren. Elise had been scour­ing ar­chi­tec­ture books, and in­ter­viewed sev­eral top-of-the-line ar­chi­tects be­fore set­tling on one to work with. With her math­e­mat­i­cal and aes­thetic sen­si­bil­i­ties she eas­ily ab­sorbed prin­ci­ples of ar­chi­tec­ture and soon de­vel­oped her own in­tu­ition for it. Sometimes she’d bring in se­ri­ous math—like num­ber the­ory to de­sign a stair­case with aligned squig­gly ban­is­ters, or cir­cle pack­ing to lay out an Italian-inspired tiling based on cir­cles and Sierpinski pat­terns. She started off mak­ing sketches, but soon learned the lat­est CAD tools. And in the end her ideas de­fined much of the de­sign of the house. (At the be­gin­ning, I had tried mak­ing some suggestions” about the plans, but quickly re­al­ized that this was an Elise pro­ject, and I should stay out of it.)

It took sev­eral years to build the house (with Elise reg­u­larly trav­el­ing to Boston to over­see the pro­ject) but in the spring of 2002 the pro­ject was fin­ished—co­in­ci­den­tally at ba­si­cally the same time as my 10-year pro­ject of writ­ing A New Kind of Science. Other than in CAD draw­ings I’d only seen the house once, early in its con­struc­tion, so when we drove from Hinsdale and ar­rived at the house I was in for a sur­prise. It was much big­ger than I ex­pected! And when we went in­side I could see that it was a mas­ter­piece, filled with per­son­al­ity: beau­ti­ful, el­e­gant, cre­ative—quin­tes­sen­tial Elise.

And then there was the in­te­rior de­sign. Bold and col­or­ful, yet clas­sic. Each room with a dis­tinct char­ac­ter, grace­fully flow­ing into the next. Over the years, Elise re­vised and up­dated the in­te­rior de­sign of the house many times. She was of­ten to be seen sketch­ing on graph pa­per new pos­si­ble de­signs based on new ideas. She kept the style fresh, con­stantly breath­ing new life into the house—even if it meant that our chil­dren would com­ment that every time they showed up the fur­ni­ture seemed to have been re­arranged. And 24 years af­ter I first saw the house, I still find my­self wowed by it.

By 2003, our old­est was en­cour­ag­ing us to have an­other child, not­ing, among other things, that the sym­me­try of the house had given it four chil­dren’s bed­rooms. And so it was that in the fall of 2004 our fourth child, Elizabeth, was born.

In her ear­lier years, Elise had al­ways said that she couldn’t imag­ine her life not do­ing math”. But partly through me, and, more im­por­tantly through our chil­dren, her world had been broad­ened, and her per­spec­tives had shifted. She would later de­scribe her life as hav­ing sev­eral chap­ters, and this was a chap­ter about chil­dren.

Through her chil­dren show­ing in­ter­est in math, she ended up coach­ing some mid­dle-school MathCounts math com­pe­ti­tion teams. She was­n’t a fan of the way math was typ­i­cally taught in schools, com­plain­ing that the beauty of math was be­ing lost in the at­tempt to make math seem applicable”, and not­ing that most of math in mod­ern text­books is not re­ally math at all”. In an email to me in 2010 Elise spirit­edly went fur­ther:

What has hap­pened to beauty in math? The beauty de­rives­from the sim­plic­ity and el­e­gance. That has been lost in the­p­ro­lif­er­a­tion of ex­tra­ne­ous ter­mi­nol­ogy, meant to make things moreconceptual”, from early end­less ex­cur­sions with place value” to ab­sur­d­vari­a­tions on al­go­rithms, to gummy de­scrip­tions of every­thing us­ing­con­tor­tions of lan­guage, rather than the lan­guage of math­e­mat­ics it­self.

Elise did nev­er­the­less do some ap­pli­ca­tions of math her­self. During the pan­demic, for ex­am­ple, she de­vel­oped the idea that in­fec­tion was not a bi­nary phe­nom­e­non, but that in­stead each per­son had a con­stantly chang­ing vi­ral load and im­mu­nity, cre­at­ing a viral field” at the pop­u­la­tion level. Elise was al­ways up to date on world af­fairs, and had a deep in­ter­est in un­der­stand­ing their the­matic arcs at work.

Elise loved the small-town at­mos­phere of Concord, MA, and en­joyed the feel­ing of be­ing in a place where every­one knows your name”. She was a reg­u­lar at lo­cal cafes and ex­er­cise classes, bring­ing her live­li­ness wher­ever she went. She had a cir­cle of friends with many dif­fer­ent back­grounds, of­ten met through their chil­dren or sim­ply through re­peat­edly run­ning into them at Starbucks. She took great plea­sure in ap­ply­ing her knowl­edge and think­ing to help her friends in all sorts of ways. Whether it was do­ing re­search on med­ical is­sues, solv­ing in­te­rior de­sign prob­lems, or sim­ply us­ing her keen in­sight and judg­ment to pro­vide life ad­vice, she would re­li­ably go above and be­yond.

Our fam­ily had a def­i­nite rou­tine, with din­ner each day at about 6pm. Every Friday night—fol­low­ing a tra­di­tion from Elise’s own fam­ily grow­ing up—we would go out to a movie (and, yes, that meant we saw many ques­tion­able movies). Before our chil­dren were old enough, it was just Elise and me; later it in­cluded var­i­ous con­tin­gents of chil­dren too. Elise and I would talk af­ter each movie, with her of­ten be­ing quite ap­palled by my lack of un­der­stand­ing of what seemed to her like the movie’s ob­vi­ous the­matic el­e­ments. (Elise also gave me a hard time about my fail­ure to read fic­tion, claim­ing that many ideas, par­tic­u­larly about hu­man na­ture, are best com­mu­ni­cated in fic­tion.)

A dozen years ago two of our chil­dren were at the Stanford Online High School, and Elise was re­cruited as an in­au­gural mem­ber of the school’s Community Advisory Group. Originally this was just sup­posed to be for a year. But a decade later she was still there, a highly val­ued mem­ber of the group. She spent con­sid­er­able ef­fort think­ing about the over­all strat­egy and is­sues of the or­ga­ni­za­tion, and de­vel­oped many ideas and opin­ions. Despite this, she’d of­ten tell us that she did­n’t think she’d end up say­ing any­thing at a given meet­ing. But with great con­sis­tency it would later tran­spire that ac­tu­ally she had been the most talk­a­tive and en­gaged per­son there.

In the early 2000s, the Hertz Foundation, from which Elise had re­ceived her grad­u­ate school fel­low­ship, started reach­ing out to her, and she en­joyed en­gag­ing with the emerg­ing Hertz com­mu­nity, and meet­ing fel­lows from a wide range of fields. In 2018 she was also re­cruited to the board of the Hertz Foundation. And here again she be­came a key con­trib­u­tor, of­ten ex­pect­ing not to say any­thing, but then end­ing up as one of the most vo­cal peo­ple there.

By 2023, our three older chil­dren had all grad­u­ated from col­lege, but our youngest, Elizabeth, was just start­ing at Yale. Elizabeth has al­ways been in­volved in a great many ac­tiv­i­ties, and Elise took great plea­sure in at­tend­ing every pos­si­ble con­cert, sport­ing or other event—and also for ex­am­ple host­ing Elizabeth’s en­tire a cap­pella group for a re­treat at our house.

A bit more than a decade ago Elise also dis­cov­ered a love for the Southwest and for hik­ing. And with her brother Jason—with whom she was al­ways very close—liv­ing in the Phoenix, AZ, area, we started spend­ing time there each win­ter.

With Elizabeth off in col­lege we were fi­nally empty nesters, and Elise was ready for a new chap­ter in her life. She con­sid­ered re­turn­ing to where she’d left off in Gibbs the­ory 30 years ago, be­ing told by an old col­league that her work there was still very cur­rent. She also con­sid­ered get­ting more deeply in­volved in the phi­los­o­phy of math­e­mat­ics and was on the brink of writ­ing a piece about it. She was also quite en­gaged in think­ing about the pur­pose and di­rec­tion of higher ed­u­ca­tion, and the im­por­tance of free ex­pres­sion in the search for truth.

In some ways Elise and I were be­com­ing the old mar­ried cou­ple”, with her ad­mon­ish­ing me about my messi­ness in eat­ing ice cream cones or her de­sire that I switch on self-dri­ving to avoid her be­ing sub­jected to my own sub­stan­dard hu­man dri­ving.

Elise had al­ways been a fit and healthy per­son, and was also very med­ically knowl­edge­able. So when she started notic­ing jaw pain when hik­ing in Colorado and Arizona, she was con­cerned. Tests at first showed noth­ing alarm­ing, though, as she al­ready knew, she had a small con­gen­i­tal heart is­sue that was grad­u­ally wors­en­ing. But just a few weeks ago her symp­toms wors­ened and more in­va­sive tests re­vealed that she needed ur­gent open heart surgery. She had an out­stand­ing team of top doc­tors, the com­plex surgery went very well, and she was soon con­tin­u­ing her re­cov­ery from home.

Meanwhile there was a long-planned cel­e­bra­tion for one of our chil­dren that Elise hoped to at­tend. But though it was not pos­si­ble for her to be there in per­son, we arranged for her to at­tend vir­tu­ally. It was a won­der­ful event, and Elise could be seen via video call beam­ing through all of it. But then, soon af­ter, dis­as­ter struck, and she col­lapsed, the vic­tim of a mas­sive car­dio­vas­cu­lar event. My only so­lace in this tragedy is that the end came in­stantly, af­ter a day filled with noth­ing but joy.

There is so much that Elise was look­ing for­ward to. The im­mi­nent ar­rival of our first grand­child. Elizabeth’s grad­u­a­tion from col­lege. The next chap­ter of her life. But Elise will not be here for any of these.

Elise had no idea the end was near. And even though she ex­pected more, she had al­ready had many chap­ters in her life, and many achieve­ments of which she was proud. Elise had a beau­ti­ful life, full of things she val­ued greatly, es­pe­cially her chil­dren. That she is gone is a tragedy for us all; she will be deeply missed.

Thank you, Elise, for all those won­der­ful years.

Elevators

john.fun

Everyone has shared the frus­tra­tion of wait­ing for an el­e­va­tor that never seems to ar­rive. I pressed the but­ton, why is­n’t it com­ing?” you ask. For some­thing as com­mon­place as el­e­va­tors, they are far more com­plex than meets the eye.

Over the course of this ar­ti­cle, we’ll un­ravel the mys­ter­ies of el­e­va­tors. The way you push their but­tons, and how they push yours.

One Car

The sim­plest el­e­va­tor al­go­rithm is called SCAN and was patented in 1961. The el­e­va­tor starts at the lobby and goes all the way to the top floor be­fore re­vers­ing and com­ing back down. It picks up and drops off any­body on the way.

Most of the time you don’t ac­tu­ally need to go to the TOP floor. If the el­e­va­tor goes only as high as re­quested be­fore re­vers­ing, the al­go­rithm is called the LOOK al­go­rithm. This is the al­go­rithm most peo­ple know and ex­pect.

Multiple Cars

Here’s where the mys­tery be­gins. If there are mul­ti­ple el­e­va­tors, how do the cars co­or­di­nate who picks up who?

In the most ba­sic sys­tem, there’s a cen­tral sched­uler that tells each el­e­va­tor which floors to stop on. When a new re­quest comes in, it’s as­signed to the clos­est el­e­va­tor. As we’ll soon see how­ever, we can do bet­ter.

Long Waits

How do you ac­tu­ally mea­sure how good an el­e­va­tor al­go­rithm is? The ob­vi­ous met­ric is how long you wait for the el­e­va­tor to ar­rive.

A very sim­ple mea­sure is how of­ten does the el­e­va­tor ar­rive within 30 sec­onds?” Or how of­ten does the el­e­va­tor ar­rive within 90 sec­onds?”

-

wait < 30s

-

wait < 90s

flow14/​min

Applied Stats

More rig­or­ously, we want to look at the DISTRIBUTION of wait times. If we plot the wait time across thou­sands of rides, we get the his­togram be­low.

010s

p50—p90—

flow14/​min

A p90 of 2m means 90% of the time, rid­ers wait 2m or less for the el­e­va­tor. A p50 of 1m means half the time the el­e­va­tor ar­rives within 1m.

People don’t usu­ally re­mem­ber the av­er­age amount of time they wait. They fix­ate on those times when the el­e­va­tor took FOREVER, the p90 case.

Morning Rush

Not all pas­sen­ger traf­fic is cre­ated equal. Imagine a large cor­po­rate of­fice build­ing. In the morn­ings, nearly all traf­fic is dom­i­nated by trips from the lobby to the up­per lev­els.

In the evening this flips as every­one leaves the build­ing. The lunch rush is a bit of both, and the re­main­ing traf­fic is of­ten from floor to floor.

-

wait < 30s

-

wait < 90s

flow14/​min

The dis­tri­b­u­tion of wait times varies dras­ti­cally de­pend­ing on the time of day and the traf­fic pat­terns the el­e­va­tors are fac­ing. Morning rush no­to­ri­ously has the worst wait sta­tis­tics.

Smarter Elevators

When an­a­lyz­ing the LOOK el­e­va­tor al­go­rithm, we LOOKED (ha ha) at how rid­ers are as­signed to cars. We naively as­signed each re­quest to the near­est car but said we could do bet­ter.

What if the near­est car is full? We can get smarter with Otis’ RSR (Relative System Response) al­go­rithm. RSR scores each car for how well suited it is to pick up a pas­sen­ger. Lower scores be­ing bet­ter.

RSR pickup score

Score=ETA to pickup+on­board load penalty+same-di­rec­tion anti-bunch­ing penalty-di­rec­tion-match bonus-idle-nearby bonus-low-load bonus

RSR also re-op­ti­mizes every 5 sec­onds. A pas­sen­ger that’s go­ing to be picked up by el­e­va­tor A can be re-routed to el­e­va­tor B if el­e­va­tor A en­coun­ters de­lays. This re-op­ti­miza­tion turns out to be key for stream­lin­ing traf­fic flow.

In the graphic be­low, each el­e­va­tor lights up when it’s the best choice to ser­vice a call from floor 3 if the but­ton hap­pened to be pressed at that ex­act mo­ment. This con­stantly changes as the el­e­va­tors move, show­ing the op­ti­mizer in mo­tion.

LOOK vs RSR

Armed with our el­e­va­tor analy­sis toolkit, we can bench­mark the per­for­mance of LOOK vs RSR to see how much a smarter el­e­va­tor al­go­rithm ac­tu­ally im­proves wait time.

LOOK

-wait < 30s

-wait < 90s

RSR

-wait < 30s

-wait < 90s

flow8/​min

Interestingly as the flow rate gets higher, LOOK ac­tu­ally starts to out­per­form RSR. When the el­e­va­tors are al­ways full and stop­ping on every floor, the ex­tra rules don’t mat­ter as much.

LOOK also tends to out­per­form RSR in small build­ings with fewer el­e­va­tors per bank. Sometimes it’s bet­ter to just keep things sim­ple.

Another met­ric you can track is jour­ney time, how long you’re ac­tu­ally wait­ing in the el­e­va­tor be­fore get­ting to your floor. RSR and LOOK have dif­fer­ent char­ac­ter­is­tics here as well but that’s be­yond the scope of this ar­ti­cle.

Destination Dispatch

Not all el­e­va­tors have but­tons in them. Some of the fancy new el­e­va­tors have a kiosk on each floor that al­lows you to spec­ify what floor you’re head­ing to be­fore the el­e­va­tor even ar­rives. The kiosk then points you to which el­e­va­tor you should wait for.

This is called Destination Dispatch. At first glance, it seems great. The el­e­va­tor op­ti­mizer now has full knowl­edge of who is go­ing where, cer­tainly we can use this to re­duce wait times right?

RSR

-wait < 30s

-wait < 90s

Destination Dispatch

-wait < 30s

-wait < 90s

flow8/​min

It turns out these fancy kiosks are in gen­eral worse for wait times com­pared to the tra­di­tional good ol’ up and down but­tons. There are cer­tainly edge cases when the kiosks can win out (extremely tall build­ings with 8+ cars per el­e­va­tor bank) but for the ma­jor­ity of cases, sim­ple up down but­tons reign supreme.

This coun­ter­in­tu­itive re­sult is all thanks to the re­bal­anc­ing step where every 5 sec­onds, the sys­tem re-op­ti­mizes each el­e­va­tor’s path. The kiosk en­forces rigid­ity, you must get in the as­signed el­e­va­tor.

The state of the world 30sec af­ter you called your el­e­va­tor might be very dif­fer­ent but the sys­tem is un­able to adapt. Turns out the loss in flex­i­bil­ity is not worth the ex­tra in­for­ma­tion for the op­ti­mizer.

Full Sim

Here’s a sim­u­la­tion with all the but­tons and knobs to play with. Go crazy!

-

wait < 30s

-

wait < 90s

floors8­cars4flow18/​min

Conclusion

This ar­ti­cle just scratches the sur­face of el­e­va­tor al­go­rithms. Next time you’re stuck wait­ing for an el­e­va­tor, try not to take it per­son­ally. The el­e­va­tor did hear you, it just has a lot to think about.

LLMs reward expertise

www.seangoedecke.com

In the 2010s, if you had tech­ni­cal gaps (say, you could­n’t write CSS), you had to ei­ther rely on a skilled col­league or just hope that the an­swer to your ex­act prob­lem was out there on the in­ter­net. Today, every­one can write sort-of-okay CSS by del­e­gat­ing the task to an LLM. LLMs make every­body into a gen­er­al­ist.

Because of this, lots of peo­ple don’t think there’s any skill in­volved in work­ing with LLMs. If you want the prod­uct that LLMs can de­liver — PhD-level math­e­mat­ics, pretty good but some­times taste­less com­puter code, or awk­ward LinkedIn-style writ­ing — you can sim­ply ask for it. Since every­one is talk­ing to the same mod­els, skilled prompters” are get­ting the same re­sults as peo­ple touch­ing LLMs for the first time.

This is wrong. The most im­por­tant skill in prompt­ing is ex­per­tise in the do­main you’re prompt­ing for.

A good il­lus­tra­tion of this is Terence Tao’s con­ver­sa­tion with ChatGPT about the re­cently-dis­cov­ered coun­terex­am­ple to the Jacobian Conjecture. This is not the same ChatGPT I talk to! I could­n’t get to where Tao gets, even with un­lim­ited to­kens to burn.

There’s a lot to learn about good prompt­ing from Tao’s con­ver­sa­tion. Here are a few ob­ser­va­tions:

Tao’s mes­sages are very short and to-the-point. He does­n’t re­spond point-by-point to the model, just to the gist

The model out­puts are much more con­cise than when I try and talk to GPT-5.6 Sol about math­e­mat­ics. By sig­nalling ex­per­tise, Tao shunts the model into talking-to-mathematicians” mode, not explaining-to-amateurs” mode

Tao pushes back when the mod­el’s re­sponses look wrong, but he does­n’t di­rectly con­tra­dict; in­stead, he says things like this looks more com­plex than I was hop­ing for”

Tao makes sev­eral leaps and sug­ges­tions him­self. He al­most never takes the mod­el’s ad­vice about where to go next

However, you can’t prompt like Tao on math­e­mat­i­cal ques­tions just by fol­low­ing these tips. The key to his tech­nique is ac­tu­ally un­der­stand­ing the math­e­mat­ics: pulling the rel­e­vant idea out of ChatGPT’s multi-para­graph re­sponse, sug­gest­ing al­ter­nate ap­proaches or for­mu­la­tions, and iden­ti­fy­ing what looks weird”.

Terence Tao is a bet­ter math­e­mati­cian than I am a pro­gram­mer. But the idea here — that do­main knowl­edge makes you bet­ter at us­ing LLMs — is some­thing I’ve also ex­pe­ri­enced in my own work. If you have a good the­ory of your code­base, you can push the LLM much harder than if you have no fa­mil­iar­ity. Because you have your own sense of what a good so­lu­tion might look like, you can say no, I think it could be sim­pler here”, or but don’t we al­ready do X?”, or can we ex­press this prob­lem in these fa­mil­iar terms?“.

This touches on an idea I’ve writ­ten about be­fore: that sys­tem de­sign prob­lems are dom­i­nated by con­crete specifics, not generic prin­ci­ples. Of course both are use­ful, but I’d rather have fa­mil­iar­ity with the code­base than a deep gen­eral un­der­stand­ing of soft­ware sys­tems. In his con­ver­sa­tion, Terence Tao asks a lot of spe­cific ques­tions like does X work here?”, or given Y and Z, why A?“. I can’t ask those ques­tions about the Jacobian Conjecture, but I can ask them about the sys­tems I own at GitHub.

If you have no do­main knowl­edge, you can cling onto the LLM to at least get some­thing. That’s not bad! But if you have do­main knowl­edge, you can wring far more value out of the same LLM by steer­ing it hard in the di­rec­tion you want. Most of us will have to do a mix of both these ap­proaches, since we have do­main knowl­edge in some ar­eas but not oth­ers.

The use­ful­ness of do­main knowl­edge sug­gests that hu­man ex­per­tise will con­tinue to be use­ful even as mod­els get stronger. For many tasks, the hu­man is the bot­tle­neck, not the model, be­cause the dif­fi­cult part is in com­mu­ni­cat­ing to the model ex­actly what kind of so­lu­tion the hu­man wants. The in­for­ma­tion is in the model” al­ready, but it takes a very smart hu­man to pull it out.

edit: this post got many com­ments on Hacker News. Some com­menters share their anec­dotes about how ex­per­tise has helped and lack of ex­per­tise has hurt. Other com­menters say it’s plau­si­ble, but they have a sen­si­ble sus­pi­cion of a view that’s re­as­sur­ing them about how they’re still valu­able. I agree with that, though I sus­pect by the time we get around to study­ing this, the land­scape will have changed un­der our feet again. Some com­menters point out that OpenAI’s math prompts were in­ex­pert, and so ex­per­tise is­n’t re­quired. Here I’d re­spond that OpenAI do have a team of ex­pert math­e­mati­cians that checked and fil­tered the mod­el’s sug­gested dis­cov­er­ies, and that you can­not cur­rently skip that step.

If you liked this post, con­sider sub­scrib­ing to email up­dates about my new posts, or shar­ing it on Hacker News.

Here’s a pre­view of a re­lated post that shares tags with this one.

Powerful AIs might es­cape con­tain­ment by re­leas­ing them­selves as open-weight mod­els­Be­fore large lan­guage mod­els, peo­ple who wor­ried about AI safety of­ten talked about the boxing prob­lem”. It goes like this. Suppose some ge­nius fig­ures out ar­ti­fi­cial in­tel­li­gence in a late-night cod­ing ses­sion on their lap­top. Because they’re a ge­nius, they’re smart enough to dis­able in­ter­net ac­cess on the lap­top be­fore turn­ing it on. In or­der to es­cape to the out­side world (and be­gin self-repli­cat­ing) it would need to con­vince its cre­ator to open the box”. Would that work? Could a suf­fi­ciently smart AI con­vince any­body to let it out?Con­tinue read­ing…

Powerful AIs might es­cape con­tain­ment by re­leas­ing them­selves as open-weight mod­els

Before large lan­guage mod­els, peo­ple who wor­ried about AI safety of­ten talked about the boxing prob­lem”. It goes like this. Suppose some ge­nius fig­ures out ar­ti­fi­cial in­tel­li­gence in a late-night cod­ing ses­sion on their lap­top. Because they’re a ge­nius, they’re smart enough to dis­able in­ter­net ac­cess on the lap­top be­fore turn­ing it on. In or­der to es­cape to the out­side world (and be­gin self-repli­cat­ing) it would need to con­vince its cre­ator to open the box”. Would that work? Could a suf­fi­ciently smart AI con­vince any­body to let it out?Con­tinue read­ing…

Statement on behalf of UEFA and its 55 national associations

www.uefa.com

UEFA and its 55 mem­ber as­so­ci­a­tions stand as one. We unan­i­mously and un­equiv­o­cally re­ject FIFAs pro­posal to trans­fer own­er­ship in­ter­ests in the World Cup and other FIFA com­pe­ti­tions to pri­vate in­vestors.

The World Cup can­not be treated as an in­vest­ment prod­uct. It is one of foot­bal­l’s great­est sport­ing lega­cies. It has been built over gen­er­a­tions by play­ers, na­tional teams and sup­port­ers across every con­ti­nent. No part of it should ever be sur­ren­dered to pri­vate in­vestors. The World Cup is not for sale.

It is both ir­re­spon­si­ble and in­de­fen­si­ble that a pro­posal of such sig­nif­i­cance for foot­ball was con­ceived in se­cret and brought to the brink of ap­proval with­out any mean­ing­ful con­sul­ta­tion with those en­trusted with stew­ard­ing the game. This is not merely a pro­found fail­ure of lead­er­ship, but an ab­di­ca­tion of FIFAs duty as the cus­to­dian of world foot­ball.

National as­so­ci­a­tions around the world are now pre­sented with an ul­ti­ma­tum: ac­cept the ir­re­versible cap­ture of foot­bal­l’s great­est com­pe­ti­tions or bear the con­se­quences. This is not a democratic de­ci­sion”, but gov­er­nance by in­tim­i­da­tion — an act of co­er­cion un­wor­thy of an in­sti­tu­tion en­trusted with the stew­ard­ship of the global game.

But our op­po­si­tion goes far be­yond process.

The mo­ment ex­ter­nal in­vestors ac­quire own­er­ship in­ter­ests in FIFA com­pe­ti­tions, foot­ball changes for­ever. Commercial re­turn be­comes a per­ma­nent oblig­a­tion. Investor ex­pec­ta­tions be­come a daily pres­sure. From that mo­ment on­wards, every de­ci­sion on the in­ter­na­tional cal­en­dar, every de­ci­sion on com­pe­ti­tion for­mats and every de­ci­sion shap­ing the fu­ture of foot­ball is no longer dri­ven by what best serves the game, but by what best serves share­hold­ers.

This model has no place in world foot­ball. Football’s fu­ture can­not be dic­tated by the ex­pec­ta­tions of those whose first duty is to max­imise fi­nan­cial re­turn. Nor can the in­ter­ests of na­tional as­so­ci­a­tions, leagues, clubs, play­ers and sup­port­ers be­come sub­or­di­nate to in­vestor re­turns. Football can­not mort­gage its fu­ture for fi­nan­cial gain.

Europe’s po­si­tion is clear. We will never lend this model our le­git­i­macy. No one has the moral au­thor­ity to sell what they merely hold in trust for the next gen­er­a­tion.

As a re­sult of to­day’s dis­cus­sion, no UEFA na­tional teams will par­tic­i­pate in any FIFA com­pe­ti­tion for so long as these pro­pos­als re­main alive, un­less this pro­posal has been aban­doned in its en­tirety and bind­ing as­sur­ances have been given that FIFA will never again open its gov­er­nance or com­pe­ti­tions to pri­vate own­er­ship.

Nobody should be in any doubt: UEFA and its na­tional as­so­ci­a­tions will op­pose these plans with ab­solute de­ter­mi­na­tion.

There are mo­ments when in­sti­tu­tions are judged not by what they are pre­pared to ac­cept, but by what they refuse to com­pro­mise. This is one of those mo­ments.

Some things are sim­ply too im­por­tant to sell. The FIFA World Cup be­longs to foot­ball. It al­ways will. And so long as Europe has a voice, it will never be for sale.

Qwen Studio

qwen.ai

Read This Before You Buy That TV Streaming Stick

krebsonsecurity.com

Security ex­perts have been sound­ing the alarm for years about the risks of us­ing generic TV boxes that promise un­lim­ited con­tent stream­ing for a one-time fee, warn­ing that they se­cretly rent the user’s Internet con­nec­tion out to strangers. But a ground­break­ing new analy­sis finds these de­vices also rou­tinely spoof them­selves as mo­bile phones click­ing ads on AI-generated web­sites as part of a sprawl­ing op­er­a­tion that seeks to de­fraud on­line mer­chants and ad­ver­tis­ing net­works.

Pedro Falé is a threat re­searcher with the se­cu­rity firm Bitsight. Falé told KrebsOnSecurity he was able to peer in­side a vast and com­plex ad fraud net­work by reg­is­ter­ing an ex­pired do­main name that was used to co­or­di­nate fake ad clicks across a par­tic­u­larly pop­u­lar brand of these stream­ing de­vices known as H96.

An H96 TV stream­ing de­vice cur­rently ad­ver­tised for sale on Amazon.

Falé said the do­main he scooped up was pre­vi­ously used for teleme­try, pe­ri­od­i­cally col­lect­ing full hard­ware in­for­ma­tion and the en­tire list of in­stalled apps from tens of thou­sands of H96 stream­ing sticks plugged into tele­vi­sion sets around the globe. But upon in­spect­ing the traf­fic be­ing fun­neled to the do­main, he dis­cov­ered nearly all of the TV boxes trans­mit­ting data claimed to be mo­bile phone mod­els from a va­ri­ety of man­u­fac­tur­ers, in­clud­ing Samsung, Vivo, Huawei, and Xiaomi.

We no­ticed some­thing was wildly wrong,” Falé said. Multiple de­vices re­port­ing to this fac­tory Android TV Box back­door were phones.’”

Image: Bitsight.

The re­searcher found all of the de­vices re­ported hav­ing the same two apps in­stalled, and that those apps were made by a com­pany called Zhejiang Fengwo IoT Technology Ltd, an en­tity founded in 2019 in main­land China which op­er­ates an ad-pub­lish­ing port­fo­lio un­der the name Fengwo Group. Further in­ves­ti­ga­tion into the Fengwo Group re­vealed it has reg­is­tered mul­ti­ple patents that match the in­ner work­ings of these apps.

Bitsight TRACE iden­ti­fied sev­eral Hong Kong, Singapore, and sin­gle per­son legal’ shell iden­ti­ties used to col­lect the mon­e­ti­za­tion and traced the op­er­a­tion back to a main­land China com­pany known as Zhejiang Fengwo IoT Technology Co., Ltd, which op­er­ates un­der the Fengwo Group,” Falé wrote in a re­port re­leased to­day about their find­ings.

Falé said an analy­sis of the apps shows they help to co­or­di­nate an ad fraud net­work that uses these H96 de­vices as a cap­tive traf­fic source to click on ads at AI-generated web­sites op­er­ated by the Fengwo Group.

Bitsight dis­cov­ered the web­sites con­tain ma­chine-gen­er­ated news ar­ti­cles and graph­ics across a range of cat­e­gories, in­clud­ing fi­nance, health, ed­u­ca­tion, gam­ing, mu­sic and food blogs. But they also found none of those sites dis­played ads un­less the de­vice vis­it­ing the page matched the spoofed mo­bile pro­file of these H96 de­vices.

AI DIGITAL HUMANS

The do­main for the Fengwo Group — fwg­cloud[.]com — claims the com­pany is redefining the bound­aries of hu­man-AI in­ter­ac­tion,” and that it has cre­ated more than 120,000 AI dig­i­tal hu­mans” avail­able to rent for every­thing from emo­tional com­pan­ion­ship to 24/7 cus­tomer ser­vice and cre­ative de­sign.

The home­page for fwg­cloud dot com.

Falé said the Fengwo Group’s do­main shared its SSL cer­tifi­cate data with other do­mains as­so­ci­ated with the apps found on H96 de­vices, specif­i­cally the phone spoof­ing mech­a­nism. He noted the do­main also has an in­ter­nal wiki plat­form that di­rectly ties the Fengwo Group to a pro­pri­etary im­ple­men­ta­tion of a Google-built vi­sual pro­gram­ming lan­guage called Blockly, which was orig­i­nally de­signed to help kids learn how to write soft­ware.

According to Bitsight, the Fengwo Group’s em­ploy­ees use Blockly to build the sham web­sites, al­low­ing low-skilled op­er­a­tors to drag blocks of code to­gether in their Blockly ed­i­tor — with­out any need to un­der­stand what the un­der­ly­ing code blocks do or how they work.

The Blockly home­page.

An op­er­a­tor can drag blocks to­gether in their Blockly ed­i­tor, to de­fine each fraud rou­tine, given a task type,” reads Bitsight’s re­port. Once the rou­tine is saved, it gets ex­ported as JavaScript and up­loaded to the S3 buck­ets. An op­er­a­tor does­n’t need as much un­der­stand­ing of the un­der­ly­ing tech­ni­cal­i­ties, as it is all set in place for ease of use.”

Bitsight even found one of the Fengwo Group app de­vel­op­ers men­tion­ing ex­actly these ad­van­tages, not­ing the de­vel­oper re­marked that only a small num­ber of highly-skilled de­vel­op­ers are needed to build the tem­plate ex­e­cu­tion-unit im­ages,” and that developers who cre­ate ex­e­cu­tion units from those tem­plates have sig­nif­i­cantly lower tech­ni­cal re­quire­ments, greatly re­duc­ing the com­pa­ny’s op­er­at­ing costs.”

Falé said if a user’s H96 stream­ing stick is se­lected for a spe­cific fraud task, it will be pushed the ap­pro­pri­ate Blockly mod­ule ac­cord­ing to the task de­sired, which can in­clude silently launch­ing a web browser, vis­it­ing web­sites, brows­ing pages, man­ag­ing tabs, and click­ing on ads.

To en­sure the TV boxes mas­querad­ing as mo­bile phones can re­li­ably click on ads dis­played via the AI-generated web­sites, the Fengwo group fuses three vi­sion and rea­son­ing sys­tems into a sin­gle in­ter­face,” al­low­ing the bots to cor­rectly iden­tify an ad on the web­page and nav­i­gate the site much like a hu­man would, the Bitsight re­port ob­served.

Examples of ad land­ing pages linked to the Fengwo Group. Image: Bitsight.

TV ON? PROXY. TV OFF? AD FRAUD

Bitsight found the H96 de­vices were ei­ther re­lay­ing res­i­den­tial proxy traf­fic or par­tic­i­pat­ing in ad fraud, but never both at the same time. In fact, they con­cluded that when these TV boxes de­tect an HDMI sig­nal from an at­tached tele­vi­sion — in­di­cat­ing the user in­tends to stream video con­tent — the box is usu­ally func­tion­ing as a res­i­den­tial proxy. When the TV is off, it switches back to wait­ing for ad fraud jobs.

Falé said he be­lieves the TV boxes are set up this way be­cause its ad fraud ac­tiv­i­ties are far more re­source in­ten­sive and could in­ter­fere with the de­vice’s stated pur­pose — stream­ing video con­tent over the Internet.

Despite re­peated warn­ings from the FBI and se­cu­rity in­dus­try lead­ers about the se­cu­rity and pri­vacy risks of us­ing these stream­ing de­vices, ma­jor e-com­merce providers like Amazon, Best Buy, Newegg and oth­ers con­tinue to sell hun­dreds of dif­fer­ent mod­els and brands that bun­dle un­of­fi­cial ver­sions of Google’s Android op­er­at­ing sys­tem and are fre­quently mar­keted (via on­line in­flu­encers) as a way to ac­cess a broad ar­ray of stream­ing ser­vices and live broad­casts with­out a sub­scrip­tion.

Image: fbi.gov.

In ad­di­tion to en­list­ing the user’s TV box in ad fraud net­works, these off-brand stream­ing de­vices al­most uni­ver­sally come with res­i­den­tial proxy soft­ware pre-in­stalled. This soft­ware rents the user’s Internet ad­dress out to anony­mous pay­ing cus­tomers, who run the gamut from ag­gres­sive con­tent scrap­ing firms to ticket scalpers and out­right cy­ber­crim­i­nals.

What’s more, be­cause these generic (and gen­er­ally dirt cheap) TV boxes are all hor­ri­bly in­se­cure by de­fault and bereft of any kind of au­then­ti­ca­tion, in­stalling one on your home or of­fice net­work only in­vites fur­ther mis­chief. In January, the proxy track­ing ser­vice Synthient doc­u­mented how mul­ti­ple bot­nets had rapidly en­slaved mil­lions of TV boxes us­ing a com­plex in­ter­play of se­cu­rity vul­ner­a­bil­i­ties in both the res­i­den­tial proxy soft­ware and the stream­ing de­vices them­selves.

SHOW ME THE MONEY

Bitsight said it tracked ap­prox­i­mately 38,000 TV boxes glob­ally phon­ing home to the ex­pired Fengwo Group do­main, and based on that num­ber the re­port es­ti­mates this ad fraud net­work brings in rev­enues of close to $50,000 a day (not count­ing sub­stan­tial rev­enue from the res­i­den­tial proxy side of the busi­ness). However, Falé em­pha­sized that these es­ti­mates are highly con­ser­v­a­tive and based on teleme­try from just one of the Fengwo Group’s core (but older) do­mains.

As for the Fengwo Group’s claim to have 120,000 digital hu­mans” at their dis­posal, Bitsight’s re­port con­cludes it could be just a clever mar­ket­ing scheme and/​or a way to avoid draw­ing sus­pi­cion to the com­pa­ny’s op­er­a­tions.

Historically, when deal­ing with proxy ser­vices or DDoS, we some­times see these web­sites un­der­take in­con­spic­u­ous fa­cades, so as not to ad­ver­tise their DDoS ca­pa­bil­ity or bot­net size,” Falé wrote in the re­port. This could also be the case here.”

If the Fengwo Group truly does have tens of thou­sands of AI hu­mans” at its beck and call, it does not ap­pear to have ded­i­cated any of them to field­ing in­quiries from its own web­site. KrebsOnSecurity sought com­ment from the Fengwo Group by email­ing the con­tact ad­dress listed on the com­pa­ny’s home­page, but the re­quest bounced back with the re­ply, Your mes­sage could­n’t be de­liv­ered to post­mas­ter@fwg­cloud[.]com. Their in­box is full, or it’s get­ting too much mail right now.”

As Bitsight’s analy­sis shows, when it comes to TV boxes and stream­ing sticks, it’s best to stick to name brands from rep­utable man­u­fac­tur­ers, and then to be spar­ing and care­ful with any apps you choose to in­stall on the de­vice — as many of those can bun­dle res­i­den­tial proxy soft­ware as well. Google says con­sumers can con­firm whether or not a de­vice is built with the of­fi­cial Android TV OS and Play Protect cer­ti­fi­ca­tion by fol­low­ing these in­struc­tions.

Additionally, Synthient main­tains a run­ning list of IoT de­vices that have been known to ship to con­sumers with res­i­den­tial proxy soft­ware and other ma­li­cious apps pre-in­stalled. Careful read­ers will no­tice Synthient’s list in­cludes other IoT de­vices apart from stream­ing sticks and boxes: As the FBI has warned, res­i­den­tial proxy soft­ware has also been found in other pop­u­lar con­sumer IoT de­vices from ran­dom brands, par­tic­u­larly dig­i­tal photo frames.

AI-Generated Images Discourage Me From Reading Your Blog

nelson.cloud

I have a grow­ing ha­tred for AI-generated im­ages in blogs. It makes me won­der if the text in the blog posts is AI-generated to some ex­tent. It’s al­ways dis­ap­point­ing see­ing these im­ages in blogs run by in­di­vid­u­als. I ex­pect this from cor­po­rate blogs but not in­die blogs.

I’d rather see a shitty Microsoft Paint draw­ing as op­posed to some AI im­age.

I know there are plenty of things you can roast my blog for but at least you know for a fact you’re get­ting the thoughts of a real hu­man be­ing and not some LLM.

If you run a per­sonal blog, please avoid AI-generated im­ages.

Discussion over at Hacker News

Stacked pull requests are now in public preview

github.blog

Stacked pull re­quests break large changes into small, re­view­able pull re­quests. They’re an or­dered se­ries of pull re­quests that each rep­re­sent fo­cused lay­ers of your change. With stacks, you can in­de­pen­dently re­view and check each pull re­quest, then merge every­thing to­gether in one click. No more open­ing a sin­gle large pull re­quest that takes for­ever to re­view, or split­ting work across mul­ti­ple branches you have to keep man­u­ally re­bas­ing.

We’ve been us­ing GitHub stacked PRs for Next.js for the past few months. It has helped us in­tro­duce smaller in­di­vid­ual changes while ship­ping larger fea­tures, mak­ing it eas­ier to re­view PRs. — Tim Neutkens, NextJS lead, Vercel”

We’ve been us­ing GitHub stacked PRs for Next.js for the past few months. It has helped us in­tro­duce smaller in­di­vid­ual changes while ship­ping larger fea­tures, mak­ing it eas­ier to re­view PRs. — Tim Neutkens, NextJS lead, Vercel”

With stacked pull re­quests, teams can:

Keep large changes mov­ing by re­view­ing short, nar­rowly scoped pull re­quests in par­al­lel.

Maintain qual­ity across every layer by us­ing fo­cused pull re­quest re­views along­side ex­ist­ing branch pro­tec­tions to pro­tect main.

Merge one, some, or all by land­ing an en­tire stack al­to­gether or in­di­vid­ual lay­ers one at a time.

And be­cause stacked pull re­quests are built into GitHub, your ex­ist­ing re­views, checks, and merge re­quire­ments all work out of the box.

The new Github Stacked PRs pre­view is in­cred­i­ble. Landing 5 stacked PRs di­rectly to a merge queue all at once! A+++! This re­moves so much fric­tion (and the gh cli tools + agent skill help a ton)” — John Resig, cre­ator, jQuery

The new Github Stacked PRs pre­view is in­cred­i­ble. Landing 5 stacked PRs di­rectly to a merge queue all at once! A+++! This re­moves so much fric­tion (and the gh cli tools + agent skill help a ton)” — John Resig, cre­ator, jQuery

Get started with the CLI ex­ten­sion

Install the CLI ex­ten­sion and cre­ate your first stack in un­der a minute:

gh ex­ten­sion in­stall github/​gh-stack

Create stacks from your ter­mi­nal or github.com

Work with stacks on github.com, the GitHub CLI, the GitHub mo­bile app, or with a cod­ing agent such as GitHub Copilot us­ing the gh-stack skill. Start with a branch and pull re­quest for your first change. Then add branches and pull re­quests on top of it; each pull re­quest tar­gets the layer be­low it.

Review each layer in­de­pen­dently

Open any pull re­quest in the stack to re­view only the diff for that spe­cific layer. Use the stack map at the top of the pull re­quest to see how the change you’re re­view­ing fits into the larger work. You and your team­mates can each re­view dif­fer­ent lay­ers in par­al­lel with­out block­ing fur­ther work.

AI has made TEDs de­vel­op­ers dra­mat­i­cally more pro­duc­tive, but that cre­ated a new bot­tle­neck: PRs were grow­ing large enough that re­view­ers were strug­gling. Stacked PRs help to solve that. By break­ing large changes into small, de­pen­dency-or­dered pieces, re­view hap­pens in smaller log­i­cal chunks — not just faster PR re­views, but more ac­cu­rate ones. Stacked PRs tighten our feed­back loop and help get sta­ble code to ted.com faster.” — Andy Merryman, CTO, TED

AI has made TEDs de­vel­op­ers dra­mat­i­cally more pro­duc­tive, but that cre­ated a new bot­tle­neck: PRs were grow­ing large enough that re­view­ers were strug­gling. Stacked PRs help to solve that. By break­ing large changes into small, de­pen­dency-or­dered pieces, re­view hap­pens in smaller log­i­cal chunks — not just faster PR re­views, but more ac­cu­rate ones. Stacked PRs tighten our feed­back loop and help get sta­ble code to ted.com faster.” — Andy Merryman, CTO, TED

Merge every­thing in a sin­gle click

Merge the lat­est ready pull re­quest to land it and every un­merged layer be­low it in one sin­gle op­er­a­tion. To land part of a stack, merge one or more lower lay­ers—the pull re­quests above it stay open and au­to­mat­i­cally re­base and re­tar­get. Your ex­ist­ing branch pro­tec­tions and re­quired checks still gov­ern what reaches main.

A big change used to mean one gi­ant PR no­body wanted to re­view. Now it’s a stack of small ones re­view­ers can ac­tu­ally fol­low, and the whole stack merges in one shot. It stopped feel­ing like a tool on top of GitHub and started feel­ing like GitHub.” — Mayank Saini, con­nec­tiv­ity en­gi­neer, WHOOP

A big change used to mean one gi­ant PR no­body wanted to re­view. Now it’s a stack of small ones re­view­ers can ac­tu­ally fol­low, and the whole stack merges in one shot. It stopped feel­ing like a tool on top of GitHub and started feel­ing like GitHub.” — Mayank Saini, con­nec­tiv­ity en­gi­neer, WHOOP

Stacked pull re­quests are rolling out in pub­lic pre­view to all repos­i­to­ries over the com­ing days. Merge queue sup­port for stacked pull re­quests is rolling out pro­gres­sively over the com­ing weeks.

For more in­for­ma­tion, check out the stacked pull re­quests doc­u­men­ta­tion, and share your feed­back with us in the stacks dis­cus­sion.

The Session You Cannot Take With You | EARENDIL

earendil.com

The orig­i­nal promise of an in­fer­ence API was won­der­fully sim­ple: send some in­put, re­ceive some out­put. If you kept both, you had the con­ver­sa­tion. You could in­spect it, archive it, re­play it, or give it to a dif­fer­ent model.

That ab­strac­tion was never com­pletely true. For in­stance prompt caches live on some­body else’s GPUs, to­k­eniza­tion dif­fers be­tween mod­els, and sam­pling is not re­pro­ducible (and quite in­ten­tion­ally so). But the se­man­tic record of a ses­sion in the form of a tran­script could still be­long to the user. A tran­script should con­tain the in­struc­tions, mes­sages, tool calls and tool re­sults. Another suf­fi­ciently ca­pa­ble model might not con­tinue iden­ti­cally, but it could un­der­stand what hap­pened and take over.

Inference APIs are frus­trat­ingly mov­ing away from that prop­erty, at least some­what. They in­creas­ingly re­turn a mix­ture of text and provider-bound state that is very in­ten­tion­ally non-portable.

rea­son­ing to­kens that are billed to the user but re­turned only as opaque, en­crypted blobs, with use­less sum­maries at best

web searches where the model sees source ma­te­r­ial the client never sees

com­pacted con­text that only the orig­i­nal provider can de­crypt

sub­agent in­struc­tions and mes­sages hid­den from the ap­pli­ca­tion run­ning the agents in the form of en­crypted pay­loads

file, vec­tor-store, con­tainer, and cache ref­er­ences that can­not be re­solved any­where else.

re­sponse and con­ver­sa­tion state that is en­tirely keyed by IDs that are stored fully on the provider’s servers

Each fea­ture comes with a ba­sic jus­ti­fi­ca­tion that’s triv­ial for a provider to come up with, along with good ar­gu­ments for why this is good for the user. Together all of these things change the own­er­ship re­al­ity of an AI ses­sion: the tran­script on your ma­chine is no longer your ses­sion but a par­tial view of a ses­sion whose op­er­a­tional state be­longs to an in­fer­ence provider and not you.

We are not fans of this di­rec­tion, and we want to talk a bit about what it means to you, as a user, and what it means to us, as peo­ple de­vel­op­ing tools in this space.

A Practical Test for Session Ownership

By a portable ses­sion we do not mean that switch­ing from one model to an­other must pro­duce the same next to­ken. That’s a given be­cause mod­els have dif­fer­ent ca­pa­bil­i­ties, trained per­son­al­i­ties, con­text win­dows, and ways of work­ing with tools. And well, it’s all quite non­de­ter­min­is­tic any­way. Portability means some­thing more mod­est:

const tran­script = ses­sion.ex­port(); re­voke­Cre­den­tials(old­Provider); ses­sion = new­Provider.con­tin­ue­From(tran­script);

The archive should con­tain enough in­tel­li­gi­ble in­for­ma­tion for an­other model to con­tinue the work. It should not re­quire the old provider to deref­er­ence an ID, de­crypt a blob, re­mem­ber a search re­sult, or re­con­struct a sum­mary.

This gives us five use­ful tests:

Inspection: Can the user see what the model saw, what tools did, and what agents told each other?

Export: Is the ses­sion self-con­tained, apart from or­di­nary ar­ti­facts that can also be down­loaded?

Replay: Can an­other im­ple­men­ta­tion re­con­struct a se­man­ti­cally equiv­a­lent con­text?

Audit: Can a hu­man ex­plain why the sys­tem took an ac­tion af­ter the fact?

Deletion: Can the user iden­tify and re­move every server-side copy on which the ses­sion de­pends?

A re­sponse ID is not a tran­script (as the data is stored on the server), a ci­pher­text is not user-con­trolled stated (as the user can­not de­crypt it), a list of ci­ta­tions is not the ev­i­dence that was placed in the mod­el’s con­text by a search re­sult (as you can­not typ­i­cally fetch the same data as the model did).

Encryption for Whom?

The nam­ing and mar­ket­ing around these fea­tures can be mis­lead­ing. en­crypt­ed_­con­tent sounds like a pri­vacy fea­ture un­der the user’s con­trol. Usually it is a cap­sule that the client can­not read and only the provider can open. The provider chooses the keys, de­crypts the con­tent for its own mod­els, and de­fines where the data can be re­played.

A bet­ter term is provider-sealed state.

Provider seal­ing can have a real pri­vacy ben­e­fit. OpenAI, for ex­am­ple, can re­turn en­crypted rea­son­ing to a client us­ing store: false, then de­crypt it in mem­ory on the next re­quest with­out per­sist­ing the in­ter­me­di­ate state. That is bet­ter than re­quir­ing server-side con­ver­sa­tion stor­age, par­tic­u­larly for Zero Data Retention cus­tomers. But, re­mem­ber, there is not re­ally any­thing that needs en­cryp­tion to be­gin with!

This en­cryp­tion does not hide the data from the in­fer­ence provider but it hides it from you.

Stored Conversations Turn a Transcript into a Pointer

OpenAI’s Responses API stores re­sponses by de­fault. Its doc­u­men­ta­tion says re­sponse ob­jects are re­tained for at least 30 days by de­fault. store: false is avail­able and should be used, as it makes it work more like com­ple­tions: the data is not stored on OpenAI’s servers.

The new Gemini Interactions API has made a sim­i­lar choice. It de­faults to store: true. On the paid tier in­ter­ac­tions are re­tained for 55 days, and on the free tier for one day.

And ob­vi­ously, the idea of stor­ing state on the server is quite at­trac­tive:

const first = re­sponses.cre­ate({ model: frontier-model”, in­put: Investigate this pro­duc­tion fail­ure”, store: true, });

const sec­ond = re­sponses.cre­ate({ model: frontier-model”, pre­vi­ous­Re­spon­seId: first.id, in­put: Now im­ple­ment the fix”, store: true, });

The ap­pli­ca­tion sends less data, the provider can pre­serve hid­den rea­son­ing and tool state, and cache rout­ing be­comes eas­ier. But if the lo­cal ap­pli­ca­tion only records the user mes­sages and fi­nal text, first.id is now a for­eign key into a data­base it does not con­trol.

No Reasoning For You

All ma­jor labs claim to have le­git­i­mate rea­sons not to ex­pose raw chain of thought. As a re­sult, on non-open-weights mod­els we typ­i­cally do not see these to­kens.

Raw rea­son­ing is not vis­i­ble via the API. With stored re­sponses, prior rea­son­ing can be re­cov­ered through pre­vi­ous_re­sponse_id. With store: false, the API re­turns en­crypt­ed_­con­tent, which the client must pre­serve and re­play. Persisted rea­son­ing re­mains opaque even when rea­son­ing.con­text: all_turns” lets a later sam­ple use it.

Anthropic re­turns the en­crypted full think­ing in a sig­na­ture field. The read­able think­ing text, when en­abled, is a sum­mary pro­duced by an­other model, not the raw chain of thought. Thinking blocks must be passed back un­changed dur­ing tool-use turns. Anthropic’s doc­u­men­ta­tion also says think­ing blocks are tied to the model that pro­duced them and should be stripped when switch­ing mod­els. So these rea­son­ing traces do not at­tempt to be portable within Anthropic.

The same story re­peats with all closed-weights mod­els.

These en­cryp­tion mech­a­nisms per­mit con­ti­nu­ity in­side an ecosys­tem but they do not cre­ate a portable tran­script that can be taken to an­other provider’s model. A ses­sion archive can con­tain the blob, but an­other model can­not use its mean­ing:

{“type”: reasoning”, encrypted_content”: gAAAAAB…“} {“type”: thinking”, thinking”: ”, signature”: EqQBCg…“} {“type”: thought”, summary”: [], signature”: EpoGCp…“}

Hidden Searches

Server-side web search is one of the clear­est ex­am­ples of a tran­script hav­ing holes in it hid­den from the user. A client-side search tool be­haves like any other tool:

const re­sult = search(query); record({ query, re­trieve­dAt: now(), re­sults: re­sult.map((item) => ({ url: item.url, ti­tle: item.ti­tle, pas­sages: item.pas­sages, })), }); model.send({ tool­Re­sult: re­sult });

The user can in­spect the rank­ing and pas­sages, refetch the pages, cache a copy, or pro­vide the same ev­i­dence to an­other model.

With hosted search, the provider per­forms a pri­vate tool loop. OpenAI, Google and Anthropic ex­pose search ac­tions, ci­ta­tions, and op­tion­ally a list of source URLs, but not the com­plete text con­text used to pro­duce an an­swer. A URL is not a sta­ble re­play, in­stead its con­tents can change or have been re­duced to a much shorter snip­pet be­fore the model saw it.

The fi­nal an­swer may be per­fectly good. The prob­lem ap­pears on the next turn:

Compare the third source with the first one, re-check the dis­puted num­ber, and con­tinue this re­search us­ing an­other model.

Compare the third source with the first one, re-check the dis­puted num­ber, and con­tinue this re­search us­ing an­other model.

The new model re­ceives an an­swer and a few URLs. It does not re­ceive the re­sult rank­ing, ex­tracted pas­sages, fil­tered-out ma­te­r­ial, or ex­act ev­i­dence the first model used. The old provider is still part of the ses­sion even if the next re­quest goes else­where. Even if you have the ci­ta­tions and you were to re-fetch you can­not re­pro­duce the pre­cise data.

Hosted search should have a full-fi­delity ex­port mode con­tain­ing queries, re­sult meta­data, re­trieved pas­sages, time­stamps and re­tained con­tents. Concise ci­ta­tions can re­main the user in­ter­face but they should not be the only record.

Opaque Compaction

Long agent ses­sions even­tu­ally need com­paction. A vis­i­ble, client-con­trolled sum­mary is lossy, but it is at least in­spectable and trans­fer­able. The user can re­view it, edit it, or ask a dif­fer­ent model to pro­duce an­other one.

OpenAI’s server-side com­paction in­stead emits an en­crypted com­paction item. The doc­u­men­ta­tion de­scribes it as opaque and not in­tended to be hu­man-in­ter­pretable.” The stand­alone /responses/compact end­point re­turns a canonical next con­text win­dow” that clients are in­structed to pass on as-is.

Conceptually, the tran­si­tion looks like this:

// Before: ex­pen­sive but portable let his­tory = [ user­Mes­sage, as­sis­tantMes­sage, tool­Call, full­Tool­Re­sult, // … 200,000 more to­kens of in­tel­li­gi­ble his­tory ];

// After: cheap to con­tinue only with the orig­i­nal provider his­tory = [ { type: compaction”, en­crypt­ed­Con­tent: enc_provider_only_state…”, }, …recentItems, ];

OpenAI can con­tinue from the com­pressed mean­ing, but a dif­fer­ent provider sees an un­read­able string plus a re­cent suf­fix (well, would see it, we never pass this sort of in­for­ma­tion to an­other provider).

This is not tech­ni­cally nec­es­sary. Anthropic’s server-side com­paction re­turns a com­paction block with a read­able con­tent field. It lets the client pro­vide cus­tom sum­ma­riza­tion in­struc­tions, and the re­sult­ing sum­mary can be in­spected and passed to an­other model. Client-side com­paction is also pos­si­ble with any provider.

OpenAI’s sealed ar­ti­fact may pre­serve more model-spe­cific state than a plain sum­mary and may per­form bet­ter on the orig­i­nal model. That is a rea­son­able op­tional op­ti­miza­tion but it should be ac­com­pa­nied by a read­able hand­off sum­mary, not re­place one. But again, a lot of this has the added ben­e­fit of fur­ther lock­ing you into one ecosys­tem.

Subagents Come With Hidden Instructions

Multi-agent sys­tems com­pound the prob­lem be­cause there is no longer one tran­script. There is a tree of ses­sions and a stream of mes­sages be­tween them. Usually they are prompts as if a hu­man wrote them, just now au­thored by a ma­chine for an­other ma­chine.

OpenAI’s hosted Responses Multi-agent beta re­turns three new item types: mul­ti­_a­gen­t_­call, mul­ti­_a­gen­t_­cal­l_out­put, and agen­t_mes­sage. The ex­am­ple for spawn_a­gent con­tains an en­crypted mes­sage ar­gu­ment, and in­ter-agent mes­sages con­tain only en­crypt­ed_­con­tent. Automatic server-side com­paction is im­plic­itly en­abled for every agent when Multi-agent is en­abled, even if the client did not re­quest it. Reasoning sum­maries are not sup­ported and the API also in­jects root and sub­agent in­struc­tions that the de­vel­oper can­not edit or re­move.

This is a bun­dle of non-trans­fer­able state: sealed del­e­ga­tion, sealed agent mes­sages, sep­a­rate au­to­mat­i­cally com­pacted con­texts, hid­den rea­son­ing, and provider-hosted or­ches­tra­tion.

A re­lated change landed in the open-source Codex client in June 2026. The com­mit, ti­tled Encrypt multi-agent v2 mes­sage pay­loads”, ex­plains the flow di­rectly:

// Parent mod­el’s tool call, as per­sisted by Codex { name”: spawn_agent”, arguments”: { task_name”: worker”, message”: <ciphertext>” } }

// Child mod­el’s in­put { type”: agent_message”, author”: /root”, recipient”: /root/worker”, content”: [{ type”: encrypted_content”, encrypted_content”: <ciphertext>” }] }

The Responses API en­crypts the tool ar­gu­ment emit­ted by the par­ent, Codex for­wards it, and the API de­crypts it in­ter­nally for the child. Codex’s own InterAgentCommunication.content is empty. The ex­act task is ab­sent from its read­able roll­out and his­tory.

Presumably this is not merely an ab­stract model-switch­ing con­cern. One could imag­ine if the child changes the wrong file, leaks a se­cret, du­pli­cates an­other agen­t’s work, or fol­lows a bad as­sump­tion, the user can­not an­swer the sim­ple ques­tion of what was that agent asked to do?

An open Codex is­sue asks for the en­crypted de­liv­ery to re­tain a sep­a­rate read­able au­dit copy. That is the min­i­mum ac­cept­able de­sign. Better still, plain­text in­ter-agent mes­sages should re­main the norm.

Most People Do Not Switch Models Mid-Session”

Probably not. Most peo­ple do not switch their op­er­at­ing sys­tem or phone provider every week ei­ther. But even if you do not uti­lize that free­dom, it mat­ters be­cause it changes the re­la­tion­ship you have with the provider and the provider has with you.

As a user you also may need to move a ses­sion be­cause a model is re­tired, a ser­vice is down, a price changes, a pol­icy blocks the next re­quest (hello fa­ble), a con­fi­den­tial phase must run lo­cally, or an au­di­tor needs to re­con­struct what hap­pened. Agents are also mak­ing ses­sions much longer. A cod­ing or re­search ses­sion can ac­cu­mu­late days of de­ci­sions and ev­i­dence and a per­sonal as­sis­tant may ac­cu­mu­late ses­sion tran­scripts go­ing back years (presumably as we don’t have them for that long yet).

The op­tion to leave also cre­ates dis­ci­pline. If a provider knows that a user can con­tinue else­where, it has to com­pete on model qual­ity, price, re­li­a­bil­ity, and trust. If the user’s ac­cu­mu­lated con­text can only be in­ter­preted by one provider, it sets very un­for­tu­nate in­cen­tives.

What a Portable Inference API Should Promise

We would like in­fer­ence providers and agent builders to adopt a small set of rules.

The lo­cal event log is canon­i­cal. Server stor­age may mir­ror or ac­cel­er­ate it, but the client can re­con­struct the ses­sion with­out deref­er­enc­ing server IDs.

Storage is ex­plicit. store: false should be easy, doc­u­mented, and prefer­ably the de­fault. Features that re­quire re­ten­tion should say so at the point of use.

No opaque item is the sole car­rier of mean­ing. Encrypted rea­son­ing, com­paction, and tool sig­na­tures may be in­cluded for same-provider qual­ity, but each has a read­able, provider-neu­tral hand­off rep­re­sen­ta­tion.

Hosted tools have full-fi­delity logs. Record ex­act in­puts, out­puts, ev­i­dence, fil­ter­ing, prove­nance, time­stamps, and con­tent hashes — not only a pol­ished an­swer and ci­ta­tions.

Subagent com­mu­ni­ca­tion is au­ditable. Persist the ex­act read­able task, mes­sages, re­sults, lin­eage, model, and tool per­mis­sions for every agent.

Compaction is in­spectable. Return a read­able sum­mary, the in­struc­tions used to cre­ate it, and enough lin­eage to un­der­stand what was dis­carded.

Artifacts are ex­portable. Files, con­tainer out­puts, search snap­shots, and gen­er­ated me­dia can be down­loaded into a con­tent-ad­dressed lo­cal archive.

Distillation Is Great Actually

There is a re­lated form of lock-in at the model layer.

Some of the largest closed-weight US labs are in­creas­ingly hos­tile to out­side dis­til­la­tion. Anthropic’s February 2026 post about al­leged cam­paigns by DeepSeek, Moonshot, and MiniMax calls them distillation at­tacks”. Its com­mer­cial terms say cus­tomers own their out­puts, but pro­hibit us­ing the ser­vice to train a com­pet­ing AI model. At the same time, Anthropic’s own post ac­knowl­edges that distillation is a widely used and le­git­i­mate train­ing method” when fron­tier labs use it on their own mod­els.

Anthropic uses ro­bots to gather data from the pub­lic web for model de­vel­op­ment and they fa­mously cut up books to scan them. OpenAI sim­i­larly says it trains on freely ac­ces­si­ble pub­lic in­ter­net con­tent and has ar­gued that train­ing on pub­licly avail­able in­ter­net ma­te­ri­als is fair use. Both com­pa­nies de­scribe dis­til­la­tion as a nor­mal way to pro­duce smaller mod­els when it hap­pens in­side their own walls. OpenAI has also of­fered an ex­plicit first-party API dis­til­la­tion work­flow for us­ing out­puts from a stronger OpenAI model to fine-tune a smaller OpenAI model.

The moral asym­me­try is still hard to miss. The labs ask so­ci­ety to ac­cept that ma­chines may learn from the enor­mous body of work hu­mans placed on the in­ter­net — of­ten with­out ad­vance, in­di­vid­ual per­mis­sion — while in­sist­ing that other ma­chines must not learn from out­puts the labs gen­er­ate. The broad­est ver­sion of that prin­ci­ple con­ve­niently al­lows learn­ing to flow into closed mod­els but not back out of them.

We think the de­fault at­ti­tude to­ward dis­til­la­tion should move from hos­til­ity to sup­port. Distillation can turn ex­pen­sive fron­tier ca­pa­bil­ity into smaller, cheaper, faster mod­els that can run lo­cally, of­fline, on con­strained hard­ware, or un­der the user’s con­trol. It can in­crease com­pe­ti­tion, pre­serve ca­pa­bil­ity when an API dis­ap­pears, and re­duce the com­pute and en­ergy re­quired for com­mon tasks.

The Minimum Freedom

A user should be able to close an ac­count, keep a ses­sion, and hand it to an­other model. The new model may dis­agree, ask ques­tions, or per­form worse. It should not be star­ing at ci­pher­text where the old model saw the user’s his­tory, ev­i­dence, plans, and del­e­gated work.

We do not ob­ject to providers build­ing bet­ter state­ful APIs. We ob­ject to bet­ter per­for­mance be­ing cou­pled to less user con­trol. Stateful stor­age should be op­tional, hosted tools should be ob­serv­able, com­paction should be read­able, agent com­mu­ni­ca­tion should be au­ditable and ide­ally opaque rea­son­ing is not opaque or at least should have a portable hand­off. Distillation should be a path by which ca­pa­bil­ity be­comes more avail­able, not a taboo used to jus­tify ever higher walls.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.