10 interesting stories served every morning and every evening.

DuckLabs to Join AWS, Projects to Remain Open Source

ducklabs.com

2026 – 08-26 | 10 min

Today, we’re an­nounc­ing that DuckLabs will join Amazon Web Services (AWS), which is ex­pected to be ef­fec­tive in early September.

Our team will re­main to­gether in Amsterdam, con­tin­u­ing our work on DuckDB, DuckLake, Quack, and the broader com­mu­nity. Joining AWS gives us the re­sources and reach to bring this tech­nol­ogy to many more de­vel­op­ers and or­ga­ni­za­tions, and to pur­sue ideas at a scale that would have been dif­fi­cult for us to reach alone.

Most im­por­tantly, the foun­da­tions of the pro­ject will re­main firmly in place. DuckDB and the other open-source com­po­nents of the Duck Stack” will re­main free and open source un­der the MIT li­cense, with the non­profit DuckDB Foundation con­tin­u­ing its stew­ard­ship of the pro­jects.

This is a sig­nif­i­cant mo­ment for all of us at DuckLabs. It marks the end of one chap­ter that we are im­mensely proud of, and the be­gin­ning of an­other that we be­lieve can take DuckDB much fur­ther.

Our Journey to This Point

We founded DuckLabs a lit­tle over five years ago to give the team be­hind DuckDB a sta­ble, long-term home.

At the time, DuckDB was be­gin­ning to gain real mo­men­tum. The first com­mer­cial con­tracts to pri­or­i­tize fea­tures were ma­te­ri­al­iz­ing, and ven­ture cap­i­tal firms were call­ing. We chose a dif­fer­ent path: a boot­strapped com­pany, fully owned by its founders and de­vel­op­ment team.

That de­ci­sion shaped DuckLabs in ways we value deeply. It gave us the free­dom to build pa­tiently, to put the tech­nol­ogy first, and to grow with­out los­ing sight of why we started. From a small group gath­ered around an am­bi­tious open-source pro­ject, we grew into a team of more than 30 peo­ple in Amsterdam, all while con­tin­u­ing to in­vest in DuckDB and the com­mu­nity tak­ing shape around it.

During those years, DuckDB trav­eled much fur­ther than we imag­ined when the pro­ject be­gan. Today, we see more than one mil­lion down­loads every day. Developers around the world use it to ex­plore data, power prod­ucts, teach new ideas, con­duct re­search, and build sys­tems we could never have an­tic­i­pated our­selves.

Watching that hap­pen has been one of the great priv­i­leges of our pro­fes­sional lives. But the lim­its of our ex­ist­ing model also be­came in­creas­ingly clear. As founders, we wor­ried that DuckDB’s growth would even­tu­ally out­pace our abil­ity to sup­port it. That our small com­pany could be­come a bot­tle­neck for the pro­ject, the team, and the peo­ple build­ing busi­nesses on top of it.

We also wor­ried that scal­ing DuckLabs into a much larger sales, sup­port, and op­er­a­tions or­ga­ni­za­tion would pull our at­ten­tion away from the tech­ni­cal work and open-source com­mu­nity that made DuckDB suc­cess­ful in the first place.

Our part­ner­ships work best with highly tech­ni­cal or­ga­ni­za­tions — of­ten com­pa­nies with sub­stan­tial data­base ex­per­tise of their own. Reaching a much broader group of users re­quires us to solve more com­plete and spe­cial­ized prob­lems, serve the needs of dif­fer­ent in­dus­tries, in­vest sub­stan­tially more in in­fra­struc­ture, and reach peo­ple who may never think to seek out an an­a­lyt­i­cal data­base di­rectly.

We be­lieve the DuckDB rev­o­lu­tion can grow by an­other cou­ple of or­ders of mag­ni­tude. To give it that op­por­tu­nity, we re­al­ized we needed a dif­fer­ent setup.

Why AWS

DuckLabs and AWS have al­ready been work­ing closely to­gether for more than a year. During that time, we came to un­der­stand how our teams col­lab­o­rate, what each brings to the table, and what might be­come pos­si­ble by com­bin­ing DuckLabs’ tech­ni­cal ex­per­tise with AWSs in­fra­struc­ture, scale, and cus­tomer reach.

That ex­pe­ri­ence gave us con­fi­dence in tak­ing this next step.

Together, we plan to use DuckDB, DuckLake, and Quack to help power a new gen­er­a­tion of data ser­vices. Our am­bi­tion is to reach peo­ple who may even­tu­ally de­pend on the Duck Stack every day, whether they in­ter­act with it di­rectly or en­counter it qui­etly in­side the prod­ucts and ser­vices they use.

For the DuckLabs team, join­ing AWS cre­ates the space to con­cen­trate on the tech­ni­cal work we care about most while op­er­at­ing at a scale we could not eas­ily achieve on our own. It al­lows us to think fur­ther ahead, be bolder, tackle harder prob­lems, and bring the ideas be­hind DuckDB to many more peo­ple.

AWS has com­mit­ted to sup­port­ing the con­tin­ued de­vel­op­ment of DuckDB and its wider com­mu­nity for the long term. That com­mit­ment mat­ters deeply to us.

Quote DuckDB is an in­cred­i­ble open source pro­ject with an amaz­ing com­mu­nity; it is broadly used and very much loved by S3 cus­tomers to­day. After about two years of work­ing closely with Mark, Hannes and the whole team at DuckLabs I’m ex­cited at the op­por­tu­nity to help the pro­ject have an even broader im­pact. Also, and maybe a lit­tle self­ishly, I’ve found the DuckLabs team to be one of the most tech­ni­cally deep, hum­ble, and high-ve­loc­ity teams that I’ve ever had a chance to work with and I’m de­lighted that we get to do a lot more of that.”

– Andy Warfield (Distinguished Engineer and Vice President, AWS)

Quote DuckDB is an in­cred­i­ble open source pro­ject with an amaz­ing com­mu­nity; it is broadly used and very much loved by S3 cus­tomers to­day. After about two years of work­ing closely with Mark, Hannes and the whole team at DuckLabs I’m ex­cited at the op­por­tu­nity to help the pro­ject have an even broader im­pact. Also, and maybe a lit­tle self­ishly, I’ve found the DuckLabs team to be one of the most tech­ni­cally deep, hum­ble, and high-ve­loc­ity teams that I’ve ever had a chance to work with and I’m de­lighted that we get to do a lot more of that.”

– Andy Warfield (Distinguished Engineer and Vice President, AWS)

What Will Remain the Same

We know that an an­nounce­ment like this brings im­por­tant — and un­der­stand­able — ques­tions for users, con­trib­u­tors, cus­tomers, and part­ners.

DuckDB, DuckLake, Quack, and the other open-source com­po­nents of the Duck Stack will re­main free and open source un­der the MIT li­cense. The non­profit DuckDB Foundation will con­tinue to stew­ard these pro­jects, and the DuckLabs team will con­tinue to con­tribute to the pro­ject and re­main to­gether in Amsterdam.

The open-source pro­jects will con­tinue to serve a broad com­mu­nity of users, con­trib­u­tors, plat­forms, and ven­dors. The open­ness that al­lowed DuckDB to flour­ish will re­main cen­tral to its fu­ture.

We care deeply about the trust this com­mu­nity has placed in us. Protecting that trust has been one of the most im­por­tant con­sid­er­a­tions through­out this process.

Quote As the CWI rep­re­sen­ta­tive on the DuckDB Foundation, I would like to con­grat­u­late Hannes, Mark and all DuckLabs em­ploy­ees with this new chap­ter. When DuckLabs spun out of CWI, we cre­ated this foun­da­tion, which holds all IP of open-source DuckDB, and will con­tinue to do so. I am de­lighted that AWS is com­mit­ted to keep ad­vanc­ing open-source DuckDB, and I an­tic­i­pate it to even ac­cel­er­ate its in­no­va­tion. Via the DuckDB Foundation we will also make sure that the voices of its sup­port­ers and all mem­bers of the open-source DuckDB com­mu­nity at large, will con­tinue to be heard. I am very proud of DuckDB and its cre­ators; this de­vel­op­ment un­der­lines it is state-of-the-art tech­nol­ogy, which re­al­izes many ideas from the data­base ar­chi­tec­ture re­search group in open source, for every­body’s ben­e­fit.”

– Peter Boncz (Database Architectures group lead at CWI Amsterdam, DuckDB Foundation board mem­ber)

Quote As the CWI rep­re­sen­ta­tive on the DuckDB Foundation, I would like to con­grat­u­late Hannes, Mark and all DuckLabs em­ploy­ees with this new chap­ter. When DuckLabs spun out of CWI, we cre­ated this foun­da­tion, which holds all IP of open-source DuckDB, and will con­tinue to do so. I am de­lighted that AWS is com­mit­ted to keep ad­vanc­ing open-source DuckDB, and I an­tic­i­pate it to even ac­cel­er­ate its in­no­va­tion. Via the DuckDB Foundation we will also make sure that the voices of its sup­port­ers and all mem­bers of the open-source DuckDB com­mu­nity at large, will con­tinue to be heard. I am very proud of DuckDB and its cre­ators; this de­vel­op­ment un­der­lines it is state-of-the-art tech­nol­ogy, which re­al­izes many ideas from the data­base ar­chi­tec­ture re­search group in open source, for every­body’s ben­e­fit.”

– Peter Boncz (Database Architectures group lead at CWI Amsterdam, DuckDB Foundation board mem­ber)

Quote Over the course of five years, DuckDB’s open na­ture, its sheer hack­a­bil­ity, and the re­mark­ably friendly com­mu­nity of de­vel­op­ers has turned the sys­tem into one of the pre­mier plat­forms for data­base re­search as well as teach­ing. For both, it is es­sen­tial that we can in­spect and tin­ker with DuckDB’s ker­nel. I am thus ex­cited to learn that DuckDB will re­main open source un­der the um­brella of AWS. An even wider range of op­por­tu­ni­ties is in reach now. We are glad to be able to be a part of this new era for DuckLabs and DuckDB.”

– Torsten Grust (Professor of Computer Science and Database Systems re­search group lead at Universität Tübingen, Germany)

Quote Over the course of five years, DuckDB’s open na­ture, its sheer hack­a­bil­ity, and the re­mark­ably friendly com­mu­nity of de­vel­op­ers has turned the sys­tem into one of the pre­mier plat­forms for data­base re­search as well as teach­ing. For both, it is es­sen­tial that we can in­spect and tin­ker with DuckDB’s ker­nel. I am thus ex­cited to learn that DuckDB will re­main open source un­der the um­brella of AWS. An even wider range of op­por­tu­ni­ties is in reach now. We are glad to be able to be a part of this new era for DuckLabs and DuckDB.”

– Torsten Grust (Professor of Computer Science and Database Systems re­search group lead at Universität Tübingen, Germany)

What We Plan to Expand

Joining AWS will give us greater ca­pac­ity to in­vest in both the tech­nol­ogy and the com­mu­nity around it. Going for­ward, the DuckDB Foundation will in­clude a tech­ni­cal ad­vi­sory board, so that lead­ing com­mu­nity mem­bers can pro­vide their in­put on the pro­jec­t’s tech­ni­cal di­rec­tion. We also plan to open the ex­ten­sion stack so that ex­ten­sions signed by other de­vel­op­ers and or­ga­ni­za­tions can run in DuckDB.

There is a great deal of work ahead, and many de­tails still to shape. We will share more as these plans de­velop.

Quote Amazon putting its weight be­hind DuckDB is go­ing to add a ton of mo­men­tum and strengthen the ecosys­tem. This is great news for those of us who be­lieve in DuckDB as the plat­form on which the fu­ture of an­a­lyt­ics is be­ing built.”

– Jordan Tigani (CEO, MotherDuck)

Quote Amazon putting its weight be­hind DuckDB is go­ing to add a ton of mo­men­tum and strengthen the ecosys­tem. This is great news for those of us who be­lieve in DuckDB as the plat­form on which the fu­ture of an­a­lyt­ics is be­ing built.”

– Jordan Tigani (CEO, MotherDuck)

Quote Amazon is the ideal home for DuckLabs. DuckDB is the cen­ter of the ven­dor-neu­tral open data stack, and Amazon has the com­mit­ment to open­ness, and the track record of work­ing with the en­tire cloud ecosys­tem, to en­able the DuckDB pro­ject to con­tinue to thrive in this role.”

– George Fraser (CEO and co-founder, Fivetran)

Quote Amazon is the ideal home for DuckLabs. DuckDB is the cen­ter of the ven­dor-neu­tral open data stack, and Amazon has the com­mit­ment to open­ness, and the track record of work­ing with the en­tire cloud ecosys­tem, to en­able the DuckDB pro­ject to con­tinue to thrive in this role.”

– George Fraser (CEO and co-founder, Fivetran)

The Next Chapter

DuckDB’s suc­cess has al­ways be­longed to a much larger com­mu­nity than the peo­ple work­ing in­side DuckLabs.

It be­longs to every­one who has used it, con­tributed code, re­ported a bug, writ­ten an ex­ten­sion, an­swered a ques­tion, pub­lished a bench­mark, taught a class, built a prod­uct, chal­lenged our as­sump­tions, or rec­om­mended DuckDB to some­one else. Every one of those acts helped the pro­ject be­come what it is to­day.

We do not take that sup­port — or the trust be­hind it — for granted.

Joining AWS gives the DuckLabs team an ex­tra­or­di­nary op­por­tu­nity to bring the Duck Stack to a much larger au­di­ence while con­tin­u­ing to in­vest in the open-source tech­nol­ogy at its heart. We en­ter this next chap­ter with the same cu­rios­ity, care, and tech­ni­cal am­bi­tion that brought us here, along­side a team that has been through the en­tire jour­ney to­gether.

To every­one who helped us reach this point: thank you. We are proud of what we have built to­gether, ex­cited by what now lies ahead, and look­ing for­ward to build­ing the next chap­ter with you.

You’re also wel­come to check out our press re­lease and the me­dia kit.

You’re also wel­come to check out our press re­lease and the me­dia kit.

z.ai

Qwen Studio

qwen.ai

Tim Curry, star of The Rocky Horror Picture Show and Stephen King’s It, dies aged 80

www.theguardian.com

Tim Curry, the ver­sa­tile per­former of screen and stage, who made his name as the flam­boy­ant Dr Frank-N-Furter in The Rocky Horror Show, has died aged 80. Reports say Curry died at his home in Los Angeles.

The ac­tor’s man­ager con­firmed to Variety that he died peace­fully. No cause of death is known as yet.

Ironically, Curry’s best-known roles all re­quired a con­sid­er­able phys­i­cal trans­for­ma­tion, to the ex­tent that he was un­recog­nis­able in all of them. Frank-N-Furter, Rocky Horror’s de­ranged sci­en­tist, saw him daubed in gar­ish makeup and wear­ing stock­ings and sus­penders, for the stage show (which de­buted in 1973) and the film adap­ta­tion (released in 1975). In the Ridley Scott-directed fan­tasy Legend (1985), he was equipped with gi­ant horns and scar­let skin as the Lord of Darkness. And in 1990, Curry was trans­formed with full-face clown pan­cake to play Pennywise in the TV minis­eries It, based on the Stephen King novel.

Alongside these show-stop­ping in­car­na­tions, Curry en­joyed a suc­cess­ful par­al­lel ca­reer in film, TV and the­atre. Born in 1946, he earned a de­gree in drama and English from the University of Birmingham in 1968 and em­barked on a stage ca­reer. His first sig­nif­i­cant role was in the orig­i­nal cast for the first West End pro­duc­tion of the con­tro­ver­sial mu­si­cal Hair, in 1968 — where he met Richard O’Brien, who would go on to cast him in his own mu­si­cal, The Rocky Horror Show. Curry went on to play Tristan Tzara in Tom Stoppard’s Travesties (taking over the role from John Hurt and Robert Powell), and Mozart in the Broadway trans­fer of Amadeus in 1980. Later high­lights in­cluded King Arthur in the Broadway run of the Monty Python mu­si­cal Spamalot.

Curry worked reg­u­larly in film, in roles in­clud­ing that of Robert Graves in Jerzy Skolimowski’s sin­is­ter The Shout, and later as Jeremy in the Ian McEwan-scripted po­lit­i­cal drama The Ploughman’s Lunch. Later, Curry found him­self work­ing largely in come­dies — in­clud­ing Home Alone 2: Lost in New York, Muppet Treasure Island, Scary Movie 2, and McHale’s Navy — be­fore pro­vid­ing voiceovers for nu­mer­ous an­i­mated films, such as Garfield: A Tail of Two Kitties, A Turtle’s Tale: Sammy’s Adventures and Ribbit.

Curry was also a reg­u­lar pres­ence on TV for more than four decades, play­ing William Shakespeare in John Mortimer’s six-part bi­og­ra­phy of the play­wright in 1978, Bill Sikes in the 1982 US TV adap­ta­tion of Oliver Twist, and the wiz­ard Trymon in the adap­ta­tion of Terry Pratchett’s The Colour of Magic in 2008. He voiced a stream of an­i­mated se­ries (Duckman, Sonic the Hedgehog, The Mask and Star Wars: The Clone Wars) and was cast in many guest roles in es­tab­lished se­ries, in­clud­ing Psych, Poirot and Criminal Minds.

I’ve had the op­por­tu­ni­ties,” he said to the Guardian in 2025, and I’m still show­ing up. You know? I think that’s what you have to do. You have to keep show­ing up.”

In 2012, Curry had a stroke while re­ceiv­ing a mas­sage and re­ceived brain surgery. He was left with mo­bil­ity is­sues and his short-term mem­ory was also af­fected. It was an odd thing, be­cause my fa­ther had a stroke and died very soon af­ter­wards,” he said. I knew I had to force my­self to re­lax and just take the op­por­tu­nity to float a lit­tle.”

In 2025, he re­leased his mem­oir Vagabond.

I’m very aware that I’m lucky,” Curry said last year. I’m as­ton­ished ac­tu­ally at how am­bi­tious I’ve been. I did­n’t think of my­self as am­bi­tious at all.”

Luke Evans, who played the role of Dr Frank-N-Furter in the re­cent Broadway pro­duc­tion of The Rocky Horror Show, paid trib­ute on Instagram. The ac­tor called him a force, a bright fierce flame” and an in­spi­ra­tion”. He added: There will only be one Tim Curry.”

Rocky Horror Picture Show co-star Susan Sarandon called him a funny sexy orig­i­nal” in an on­line post. Even wheel­chair bound he con­tin­ued to be so sweet and gen­er­ous with his fans,” she wrote. Thank you Tim for mean­ing so much to so many peo­ple.”

Carol Burnett, who starred with Curry in Annie, also paid trib­ute on­line. Nobody could play lov­able vil­lains bet­ter than he could,” she wrote. He was a dear friend. I was blessed to know him.”

Curry’s Clue co-star Michael McKean shared on X: My last con­ver­sa­tion with Tim Curry was about life and love and laugh­ter. The grim phys­i­cal state he found him­self in has now ended, a bless­ing. Our bless­ing was that we had Tim Curry in our lives. RIP, my friend.”

GitHub - tailscale/tailcat: like netcat, but over Tailscale's data plane, without Tailscale's control plane

github.com

Tailscale with­out Tailscale, by Tailscale”

Tailcat is a remix of Tailscale open source pieces to act like net­cat, but over Tailscale’s data plane, with­out Tailscale’s con­trol plane. Tailscale’s data plane (magicsock, in­ter­nally) gives you point-to-point WireGuard®-encrypted tun­nels be­tween two ma­chines with DERP as the NAT-hole-punching com­mu­ni­ca­tion side chan­nel and the ul­ti­mate re­lay-of-last-re­sort if NAT tra­ver­sal fails. Instead of us­ing the Tailscale con­trol plane, all tail­cat con­nec­tion meta­data is ex­changed out of band, how­ever you want.

The tail­cat CLI (in cmd/​tail­cat) is built on the tail­cat Go li­brary (importable as github.com/​tailscale/​tail­cat).

Whether you use tail­cat as a CLI tool or li­brary, one side runs a tail­cat server (listener) and gets back a short con­nec­tion to­ken. The other side passes that to­ken to tail­cat’s client side to con­nect. All traf­fic be­tween the two is en­crypted end-to-end with WireGuard. The ini­tial con­nec­tion boot­straps through a DERP server (see be­low), and then mag­ic­sock per­forms NAT tra­ver­sal to up­grade to a di­rect peer-to-peer UDP con­nec­tion when pos­si­ble (usually!).

You don’t need a Tailscale ac­count, root/​ad­min ac­cess on the ma­chine (it does­n’t al­ter your ma­chine’s rout­ing ta­bles, DNS, etc.). It’s just a user­space li­brary and CLI tool.

And it’s all open source.

You can use our free rate-lim­ited DERP re­lays (the de­fault DERP map is https://​tail­cat.dev/​derpmap.json) or you can run your own.

There’s also an ex­per­i­men­tal in-browser web demo (tailcat com­piled to WebAssembly) at https://​tailscale.github.io/​tail­cat/ that can send and re­ceive files or text, in­ter­op­er­at­ing with the CLI. Browser traf­fic is re­layed over DERP only, with no di­rect con­nec­tions un­til WebRTC sup­port (#4).

Install

$ go in­stall github.com/​tailscale/​tail­cat/​cmd/​tail­cat@lat­est

Or with Nix flakes, run it di­rectly or in­stall it:

$ nix run github:tailscale/​tail­cat $ nix pro­file in­stall github:tailscale/​tail­cat

Usage

Pipe stdin/​std­out be­tween two ma­chines

Server starts, print­ing out its ephemeral ad­dress:

$ tail­cat # Selected boot­strap re­lay re­gion 302, San Francisco # 🐈 Server lis­ten­ing with new ad­dress: tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNGF­pGQEu (hangs, wait­ing…)

And then the client can:

$ echo hello | tail­cat tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNGF­pGQEu $

Then the server un­blocks:

$ tail­cat # Selected boot­strap re­lay re­gion 302, San Francisco # 🐈 Server lis­ten­ing with new ad­dress: tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNGF­pGQEu hello $

Expose lo­cal ports through the tun­nel

Or you can serve a lo­cal TCP port, for­warded to lo­cal­host:

$ tail­cat –serve=8080,8443 # or –serve=all # 🐈 Server lis­ten­ing with new ad­dress: tcXXXXXXXXX

And then the client:

$ tail­cat tcXXXXXXXXX 8080 GET / HTTP/1.1 Host: foo

HTTP/1.1 200 OK ….

Auth-free SSH server

On Linux and ma­cOS, you can run an SSH server too with no auth. (If you want auth, you can just tail­cat –serve=22 and proxy to your sys­tem SSH server)

$ tail­cat –serve=no-auth-ssh # 🐈 Server lis­ten­ing with new ad­dress: tcXXXXXXXXX

And on the client side:

$ tail­cat ssh tcXXXXXXXXX $ tail­cat ssh tcXXXXXXXXX ls -la

Misc com­mands

Ping to test con­nec­tiv­ity; each pong re­ports whether it ar­rived via a DERP re­lay or a di­rect path. –until-direct keeps ping­ing (up to –timeout, de­fault 10s) un­til a di­rect path works, ex­it­ing non-zero if one does­n’t:

$ tail­cat ping –until-direct <token> pong in 42.1ms via DERP(sfo) pong in 1.2ms via 203.0.113.7:41641

Run a com­mand through a SOCKS5 proxy routed over the tun­nel:

$ tail­cat socks <token> curl http://​server.tail­cat:8081/

Tokens also work di­rectly as URL host­names: the SOCKS proxy rec­og­nizes and di­als them, so the to­ken ar­gu­ment is op­tional. (Tokens are case-sen­si­tive; this works with curl and most CLI tools, but not with browsers, which low­er­case host­names.)

$ tail­cat socks curl http://&​lt;to­ken>:8081/

Act as an exit node so the client can reach the server’s net­work:

$ tail­cat –serve=exit-node

Parse a con­nec­tion to­ken and print its con­tents (the server’s WireGuard pub­lic key and DERP info) as JSON, with­out con­nect­ing to any­thing:

$ tail­cat parse tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNGF­pGQEu { ServerPublic”: nodekey:9c8d2e6728da80a1dd37e275a82595b42d9a838610bc53f74a7670d1610f2e34″, RegionID”: 302 }

Resolve a short to­ken (which ref­er­ences a DERP re­gion by ID, re­quir­ing clients to fetch the DERP map) into a longer self-con­tained one with the DERP server info em­bed­ded, let­ting clients con­nect more quickly:

$ tail­cat re­solve tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNGF­pGQEu tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNG­Fy­gaFh­ToGjY­WhudG­MzMD­JhLml­w­bi5kZXZh­NG0yMDguM­TExLjM5LjM4YTZzMjY­wNzpmNzQ­wO­jA6M2Y6O­j­cyMA

Parsing that re­solved to­ken shows the em­bed­ded DERP info:

$ tail­cat parse tcom­FwW­C­C­cjS5nKN­qAod034n­Wo­JZW0LZqD­hhC8U_d­Kd­nDRYQ8uNG­Fy­gaFh­ToGjY­WhudG­MzMD­JhLml­w­bi5kZXZh­NG0yMDguM­TExLjM5LjM4YTZzMjY­wNzpmNzQ­wO­jA6M2Y6O­j­cyMA { ServerPublic”: nodekey:9c8d2e6728da80a1dd37e275a82595b42d9a838610bc53f74a7670d1610f2e34″, Region”: [ { Nodes”: [ { HostName”: tc302a.ipn.dev”, IPv4”: 208.111.39.38″, IPv6”: 2607:f740:0:3f::720″ } ] } ] }

A server can print the long self-con­tained form di­rectly with the –full-address flag.

Key Management

A server’s ad­dress (connection to­ken) is de­rived from its WireGuard key, so the key you use de­ter­mines who can reach you:

Ephemeral keys (the de­fault): each server run gen­er­ates a fresh key in mem­ory and prints an ad­dress no­body has ever seen. When the process ex­its, the key is dis­carded and the ad­dress is dead for­ever. This is the safe de­fault: shar­ing that ad­dress only ever refers to that one run.

Ephemeral keys (the de­fault): each server run gen­er­ates a fresh key in mem­ory and prints an ad­dress no­body has ever seen. When the process ex­its, the key is dis­carded and the ad­dress is dead for­ever. This is the safe de­fault: shar­ing that ad­dress only ever refers to that one run.

Saved keys: tail­cat genkey gen­er­ates a key saved to disk so the ad­dress stays sta­ble across restarts. The flip side: any­one you’ve ever shared that ad­dress with can con­nect to any fu­ture server us­ing that key, un­less you re­strict clients with –allow (see tail­cat genkey –client).

Saved keys: tail­cat genkey gen­er­ates a key saved to disk so the ad­dress stays sta­ble across restarts. The flip side: any­one you’ve ever shared that ad­dress with can con­nect to any fu­ture server us­ing that key, un­less you re­strict clients with –allow (see tail­cat genkey –client).

The CLI says at startup which kind it’s us­ing, so you know whether you’re start­ing a fresh sin­gle-use server or re-lis­ten­ing on an ad­dress you may have shared in the past.

$ tail­cat genkey –region=nyc # prints the to­ken; key saved to ~/.config/tailcat/keys/default.private.json

# later; the key named default” is used au­to­mat­i­cally once it ex­ists: $ tail­cat –serve=8080 # 🐈 Server lis­ten­ing with saved key default”: tcXXXXXXXXX

# … un­less you force a one-off ephemeral key: $ tail­cat –serve=8080 –key=new # 🐈 Server lis­ten­ing with new ad­dress: tcXXXXXXXXX

That is, de­fault is a magic key name: once it ex­ists, plain tail­cat silently uses it in­stead of gen­er­at­ing an ephemeral key, and the startup line above is what tells you which hap­pened. Use –key=new to get an ephemeral key any­way, –key=<name> to use a dif­fer­ent saved key, or tail­cat genkey –delete –key=default to re­move the saved de­fault key. tail­cat genkey –list lists your saved keys.

Tokens can also be pub­lished as DNS TXT records and looked up by name; a DNS name works any­where the CLI takes a to­ken:

# If ex­am­ple.com has a TXT record tailcat=tc…” $ tail­cat ex­am­ple.com 8080 $ tail­cat ssh ex­am­ple.com $ tail­cat ping ex­am­ple.com

Examples

Protected SSH server over DNS

Who needs port for­ward­ing or port knock­ing? This runs an SSH server reach­able from any­where by name, with no open in­bound ports on the server, where WireGuard au­then­ti­cates the client be­fore the SSH server ever sees a packet.

On the client ma­chine, gen­er­ate a client iden­tity key­pair. It prints the pub­lic key, which is all the server needs to know:

client$ tail­cat genkey –client # wrote file to ~/.config/tailcat/keys/client-default.private.json nodekey:cf­b6b­fa77a0654d7450947fd6ace­f17d2cd848­da1d30b2540b13­dac272ddfd16

On the server, gen­er­ate a server key­pair pinned to its near­est DERP re­gion (see why be­low), then serve SSH to only that client:

server$ tail­cat genkey –fixed-region # wrote file to ~/.config/tailcat/keys/default.private.json tcXXXXXXXXX

server$ tail­cat –serve=22 –allow=nodekey:cfb6bf…ddfd16 # 🐈 Server lis­ten­ing with saved key default”: tcXXXXXXXXX

Publish the to­ken in DNS as a TXT record:

my-server.ex­am­ple.com. 300 IN TXT tailcat=tcXXXXXXXXX”

And then the client side is just:

client$ tail­cat ssh my-server.ex­am­ple.com

Client modes au­to­mat­i­cally use the saved client-de­fault key when it ex­ists, so no ex­tra flags are needed to pre­sent the al­lowed iden­tity. Anyone else’s hand­shake is silently ig­nored: they can’t reach the SSH server, or even learn that one is run­ning.

Why –fixed-region: it dis­cov­ers the near­est DERP re­gion once, at genkey time, and bakes its ID into both the printed to­ken and the saved key file, so server restarts bind to the same re­gion (keeping the pub­lished to­ken valid) with­out re-prob­ing. Plain tail­cat genkey de­faults to –region=auto, which in­stead bakes in pick at startup”: fine for one-off use, but a to­ken pub­lished in DNS should name a fixed re­gion so clients and fu­ture server restarts all ren­dezvous in the same place. (–region=<name> pins an ex­plicit one in­stead; –region=list shows the choices.)

TODO: make the client more ro­bust here if the DERP map changes over time: #7

Bring your own DERP re­lay

Nothing re­quires Tailscale’s re­lays: run your own DERP server (it needs a host­name with a TLS cer­tifi­cate, which der­per can get it­self via Let’s Encrypt), then gen­er­ate a server key that uses it by pass­ing its host­name (or sev­eral, comma-sep­a­rated) as the re­gion:

server$ tail­cat genkey –region=derp.example.com tcom­FwW­C­CAIsKO­qPU­ux6­ClG2R­M4A_vO­q4VBzG­gHG­Gjq9OsJuFK­SWFy­gaFh­ToGhY­Wh­wZGVy­c­C5leGFtcGxlLm­N­vbQ

server$ tail­cat –serve=22

The to­ken em­beds your re­lay’s host­name:

$ tail­cat parse tcom­FwW­C­CAIsKO­qPU­ux6­ClG2R­M4A_vO­q4VBzG­gHG­Gjq9OsJuFK­SWFy­gaFh­ToGhY­Wh­wZGVy­c­C5leGFtcGxlLm­N­vbQ { ServerPublic”: nodekey:8022c28ea8f52ec7a0a51b644ce00fef3aae150731a01c61a3abd3ac26e14a49″, Region”: [ { Nodes”: [ { HostName”: derp.example.com” } ] } ] }

so clients need no ex­tra flags and never con­tact Tailscale’s DERP map server or re­lays, and the only rate lim­its are yours. Alternatively, if you run a whole fleet of re­lays, serve your own DERP map JSON and point both sides at it with –derpmap-url.

Go li­brary

A min­i­mal server that an­swers any TCP port through the tun­nel and prints its to­ken. The zero value Server picks de­faults for any­thing un­set: a fresh ephemeral key, the near­est re­gion of the de­fault DERP map, and log.Printf log­ging (set Logf to log­ger.Dis­card for quiet):

pack­age main

im­port ( fmt” log” net”

github.com/​tailscale/​tail­cat )

func main() { s := &tailcat.Server{ OnTCP: func(port uin­t16) func(net.Conn) { re­turn func(c net.Conn) { fmt.Fprintf(c, hello from port %v\n”, port) c.Close() } }, } if err := s.Start(); err != nil { log.Fa­tal(err) } fmt.Println(s.ConnBlob()) se­lect {} }

And a min­i­mal client that di­als it, given that to­ken as its ar­gu­ment. Like Server, the Client zero value works with just its Server to­ken field set (tailcat.NewClient is short­hand for ex­actly that), and the tun­nel is es­tab­lished lazily by the first dial:

pack­age main

im­port ( context” io” log” os”

github.com/​tailscale/​tail­cat )

func main() { cl := tail­cat.New­Client(tail­cat.ConnBlob(os.Args[1])) de­fer cl.Close() c, err := cl.Di­alTCP­Port(con­text.Back­ground(), 80) if err != nil { log.Fa­tal(err) } io.Copy(os.Std­out, c) }

$ ./client tcom­FwW­CAWf933BLELdzd3RkHiOufJ… hello from port 80

How it works

Connection to­kens

A Tailcat server is iden­ti­fied by a con­nec­tion to­ken (called a ConnBlob in­ter­nally). It looks like tcXYZ… and is a tc” pre­fix fol­lowed by base64-en­coded CBOR con­tain­ing:

The server’s WireGuard pub­lic key (Curve25519, 32 bytes)

DERP info. Either:

a small in­te­ger ref­er­enc­ing one of the de­fault Tailscale-run tail­cat servers), or full DERP server meta­data, to ei­ther use a cus­tom DERP server, or to avoid the client need­ing a po­ten­tial round-trip to fetch the lat­est DERP map (the server’s –full-address flag and the tail­cat re­solve sub­com­mand pro­duce this form)

Nvidia has been in talks to acquire Hugging Face for more than $13 billion

www.businessinsider.com

Nvidia has been in talks to ac­quire Hugging Face, the pop­u­lar AI plat­form for shar­ing and build­ing with open-source mod­els, in what could be one of the chip gi­ant’s biggest deals yet.

The two par­ties have had ac­qui­si­tion con­ver­sa­tions in re­cent weeks about a deal that would value Hugging Face at more than $13 bil­lion, ac­cord­ing to a per­son fa­mil­iar with the mat­ter. The com­pa­nies have not yet reached a deal, and the talks could still fall apart, the per­son said. Business Insider on Sunday was the first to re­port that Hugging Face was field­ing takeover in­ter­est.

Nvidia and Hugging Face did not re­spond to re­quests for com­ment.

Nvidia has in­creas­ingly ramped up deal­mak­ing with its enor­mous cash pile. The com­pany said Wednesday that it has $18 bil­lion com­mit­ted to eq­uity in­vest­ments for the rest of its fis­cal year, on top of $47.9 bil­lion it al­ready holds in pri­vate com­pa­nies.

Microsoft also met with Hugging Face, but the per­son fa­mil­iar and a sec­ond per­son said talks are not on­go­ing.

Nvidia al­ready has a re­la­tion­ship with Hugging Face. The chip­maker par­tic­i­pated in its $235 mil­lion fund­ing round in 2023 that val­ued it at $4.5 bil­lion.

Hugging Face turned down a $500 mil­lion in­vest­ment of­fer from Nvidia late last year that would have val­ued it at $7 bil­lion, the Financial Times pre­vi­ously re­ported. Hugging Face said at the time it did not want a dom­i­nant in­vestor that could sway de­ci­sions.

Hugging Face sits at the cen­ter of the open-source AI ecosys­tem, host­ing mil­lions of AI mod­els and datasets that de­vel­op­ers can build on. Owning the plat­form could give Nvidia a big­ger foothold with those de­vel­op­ers — and po­ten­tially drive more work­loads onto its chips.

Hugging Face was founded in 2016 by French en­tre­pre­neurs Clément Delangue, Julien Chaumond, and Thomas Wolf.

But Nvidia own­er­ship could also com­pli­cate one of Hugging Face’s strengths: its neu­tral­ity. The plat­form sup­ports mod­els and hard­ware from across the in­dus­try, in­clud­ing Nvidia com­peti­tors such as AMD and Intel.

Have a tip?

Contact Katie Roof via email at kroof@busi­nessin­sider.com or Signal at @kroof.26

Contact Geoff Weiss via email at gweiss@busi­nessin­sider.com or Signal at @geoffweiss.25.

Contact Ashley Stewart via email at astew­art@busi­nessin­sider.com or Signal at +1 – 425-344 – 8242.

Use a per­sonal email ad­dress and a non­work de­vice; here’s our guide to shar­ing in­for­ma­tion se­curely.

Read next

Geoff Weiss

You’re cur­rently fol­low­ing this au­thor! Want to un­fol­low? Unsubscribe via the link in your email.

Geoff Weiss is a se­nior re­porter on Business Insider’s tech team, where he writes about AI star­tups and Y Combinator, the in­ter­sec­tion of AI and the me­dia in­dus­try, and work­place dy­nam­ics within top AI labs and chip com­pa­nies.Pre­vi­ously, Geoff was on the me­dia desk, cov­er­ing YouTube and Netflix, and themes like the in­ter­sec­tion of Hollywood and the cre­ator econ­omy. His work on Netflix’s video pod­cast­ing am­bi­tions and Mr Beast’s lessons for Hollywood won sec­ond and first prize, re­spec­tively, at the 2025 LA Press Club Awards.Prior to join­ing Business Insider, Geoff was the se­nior ed­i­tor of Tubefilter and a staff writer at Entrepreneur. He grad­u­ated from New York University with a de­gree in English Literature.He can be reached at gweiss@busi­nessin­sider.com, on Signal @geoffweiss.25, and on LinkedIn. Have a tip? Use a per­sonal email ad­dress and a non­work de­vice; here’s our guide to shar­ing in­for­ma­tion se­curely.Se­lected sto­ries:Nvidia crushed its quar­ter — and CEO Jensen Huang said in a leaked all-hands that the mar­ket did not ap­pre­ci­ate it’N­vidia will foot the bill for Trump’s new visa fees. Here’s what CEO Jensen Huang told staff.Mas­sive AI salaries and RTO are fu­el­ing a real es­tate boom in San Francisco: It’s go­ing to rain mon­ey’The AI tal­ent wars are ric­o­chet­ing across star­tups. Here’s how they’re com­pet­ing with Big Tech.

Ashley Stewart

You’re cur­rently fol­low­ing this au­thor! Want to un­fol­low? Unsubscribe via the link in your email.

AI

Big Tech

Venture Capital

More

Exclusive

reuters.com

www.reuters.com

Please en­able JS and dis­able any ad blocker

6 RAG Architectures — and How to Avoid Over-Engineering

www.lighthousenewsletter.com

Hello, Rafael here - every week I cover in­ter­est­ing chal­lenges and de­vel­op­ments that I’ve come across re­cently through the lens of an en­gi­neer build­ing AI sys­tems.

Subscribe and get my weekly takes 👇

Nowadays, most peo­ple seem to over-en­gi­neer their RAG stack. They jump straight to em­bed­dings, vec­tor data­bases, and rerank­ing pipelines. Meanwhile, their users just want to find the doc that says How to re­set my pass­word.”

In en­gi­neer­ing, there’s al­ways the right tool for the right prob­lem. In AI Retrieval Systems it’s not dif­fer­ent.

Before we dive into recipes, let’s es­tab­lish when you should use each ap­proach. The key fac­tors are:

1. Data Freshness Requirements - Real-time up­dates (news, so­cial me­dia) fa­vor ap­proaches with easy re-in­dex­ing. Daily or weekly up­dates work well with hy­brid ap­proaches. A sta­ble cor­pus (monthly or quar­terly up­dates) makes pre-em­bed­ding sen­si­ble.

2. Corpus Characteristics - High churn (more than 10% changes daily) means you should avoid full pre-em­bed­ding. Stable doc­u­ments work fine with pre-em­bed­ding. Long-tail dis­tri­b­u­tion (90% never ac­cessed) means on-the-fly wins.

3. Query Patterns - Keyword-heavy queries should start with full-text search. Semantic or con­ver­sa­tional queries ben­e­fit from em­bed­dings. Mixed pat­terns need hy­brid ap­proaches.

4. Scale & Performance - Less than 1000 queries per day means sim­ple ap­proaches are suf­fi­cient. 1K to 10K queries per day re­quires se­lec­tive op­ti­miza­tion. More than 10K queries per day jus­ti­fies full op­ti­miza­tion.

5. Team Capabilities - No ML ex­per­tise means stay with full-text plus query rewrit­ing. Some ML ex­pe­ri­ence makes hy­brid search man­age­able. Having an ML team avail­able makes ad­vanced ap­proaches vi­able.

Now, let’s look at the recipe book. Start at the top. Move down only when you have data prov­ing you need to.

Good old BM25. Elasticsearch. Postgres full-text search. The stuff that ex­isted be­fore embedding” be­came a verb.

You’re just start­ing out. Your users write key­word-style queries (”pandas merge dataframe”). Exact matches mat­ter (”invoice #12345”). You want zero ML com­plex­ity. Your cor­pus has pro­pri­etary ter­mi­nol­ogy (more on this later).

Zero API costs. Fast (under 10ms). Easy to de­bug (you can see ex­actly why a doc­u­ment matched). Surprisingly ef­fec­tive (handles many use cases). No chunk­ing strat­egy needed — works with full doc­u­ments. No eval­u­a­tion com­plex­ity — easy to test and val­i­date. No model dep­re­ca­tion risk (BM25 does­n’t change).

Misses syn­onyms (”car” vs automobile”). Fails on se­man­tic queries (”How do I…?”). Can’t un­der­stand in­tent be­yond key­words.

In my ex­pe­ri­ence, this han­dles a sig­nif­i­cant por­tion of use cases. Don’t skip this step. You might be sur­prised how far you can get.

When you jump straight to em­bed­dings, you im­me­di­ately face ques­tions like: What chunk size? (512 to­kens? 1024?) What over­lap? (50 to­kens? 100?) Semantic chunk­ing or fixed-size? How do I eval­u­ate if my chunk­ing is good?

With full-text search, you skip all of this. Your doc­u­ments are your doc­u­ments. Search just works.

Thanks for read­ing Lighthouse AI! This post is pub­lic so feel free to share it.

Share

Use an LLM to trans­form messy user queries into clean key­word searches.

Most semantic search” prob­lems are ac­tu­ally query for­mu­la­tion prob­lems.

Users ask ques­tions con­ver­sa­tion­ally. Vocabulary mis­match (users say fix bugs”, docs say debugging”). You have in­ter­nal jar­gon (your frame­work called Atlas”). You want flex­i­bil­ity to it­er­ate quickly on query strate­gies.

~$0.001 per query (using GPT-4o-mini for query rewrit­ing)

An LLM can re­move stop­words (”how do I” be­comes noth­ing). It can add syn­onyms (”car” be­comes car au­to­mo­bile ve­hi­cle”). It can trans­late do­main terms (”speed up code” be­comes optimize per­for­mance”). It can de­com­pose com­plex queries (”read CSV and plot” be­comes [”read CSV, plot data”]). It can learn from your glos­sary (via sys­tem prompt).

With em­bed­dings, if re­sults aren’t good, you need to ad­just chunk­ing strat­egy, re-em­bed en­tire cor­pus, run re­gres­sion tests on your eval set, and hope it im­proved.

With query rewrit­ing, if re­sults aren’t good, you ad­just the sys­tem prompt. That’s it. Test im­me­di­ately.

Even bet­ter, you can cre­ate a loop:

The agent can it­er­ate, learn, and adapt — all with­out re-em­bed­ding any­thing.

Say your com­pany has a Python frame­work called Atlas.” If you use gen­eral-pur­pose em­bed­dings:

General em­bed­ding model (trained on in­ter­net): Atlas” = [vectors point­ing to­ward: Greek mythol­ogy, maps, ge­og­ra­phy] Your ac­tual Atlas docs = [vectors about data pro­cess­ing] Similarity score: 0.15 (terrible!)

The model has no idea your Atlas” ex­ists. It falls back to what it learned in train­ing. But with query rewrit­ing:

For pro­pri­etary terms, ex­act key­word match­ing beats se­man­tic un­der­stand­ing.

Use BM25 to get can­di­dates (top 50 – 100), then rerank with em­bed­dings (top 10).

BM25 is fast and great at key­word match­ing. Embeddings are good at se­man­tic un­der­stand­ing. Together, they cover each oth­er’s weak­nesses.

Users ask se­man­tic ques­tions (”find al­ter­na­tives to X”). BM25 plus query rewrit­ing alone is­n’t cut­ting it (you have data prov­ing this). You can tol­er­ate 100 – 500ms la­tency. Your cor­pus is rel­a­tively sta­ble (not chang­ing every minute).

Let’s do the math with cur­rent pric­ing (OpenAI text-em­bed­ding-3-small at $0.02 per 1M to­kens):

Embedding 50 docs per query (avg 500 to­kens each) means 50 docs × 500 to­kens = 25,000 to­kens

Embedding 50 docs per query (avg 500 to­kens each) means 50 docs × 500 to­kens = 25,000 to­kens

Cost: 25,000 × $0.00002 = ~$0.0005 per query. At 1,000 queries per day × 30 days = ~$15 per month.

Cost: 25,000 × $0.00002 = ~$0.0005 per query. At 1,000 queries per day × 30 days = ~$15 per month.

Actually pretty rea­son­able. But there’s a catch: la­tency.

Embedding 50 doc­u­ments on-the-fly adds 200 – 500ms per query. For user-fac­ing search, that’s no­tice­able. This is where the real trade-off lives — not cost, but speed.

When you in­tro­duce em­bed­dings, you need to de­cide how to chunk your doc­u­ments (fixed-size? se­man­tic? by sec­tion?). You need to de­ter­mine what chunk size and over­lap to use. You need to han­dle chunks that span im­por­tant con­text.

This adds com­plex­ity that pure full-text search avoids.

If your data changes fre­quently, why pay to re-em­bed every­thing?

High doc­u­ment churn (more than 10% of docs up­dated daily). Real-time con­tent (news, so­cial me­dia, live up­dates). You’re ex­per­i­ment­ing with em­bed­ding mod­els (no re-in­dex­ing needed). Data fresh­ness is crit­i­cal (documents must be up-to-date). Small K for rerank­ing (20 – 50 docs).

On-the-fly / on­line (1000 queries/​day, 50 docs/​query): - Embedding cost: ~$15/month (ongoing) - Storage: $0 (just store text) - Latency: 200 – 500ms per query - Freshness: Perfect (always cur­rent) - Model switch­ing: Easy (just change the API call)

Embedding mod­els get dep­re­cated.

OpenAI dep­re­cated text-em­bed­ding-ada-002 in fa­vor of text-em­bed­ding-3. If you pre-em­bed­ded 10 mil­lion doc­u­ments with the old model, you now need to re-em­bed all 10 mil­lion doc­u­ments with the new model, up­date your vec­tor data­base, run re­gres­sion tests on your eval­u­a­tion set, val­i­date that qual­ity did­n’t de­grade, han­dle the cu­tover pe­riod, and deal with any API changes.

You lit­er­ally just change one line of code. Done.

Latency. You’re em­bed­ding doc­u­ments on every query. This is only vi­able if you’re okay with 200 – 500ms la­tency, K is small (reranking 20 – 50 docs, not 500), and your use case fa­vors fresh­ness over speed.

Pre-embed fre­quently ac­cessed doc­u­ments (”hot tier”), em­bed rarely-ac­cessed doc­u­ments on-the-fly (”cold tier”).

Access pat­terns fol­low Pareto dis­tri­b­u­tion. 20% of docs get 80% of traf­fic.

Clear ac­cess pat­terns (some docs are ac­cessed way more than oth­ers). Medium-to-large cor­pus (more than 100K doc­u­ments). Mix of sta­ble and chang­ing con­tent. Need good la­tency for com­mon queries. Want to min­i­mize re-em­bed­ding on model up­dates.

Fast for 80% of queries (hit pre-em­bed­ded cache). Fresh for rarely-ac­cessed docs. Only re-em­bed hot tier when switch­ing mod­els (20% of cor­pus). Adapts to chang­ing ac­cess pat­terns. Best la­tency/​cost/​flex­i­bil­ity trade-off.

When your em­bed­ding model gets dep­re­cated:

Full pre-em­bed­ding: Re-embed 1M docs × $0.01 = $10,000 + down­time Hot/cold tiers: Re-embed 200K docs × $0.01 = $2,000 + min­i­mal down­time On-the-fly: Change one line of code = $0 + zero down­time

Embed every­thing up­front. Store in vec­tor data­base. Search with ANN (approximate near­est neigh­bors).

Very high query vol­ume (more than 10K queries per day). Need un­der 50ms la­tency. Very sta­ble cor­pus (under 5% churn per month). Access pat­tern is broad (no long tail). You have ML team to man­age in­fra­struc­ture.

Pre-embedding (1M docs): - One-time em­bed­ding: 1M docs × 500 to­kens × $0.00002 = $10 - Storage: 1M × 1536 dims × 4 bytes = 6GB (~$10 – 30/month) - Search la­tency: un­der 50ms (blazing fast!) - Freshness: Only as fresh as last re-in­dex

Documents change fre­quently (more than 10% per week). You’re ex­per­i­ment­ing with em­bed­ding mod­els. Low query vol­ume (under 1K queries per day). You haven’t tried sim­pler ap­proaches first.

This is where full pre-em­bed­ding hurts the most. When you need to switch mod­els, you face down­time (your search is de­graded while re-em­bed­ding), com­pute cost (re-embedding mil­lions of doc­u­ments), test­ing bur­den (full re­gres­sion test suite on new em­bed­dings), chunk­ing reeval­u­a­tion (maybe new model works bet­ter with dif­fer­ent chunk sizes?), and risk (what if the new model is worse for your do­main?).

This is overkill for most sys­tems. I’ve seen teams spend months op­ti­miz­ing their vec­tor data­base setup when query rewrit­ing would have solved 90% of their prob­lems.

This is overkill for most sys­tems. I’ve seen teams spend months op­ti­miz­ing their vec­tor data­base setup when query rewrit­ing would have solved 90% of their prob­lems.

But if you’re Pinterest, Shopify, or han­dling mas­sive scale with a sta­ble cor­pus, this is where you end up.

We’ve been dis­cussing sin­gle-in­tent queries: How do I merge dataframes?”

But real users ask stuff like: How do I read a CSV file, clean miss­ing data, and plot the re­sults?”

That’s three sep­a­rate in­tents. Searching for this as one query is like try­ing to find a restau­rant that serves pizza, sushi, and tacos. Good luck.

Modern agen­tic RAG sys­tems (Perplexity, ChatGPT search) han­dle this el­e­gantly:

Break down the query.

Route each sub-query op­ti­mally

Combine re­sults into co­her­ent an­swer

Each sub-query is fo­cused and pre­cise, lead­ing to bet­ter re­trieval. Parallel ex­e­cu­tion means lower la­tency (max, not sum). Adaptive rout­ing re­sults in lower cost (only com­plex queries pay for LLM). Structured out­put pro­vides bet­ter UX.

Without de­com­po­si­tion

LLM rewrit­ing en­tire com­plex query: $0.005

LLM rewrit­ing en­tire com­plex query: $0.005

Embedding 50 docs: $0.025

Embedding 50 docs: $0.025

Total: $0.03

Total: $0.03

With de­com­po­si­tion

Decompose: $0.001

Decompose: $0.001

Sub-query 1 (simple): $0

Sub-query 1 (simple): $0

Sub-query 2 (simple): $0

Sub-query 2 (simple): $0

Sub-query 3 (complex): $0.001

Sub-query 3 (complex): $0.001

Total: $0.002

Total: $0.002

15x cheaper, bet­ter qual­ity.

This is where agen­tic re­trieval re­ally shines. The agent can in­tel­li­gently de­cide which sub-queries need ex­pen­sive pro­cess­ing (embeddings) and which can be han­dled with cheap meth­ods (simple pre­pro­cess­ing + BM25).

Okay, you’ve read this far. You just want to know: What should I build?”

Start here: Do you have search at all? If not, build BM25 first. Seriously. Stop read­ing and build it. If you do have search, con­tinue.

Measure your base­line. Run your cur­rent search for 2 – 4 weeks and col­lect user feed­back. Are users happy with the re­sults? If yes, stop. You’re done. Go ship fea­tures. If no, con­tinue.

What’s the main com­plaint?

If users say Can’t find docs that clearly ex­ist,” try query rewrit­ing first. At $0.001 per query with zero re-in­dex­ing, it’s worth test­ing. Run an A/B test for 2 weeks. If you see good im­prove­ment, keep it and you’re done. If it’s not enough, con­tinue.

If users say Results are okay but not great,” A/B test hy­brid search (sparse plus em­bed­ding rerank). Is the added la­tency worth it? If yes, de­cide on im­ple­men­ta­tion. If your data changes fre­quently, use on-the-fly em­bed­ding. If you have clear hot docs, use hot/​cold tiers. If you have a sta­ble cor­pus and high scale, use full pre-em­bed­ding. If the la­tency is­n’t worth it, op­ti­mize query rewrit­ing fur­ther in­stead.

Bloomberg - Are you a robot?

www.bloomberg.com

We’ve de­tected un­usual ac­tiv­ity from your com­puter net­work

To con­tinue, please click the box be­low to let us know you’re not a ro­bot.

Why did this hap­pen?

Please make sure your browser sup­ports JavaScript and cook­ies and that you are not block­ing them from load­ing. For more in­for­ma­tion you can re­view our Terms of Service and Cookie Policy.

Need Help?

For in­quiries re­lated to this mes­sage please con­tact our sup­port team and pro­vide the ref­er­ence ID be­low.

Block ref­er­ence ID:821dcd31-a1cb-11f1-a0d9 – 2f0c4f1083f6

Get the most im­por­tant global mar­kets news at your fin­ger­tips with a Bloomberg.com sub­scrip­tion.

Nebula Sans

www.nebulasans.com

Introducing

NebulaSans

A ver­sa­tile, mod­ern, hu­man­ist sans-serif with a neu­tral aes­thetic, de­signed for leg­i­bil­ity in both dig­i­tal and print ap­pli­ca­tions.

Based on Source Sans by Paul D. Hunt for Adobe Fonts.

Nebula Sans is the new brand type­face for Nebula, the pre­mium stream­ing ser­vice from in­de­pen­dent cre­ators. Based on Source Sans and de­signed to be a drop-in al­ter­na­tive to Whitney SSm, Nebula Sans is avail­able for any­one to use un­der the SIL Open Font License.

Watch our short doc­u­men­tary film about the story be­hind Nebula Sans, writ­ten & di­rected by David Friedman.

Featuring two styles in six weights, Nebula Sans is well-suited for use in in­ter­faces, print, and for any other graph­i­cal, dig­i­tal, phys­i­cal, meta­phys­i­cal, metaphor­i­cal, or al­le­gor­i­cal type­face needs.

Download View font li­cense

Nebula Sans Light

I’d take the awe of un­der­stand­ing over the awe of ig­no­rance any day

Nebula Sans Book

The tv” in neb­ula.tv stands for Taylor’s Version”

Nebula Sans Medium

Don’t use seven words when four will do

Nebula Sans Semibold

Introducing: Facts and fic­tion

Nebula Sans Bold

An in­die stream­ing ser­vice

Nebula Sans Black

Powered by hu­mans

Nebula Sans Black Italic

Enter the Snack Zone

Nebula Sans Bold Italic

There’s no place like home

Nebula Sans Semibold Italic

Charl is the key to our suc­cess

Nebula Sans Medium Italic

We’re as­sem­bling a crew for a heist

Nebula Sans Book Italic

We be­lieve in facts, sci­ence, and hu­man rights

Nebula Sans Light Italic

I’ve been nav­i­gat­ing based on car­di­nal di­rec­tions…and vibes

Why we made this

We built our own type­face for a few key rea­sons:

Personalization: we can cus­tomize the fonts to align with our pref­er­ences.

Features: we can in­te­grate ad­vanced ty­pog­ra­phy fea­tures tai­lored to our use cases.

Sustainability: the cost of li­cens­ing com­mer­cial type­faces in­creases as we grow.

Source Sans was the per­fect foun­da­tion for Nebula Sans be­cause it shares many pri­mary char­ac­ter­is­tics with Whitney SSm, our pre­vi­ous brand type­face — both were de­signed to bridge the gap be­tween American gothic and European hu­man­ist type­faces, with a strong em­pha­sis on read­abil­ity. The ma­jor­ity of the ad­just­ments we made were to adapt the met­rics of Source Sans to bet­ter match those of Whitney SSm, since Source Sans is smaller and nar­rower by de­fault.

Typographical Details

Punctuation

The de­fault punc­tu­a­tion marks in Whitney SSm were, to our taste, too straight. Nebula Sans uses beau­ti­ful curly glyphs from Source Sans.

Stylistic Alternates

Nebula Sans fea­tures the same styl­is­tic al­ter­nates as Source Sans, with the de­faults aligned with those of Whitney SSm.

font-fea­ture-set­tings: ss01’;

font-fea­ture-set­tings: ss02’;

font-fea­ture-set­tings: ss03’;

Asterisk

In ty­pog­ra­phy, the as­ter­isk sym­bol was named as such be­cause it re­sem­bles a star. We love stars, so how could we not put our own spin on this lit­tle glyph?

Tabular Figures

The de­fault ver­sion of Whitney SSm lacks sup­port for tab­u­lar lin­ing fig­ures, so we were thrilled to be able to in­clude them in Nebula Sans. Tabular fig­ures (or mono­spaced nu­mer­als) al­low us to do things like in­cre­ment the time­stamp in the video player while keep­ing the dig­its from jump­ing around as they change.

Wait now it’s work­ing

Why did it not work 5 min­utes ago

Internet weather

Sent 2:57pm

@&%!$#?!*

Sent 3:06pm

Lorem ip­sum do­lor sit amet, con­secte­tur adip­isc­ing elit.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.