10 interesting stories served every morning and every evening.

DuckLabs to Join AWS, Projects to Remain Open Source

ducklabs.com

2026 – 08-26 | 10 min

Today, we’re an­nounc­ing that DuckLabs will join Amazon Web Services (AWS), which is ex­pected to be ef­fec­tive in early September.

Our team will re­main to­gether in Amsterdam, con­tin­u­ing our work on DuckDB, DuckLake, Quack, and the broader com­mu­nity. Joining AWS gives us the re­sources and reach to bring this tech­nol­ogy to many more de­vel­op­ers and or­ga­ni­za­tions, and to pur­sue ideas at a scale that would have been dif­fi­cult for us to reach alone.

Most im­por­tantly, the foun­da­tions of the pro­ject will re­main firmly in place. DuckDB and the other open-source com­po­nents of the Duck Stack” will re­main free and open source un­der the MIT li­cense, with the non­profit DuckDB Foundation con­tin­u­ing its stew­ard­ship of the pro­jects.

This is a sig­nif­i­cant mo­ment for all of us at DuckLabs. It marks the end of one chap­ter that we are im­mensely proud of, and the be­gin­ning of an­other that we be­lieve can take DuckDB much fur­ther.

Our Journey to This Point

We founded DuckLabs a lit­tle over five years ago to give the team be­hind DuckDB a sta­ble, long-term home.

At the time, DuckDB was be­gin­ning to gain real mo­men­tum. The first com­mer­cial con­tracts to pri­or­i­tize fea­tures were ma­te­ri­al­iz­ing, and ven­ture cap­i­tal firms were call­ing. We chose a dif­fer­ent path: a boot­strapped com­pany, fully owned by its founders and de­vel­op­ment team.

That de­ci­sion shaped DuckLabs in ways we value deeply. It gave us the free­dom to build pa­tiently, to put the tech­nol­ogy first, and to grow with­out los­ing sight of why we started. From a small group gath­ered around an am­bi­tious open-source pro­ject, we grew into a team of more than 30 peo­ple in Amsterdam, all while con­tin­u­ing to in­vest in DuckDB and the com­mu­nity tak­ing shape around it.

During those years, DuckDB trav­eled much fur­ther than we imag­ined when the pro­ject be­gan. Today, we see more than one mil­lion down­loads every day. Developers around the world use it to ex­plore data, power prod­ucts, teach new ideas, con­duct re­search, and build sys­tems we could never have an­tic­i­pated our­selves.

Watching that hap­pen has been one of the great priv­i­leges of our pro­fes­sional lives. But the lim­its of our ex­ist­ing model also be­came in­creas­ingly clear. As founders, we wor­ried that DuckDB’s growth would even­tu­ally out­pace our abil­ity to sup­port it. That our small com­pany could be­come a bot­tle­neck for the pro­ject, the team, and the peo­ple build­ing busi­nesses on top of it.

We also wor­ried that scal­ing DuckLabs into a much larger sales, sup­port, and op­er­a­tions or­ga­ni­za­tion would pull our at­ten­tion away from the tech­ni­cal work and open-source com­mu­nity that made DuckDB suc­cess­ful in the first place.

Our part­ner­ships work best with highly tech­ni­cal or­ga­ni­za­tions — of­ten com­pa­nies with sub­stan­tial data­base ex­per­tise of their own. Reaching a much broader group of users re­quires us to solve more com­plete and spe­cial­ized prob­lems, serve the needs of dif­fer­ent in­dus­tries, in­vest sub­stan­tially more in in­fra­struc­ture, and reach peo­ple who may never think to seek out an an­a­lyt­i­cal data­base di­rectly.

We be­lieve the DuckDB rev­o­lu­tion can grow by an­other cou­ple of or­ders of mag­ni­tude. To give it that op­por­tu­nity, we re­al­ized we needed a dif­fer­ent setup.

Why AWS

DuckLabs and AWS have al­ready been work­ing closely to­gether for more than a year. During that time, we came to un­der­stand how our teams col­lab­o­rate, what each brings to the table, and what might be­come pos­si­ble by com­bin­ing DuckLabs’ tech­ni­cal ex­per­tise with AWSs in­fra­struc­ture, scale, and cus­tomer reach.

That ex­pe­ri­ence gave us con­fi­dence in tak­ing this next step.

Together, we plan to use DuckDB, DuckLake, and Quack to help power a new gen­er­a­tion of data ser­vices. Our am­bi­tion is to reach peo­ple who may even­tu­ally de­pend on the Duck Stack every day, whether they in­ter­act with it di­rectly or en­counter it qui­etly in­side the prod­ucts and ser­vices they use.

For the DuckLabs team, join­ing AWS cre­ates the space to con­cen­trate on the tech­ni­cal work we care about most while op­er­at­ing at a scale we could not eas­ily achieve on our own. It al­lows us to think fur­ther ahead, be bolder, tackle harder prob­lems, and bring the ideas be­hind DuckDB to many more peo­ple.

AWS has com­mit­ted to sup­port­ing the con­tin­ued de­vel­op­ment of DuckDB and its wider com­mu­nity for the long term. That com­mit­ment mat­ters deeply to us.

Quote DuckDB is an in­cred­i­ble open source pro­ject with an amaz­ing com­mu­nity; it is broadly used and very much loved by S3 cus­tomers to­day. After about two years of work­ing closely with Mark, Hannes and the whole team at DuckLabs I’m ex­cited at the op­por­tu­nity to help the pro­ject have an even broader im­pact. Also, and maybe a lit­tle self­ishly, I’ve found the DuckLabs team to be one of the most tech­ni­cally deep, hum­ble, and high-ve­loc­ity teams that I’ve ever had a chance to work with and I’m de­lighted that we get to do a lot more of that.”

– Andy Warfield (Distinguished Engineer and Vice President, AWS)

Quote DuckDB is an in­cred­i­ble open source pro­ject with an amaz­ing com­mu­nity; it is broadly used and very much loved by S3 cus­tomers to­day. After about two years of work­ing closely with Mark, Hannes and the whole team at DuckLabs I’m ex­cited at the op­por­tu­nity to help the pro­ject have an even broader im­pact. Also, and maybe a lit­tle self­ishly, I’ve found the DuckLabs team to be one of the most tech­ni­cally deep, hum­ble, and high-ve­loc­ity teams that I’ve ever had a chance to work with and I’m de­lighted that we get to do a lot more of that.”

– Andy Warfield (Distinguished Engineer and Vice President, AWS)

What Will Remain the Same

We know that an an­nounce­ment like this brings im­por­tant — and un­der­stand­able — ques­tions for users, con­trib­u­tors, cus­tomers, and part­ners.

DuckDB, DuckLake, Quack, and the other open-source com­po­nents of the Duck Stack will re­main free and open source un­der the MIT li­cense. The non­profit DuckDB Foundation will con­tinue to stew­ard these pro­jects, and the DuckLabs team will con­tinue to con­tribute to the pro­ject and re­main to­gether in Amsterdam.

The open-source pro­jects will con­tinue to serve a broad com­mu­nity of users, con­trib­u­tors, plat­forms, and ven­dors. The open­ness that al­lowed DuckDB to flour­ish will re­main cen­tral to its fu­ture.

We care deeply about the trust this com­mu­nity has placed in us. Protecting that trust has been one of the most im­por­tant con­sid­er­a­tions through­out this process.

Quote As the CWI rep­re­sen­ta­tive on the DuckDB Foundation, I would like to con­grat­u­late Hannes, Mark and all DuckLabs em­ploy­ees with this new chap­ter. When DuckLabs spun out of CWI, we cre­ated this foun­da­tion, which holds all IP of open-source DuckDB, and will con­tinue to do so. I am de­lighted that AWS is com­mit­ted to keep ad­vanc­ing open-source DuckDB, and I an­tic­i­pate it to even ac­cel­er­ate its in­no­va­tion. Via the DuckDB Foundation we will also make sure that the voices of its sup­port­ers and all mem­bers of the open-source DuckDB com­mu­nity at large, will con­tinue to be heard. I am very proud of DuckDB and its cre­ators; this de­vel­op­ment un­der­lines it is state-of-the-art tech­nol­ogy, which re­al­izes many ideas from the data­base ar­chi­tec­ture re­search group in open source, for every­body’s ben­e­fit.”

– Peter Boncz (Database Architectures group lead at CWI Amsterdam, DuckDB Foundation board mem­ber)

Quote As the CWI rep­re­sen­ta­tive on the DuckDB Foundation, I would like to con­grat­u­late Hannes, Mark and all DuckLabs em­ploy­ees with this new chap­ter. When DuckLabs spun out of CWI, we cre­ated this foun­da­tion, which holds all IP of open-source DuckDB, and will con­tinue to do so. I am de­lighted that AWS is com­mit­ted to keep ad­vanc­ing open-source DuckDB, and I an­tic­i­pate it to even ac­cel­er­ate its in­no­va­tion. Via the DuckDB Foundation we will also make sure that the voices of its sup­port­ers and all mem­bers of the open-source DuckDB com­mu­nity at large, will con­tinue to be heard. I am very proud of DuckDB and its cre­ators; this de­vel­op­ment un­der­lines it is state-of-the-art tech­nol­ogy, which re­al­izes many ideas from the data­base ar­chi­tec­ture re­search group in open source, for every­body’s ben­e­fit.”

– Peter Boncz (Database Architectures group lead at CWI Amsterdam, DuckDB Foundation board mem­ber)

Quote Over the course of five years, DuckDB’s open na­ture, its sheer hack­a­bil­ity, and the re­mark­ably friendly com­mu­nity of de­vel­op­ers has turned the sys­tem into one of the pre­mier plat­forms for data­base re­search as well as teach­ing. For both, it is es­sen­tial that we can in­spect and tin­ker with DuckDB’s ker­nel. I am thus ex­cited to learn that DuckDB will re­main open source un­der the um­brella of AWS. An even wider range of op­por­tu­ni­ties is in reach now. We are glad to be able to be a part of this new era for DuckLabs and DuckDB.”

– Torsten Grust (Professor of Computer Science and Database Systems re­search group lead at Universität Tübingen, Germany)

Quote Over the course of five years, DuckDB’s open na­ture, its sheer hack­a­bil­ity, and the re­mark­ably friendly com­mu­nity of de­vel­op­ers has turned the sys­tem into one of the pre­mier plat­forms for data­base re­search as well as teach­ing. For both, it is es­sen­tial that we can in­spect and tin­ker with DuckDB’s ker­nel. I am thus ex­cited to learn that DuckDB will re­main open source un­der the um­brella of AWS. An even wider range of op­por­tu­ni­ties is in reach now. We are glad to be able to be a part of this new era for DuckLabs and DuckDB.”

– Torsten Grust (Professor of Computer Science and Database Systems re­search group lead at Universität Tübingen, Germany)

What We Plan to Expand

Joining AWS will give us greater ca­pac­ity to in­vest in both the tech­nol­ogy and the com­mu­nity around it. Going for­ward, the DuckDB Foundation will in­clude a tech­ni­cal ad­vi­sory board, so that lead­ing com­mu­nity mem­bers can pro­vide their in­put on the pro­jec­t’s tech­ni­cal di­rec­tion. We also plan to open the ex­ten­sion stack so that ex­ten­sions signed by other de­vel­op­ers and or­ga­ni­za­tions can run in DuckDB.

There is a great deal of work ahead, and many de­tails still to shape. We will share more as these plans de­velop.

Quote Amazon putting its weight be­hind DuckDB is go­ing to add a ton of mo­men­tum and strengthen the ecosys­tem. This is great news for those of us who be­lieve in DuckDB as the plat­form on which the fu­ture of an­a­lyt­ics is be­ing built.”

– Jordan Tigani (CEO, MotherDuck)

Quote Amazon putting its weight be­hind DuckDB is go­ing to add a ton of mo­men­tum and strengthen the ecosys­tem. This is great news for those of us who be­lieve in DuckDB as the plat­form on which the fu­ture of an­a­lyt­ics is be­ing built.”

– Jordan Tigani (CEO, MotherDuck)

Quote Amazon is the ideal home for DuckLabs. DuckDB is the cen­ter of the ven­dor-neu­tral open data stack, and Amazon has the com­mit­ment to open­ness, and the track record of work­ing with the en­tire cloud ecosys­tem, to en­able the DuckDB pro­ject to con­tinue to thrive in this role.”

– George Fraser (CEO and co-founder, Fivetran)

Quote Amazon is the ideal home for DuckLabs. DuckDB is the cen­ter of the ven­dor-neu­tral open data stack, and Amazon has the com­mit­ment to open­ness, and the track record of work­ing with the en­tire cloud ecosys­tem, to en­able the DuckDB pro­ject to con­tinue to thrive in this role.”

– George Fraser (CEO and co-founder, Fivetran)

The Next Chapter

DuckDB’s suc­cess has al­ways be­longed to a much larger com­mu­nity than the peo­ple work­ing in­side DuckLabs.

It be­longs to every­one who has used it, con­tributed code, re­ported a bug, writ­ten an ex­ten­sion, an­swered a ques­tion, pub­lished a bench­mark, taught a class, built a prod­uct, chal­lenged our as­sump­tions, or rec­om­mended DuckDB to some­one else. Every one of those acts helped the pro­ject be­come what it is to­day.

We do not take that sup­port — or the trust be­hind it — for granted.

Joining AWS gives the DuckLabs team an ex­tra­or­di­nary op­por­tu­nity to bring the Duck Stack to a much larger au­di­ence while con­tin­u­ing to in­vest in the open-source tech­nol­ogy at its heart. We en­ter this next chap­ter with the same cu­rios­ity, care, and tech­ni­cal am­bi­tion that brought us here, along­side a team that has been through the en­tire jour­ney to­gether.

To every­one who helped us reach this point: thank you. We are proud of what we have built to­gether, ex­cited by what now lies ahead, and look­ing for­ward to build­ing the next chap­ter with you.

You’re also wel­come to check out our press re­lease and the me­dia kit.

You’re also wel­come to check out our press re­lease and the me­dia kit.

z.ai

Qwen Studio

qwen.ai

reuters.com

www.reuters.com

Please en­able JS and dis­able any ad blocker

Tim Curry, star of The Rocky Horror Picture Show and Stephen King’s It, dies aged 80

www.theguardian.com

Tim Curry, the ver­sa­tile per­former of screen and stage, who made his name as the flam­boy­ant Dr Frank-N-Furter in The Rocky Horror Show, has died aged 80. Reports say Curry died at his home in Los Angeles.

The ac­tor’s man­ager con­firmed to Variety that he died peace­fully. No cause of death is known as yet.

Ironically, Curry’s best-known roles all re­quired a con­sid­er­able phys­i­cal trans­for­ma­tion, to the ex­tent that he was un­recog­nis­able in all of them. Frank-N-Furter, Rocky Horror’s de­ranged sci­en­tist, saw him daubed in gar­ish makeup and wear­ing stock­ings and sus­penders, for the stage show (which de­buted in 1973) and the film adap­ta­tion (released in 1975). In the Ridley Scott-directed fan­tasy Legend (1985), he was equipped with gi­ant horns and scar­let skin as the Lord of Darkness. And in 1990, Curry was trans­formed with full-face clown pan­cake to play Pennywise in the TV minis­eries It, based on the Stephen King novel.

Alongside these show-stop­ping in­car­na­tions, Curry en­joyed a suc­cess­ful par­al­lel ca­reer in film, TV and the­atre. Born in 1946, he earned a de­gree in drama and English from the University of Birmingham in 1968 and em­barked on a stage ca­reer. His first sig­nif­i­cant role was in the orig­i­nal cast for the first West End pro­duc­tion of the con­tro­ver­sial mu­si­cal Hair, in 1968 — where he met Richard O’Brien, who would go on to cast him in his own mu­si­cal, The Rocky Horror Show. Curry went on to play Tristan Tzara in Tom Stoppard’s Travesties (taking over the role from John Hurt and Robert Powell), and Mozart in the Broadway trans­fer of Amadeus in 1980. Later high­lights in­cluded King Arthur in the Broadway run of the Monty Python mu­si­cal Spamalot.

Curry worked reg­u­larly in film, in roles in­clud­ing that of Robert Graves in Jerzy Skolimowski’s sin­is­ter The Shout, and later as Jeremy in the Ian McEwan-scripted po­lit­i­cal drama The Ploughman’s Lunch. Later, Curry found him­self work­ing largely in come­dies — in­clud­ing Home Alone 2: Lost in New York, Muppet Treasure Island, Scary Movie 2, and McHale’s Navy — be­fore pro­vid­ing voiceovers for nu­mer­ous an­i­mated films, such as Garfield: A Tail of Two Kitties, A Turtle’s Tale: Sammy’s Adventures and Ribbit.

Curry was also a reg­u­lar pres­ence on TV for more than four decades, play­ing William Shakespeare in John Mortimer’s six-part bi­og­ra­phy of the play­wright in 1978, Bill Sikes in the 1982 US TV adap­ta­tion of Oliver Twist, and the wiz­ard Trymon in the adap­ta­tion of Terry Pratchett’s The Colour of Magic in 2008. He voiced a stream of an­i­mated se­ries (Duckman, Sonic the Hedgehog, The Mask and Star Wars: The Clone Wars) and was cast in many guest roles in es­tab­lished se­ries, in­clud­ing Psych, Poirot and Criminal Minds.

I’ve had the op­por­tu­ni­ties,” he said to the Guardian in 2025, and I’m still show­ing up. You know? I think that’s what you have to do. You have to keep show­ing up.”

In 2012, Curry had a stroke while re­ceiv­ing a mas­sage and re­ceived brain surgery. He was left with mo­bil­ity is­sues and his short-term mem­ory was also af­fected. It was an odd thing, be­cause my fa­ther had a stroke and died very soon af­ter­wards,” he said. I knew I had to force my­self to re­lax and just take the op­por­tu­nity to float a lit­tle.”

In 2025, he re­leased his mem­oir Vagabond.

I’m very aware that I’m lucky,” Curry said last year. I’m as­ton­ished ac­tu­ally at how am­bi­tious I’ve been. I did­n’t think of my­self as am­bi­tious at all.”

Luke Evans, who played the role of Dr Frank-N-Furter in the re­cent Broadway pro­duc­tion of The Rocky Horror Show, paid trib­ute on Instagram. The ac­tor called him a force, a bright fierce flame” and an in­spi­ra­tion”. He added: There will only be one Tim Curry.”

Carol Burnett, who starred with Curry in Annie, also paid trib­ute on­line. Nobody could play lov­able vil­lains bet­ter than he could,” she wrote. He was a dear friend. I was blessed to know him.”

Curry’s Clue co-star Michael McKean shared on X: My last con­ver­sa­tion with Tim Curry was about life and love and laugh­ter. The grim phys­i­cal state he found him­self in has now ended, a bless­ing. Our bless­ing was that we had Tim Curry in our lives. RIP, my friend.”

Bloomberg - Are you a robot?

www.bloomberg.com

We’ve de­tected un­usual ac­tiv­ity from your com­puter net­work

To con­tinue, please click the box be­low to let us know you’re not a ro­bot.

Why did this hap­pen?

Please make sure your browser sup­ports JavaScript and cook­ies and that you are not block­ing them from load­ing. For more in­for­ma­tion you can re­view our Terms of Service and Cookie Policy.

Need Help?

For in­quiries re­lated to this mes­sage please con­tact our sup­port team and pro­vide the ref­er­ence ID be­low.

Block ref­er­ence ID:e143ee5b-a18f-11f1-bd9f-95489b1a47b3

Get the most im­por­tant global mar­kets news at your fin­ger­tips with a Bloomberg.com sub­scrip­tion.

6 RAG Architectures — and How to Avoid Over-Engineering

www.lighthousenewsletter.com

Hello, Rafael here - every week I cover in­ter­est­ing chal­lenges and de­vel­op­ments that I’ve come across re­cently through the lens of an en­gi­neer build­ing AI sys­tems.

Subscribe and get my weekly takes 👇

Nowadays, most peo­ple seem to over-en­gi­neer their RAG stack. They jump straight to em­bed­dings, vec­tor data­bases, and rerank­ing pipelines. Meanwhile, their users just want to find the doc that says How to re­set my pass­word.”

In en­gi­neer­ing, there’s al­ways the right tool for the right prob­lem. In AI Retrieval Systems it’s not dif­fer­ent.

Before we dive into recipes, let’s es­tab­lish when you should use each ap­proach. The key fac­tors are:

1. Data Freshness Requirements - Real-time up­dates (news, so­cial me­dia) fa­vor ap­proaches with easy re-in­dex­ing. Daily or weekly up­dates work well with hy­brid ap­proaches. A sta­ble cor­pus (monthly or quar­terly up­dates) makes pre-em­bed­ding sen­si­ble.

2. Corpus Characteristics - High churn (more than 10% changes daily) means you should avoid full pre-em­bed­ding. Stable doc­u­ments work fine with pre-em­bed­ding. Long-tail dis­tri­b­u­tion (90% never ac­cessed) means on-the-fly wins.

3. Query Patterns - Keyword-heavy queries should start with full-text search. Semantic or con­ver­sa­tional queries ben­e­fit from em­bed­dings. Mixed pat­terns need hy­brid ap­proaches.

4. Scale & Performance - Less than 1000 queries per day means sim­ple ap­proaches are suf­fi­cient. 1K to 10K queries per day re­quires se­lec­tive op­ti­miza­tion. More than 10K queries per day jus­ti­fies full op­ti­miza­tion.

5. Team Capabilities - No ML ex­per­tise means stay with full-text plus query rewrit­ing. Some ML ex­pe­ri­ence makes hy­brid search man­age­able. Having an ML team avail­able makes ad­vanced ap­proaches vi­able.

Now, let’s look at the recipe book. Start at the top. Move down only when you have data prov­ing you need to.

Good old BM25. Elasticsearch. Postgres full-text search. The stuff that ex­isted be­fore embedding” be­came a verb.

You’re just start­ing out. Your users write key­word-style queries (”pandas merge dataframe”). Exact matches mat­ter (”invoice #12345”). You want zero ML com­plex­ity. Your cor­pus has pro­pri­etary ter­mi­nol­ogy (more on this later).

Zero API costs. Fast (under 10ms). Easy to de­bug (you can see ex­actly why a doc­u­ment matched). Surprisingly ef­fec­tive (handles many use cases). No chunk­ing strat­egy needed — works with full doc­u­ments. No eval­u­a­tion com­plex­ity — easy to test and val­i­date. No model dep­re­ca­tion risk (BM25 does­n’t change).

Misses syn­onyms (”car” vs automobile”). Fails on se­man­tic queries (”How do I…?”). Can’t un­der­stand in­tent be­yond key­words.

In my ex­pe­ri­ence, this han­dles a sig­nif­i­cant por­tion of use cases. Don’t skip this step. You might be sur­prised how far you can get.

When you jump straight to em­bed­dings, you im­me­di­ately face ques­tions like: What chunk size? (512 to­kens? 1024?) What over­lap? (50 to­kens? 100?) Semantic chunk­ing or fixed-size? How do I eval­u­ate if my chunk­ing is good?

With full-text search, you skip all of this. Your doc­u­ments are your doc­u­ments. Search just works.

Thanks for read­ing Lighthouse AI! This post is pub­lic so feel free to share it.

Share

Use an LLM to trans­form messy user queries into clean key­word searches.

Most semantic search” prob­lems are ac­tu­ally query for­mu­la­tion prob­lems.

Users ask ques­tions con­ver­sa­tion­ally. Vocabulary mis­match (users say fix bugs”, docs say debugging”). You have in­ter­nal jar­gon (your frame­work called Atlas”). You want flex­i­bil­ity to it­er­ate quickly on query strate­gies.

~$0.001 per query (using GPT-4o-mini for query rewrit­ing)

An LLM can re­move stop­words (”how do I” be­comes noth­ing). It can add syn­onyms (”car” be­comes car au­to­mo­bile ve­hi­cle”). It can trans­late do­main terms (”speed up code” be­comes optimize per­for­mance”). It can de­com­pose com­plex queries (”read CSV and plot” be­comes [”read CSV, plot data”]). It can learn from your glos­sary (via sys­tem prompt).

With em­bed­dings, if re­sults aren’t good, you need to ad­just chunk­ing strat­egy, re-em­bed en­tire cor­pus, run re­gres­sion tests on your eval set, and hope it im­proved.

With query rewrit­ing, if re­sults aren’t good, you ad­just the sys­tem prompt. That’s it. Test im­me­di­ately.

Even bet­ter, you can cre­ate a loop:

The agent can it­er­ate, learn, and adapt — all with­out re-em­bed­ding any­thing.

Say your com­pany has a Python frame­work called Atlas.” If you use gen­eral-pur­pose em­bed­dings:

General em­bed­ding model (trained on in­ter­net): Atlas” = [vectors point­ing to­ward: Greek mythol­ogy, maps, ge­og­ra­phy] Your ac­tual Atlas docs = [vectors about data pro­cess­ing] Similarity score: 0.15 (terrible!)

The model has no idea your Atlas” ex­ists. It falls back to what it learned in train­ing. But with query rewrit­ing:

For pro­pri­etary terms, ex­act key­word match­ing beats se­man­tic un­der­stand­ing.

Use BM25 to get can­di­dates (top 50 – 100), then rerank with em­bed­dings (top 10).

BM25 is fast and great at key­word match­ing. Embeddings are good at se­man­tic un­der­stand­ing. Together, they cover each oth­er’s weak­nesses.

Users ask se­man­tic ques­tions (”find al­ter­na­tives to X”). BM25 plus query rewrit­ing alone is­n’t cut­ting it (you have data prov­ing this). You can tol­er­ate 100 – 500ms la­tency. Your cor­pus is rel­a­tively sta­ble (not chang­ing every minute).

Let’s do the math with cur­rent pric­ing (OpenAI text-em­bed­ding-3-small at $0.02 per 1M to­kens):

Embedding 50 docs per query (avg 500 to­kens each) means 50 docs × 500 to­kens = 25,000 to­kens

Embedding 50 docs per query (avg 500 to­kens each) means 50 docs × 500 to­kens = 25,000 to­kens

Cost: 25,000 × $0.00002 = ~$0.0005 per query. At 1,000 queries per day × 30 days = ~$15 per month.

Cost: 25,000 × $0.00002 = ~$0.0005 per query. At 1,000 queries per day × 30 days = ~$15 per month.

Actually pretty rea­son­able. But there’s a catch: la­tency.

Embedding 50 doc­u­ments on-the-fly adds 200 – 500ms per query. For user-fac­ing search, that’s no­tice­able. This is where the real trade-off lives — not cost, but speed.

When you in­tro­duce em­bed­dings, you need to de­cide how to chunk your doc­u­ments (fixed-size? se­man­tic? by sec­tion?). You need to de­ter­mine what chunk size and over­lap to use. You need to han­dle chunks that span im­por­tant con­text.

This adds com­plex­ity that pure full-text search avoids.

If your data changes fre­quently, why pay to re-em­bed every­thing?

High doc­u­ment churn (more than 10% of docs up­dated daily). Real-time con­tent (news, so­cial me­dia, live up­dates). You’re ex­per­i­ment­ing with em­bed­ding mod­els (no re-in­dex­ing needed). Data fresh­ness is crit­i­cal (documents must be up-to-date). Small K for rerank­ing (20 – 50 docs).

On-the-fly / on­line (1000 queries/​day, 50 docs/​query): - Embedding cost: ~$15/month (ongoing) - Storage: $0 (just store text) - Latency: 200 – 500ms per query - Freshness: Perfect (always cur­rent) - Model switch­ing: Easy (just change the API call)

Embedding mod­els get dep­re­cated.

OpenAI dep­re­cated text-em­bed­ding-ada-002 in fa­vor of text-em­bed­ding-3. If you pre-em­bed­ded 10 mil­lion doc­u­ments with the old model, you now need to re-em­bed all 10 mil­lion doc­u­ments with the new model, up­date your vec­tor data­base, run re­gres­sion tests on your eval­u­a­tion set, val­i­date that qual­ity did­n’t de­grade, han­dle the cu­tover pe­riod, and deal with any API changes.

You lit­er­ally just change one line of code. Done.

Latency. You’re em­bed­ding doc­u­ments on every query. This is only vi­able if you’re okay with 200 – 500ms la­tency, K is small (reranking 20 – 50 docs, not 500), and your use case fa­vors fresh­ness over speed.

Pre-embed fre­quently ac­cessed doc­u­ments (”hot tier”), em­bed rarely-ac­cessed doc­u­ments on-the-fly (”cold tier”).

Access pat­terns fol­low Pareto dis­tri­b­u­tion. 20% of docs get 80% of traf­fic.

Clear ac­cess pat­terns (some docs are ac­cessed way more than oth­ers). Medium-to-large cor­pus (more than 100K doc­u­ments). Mix of sta­ble and chang­ing con­tent. Need good la­tency for com­mon queries. Want to min­i­mize re-em­bed­ding on model up­dates.

Fast for 80% of queries (hit pre-em­bed­ded cache). Fresh for rarely-ac­cessed docs. Only re-em­bed hot tier when switch­ing mod­els (20% of cor­pus). Adapts to chang­ing ac­cess pat­terns. Best la­tency/​cost/​flex­i­bil­ity trade-off.

When your em­bed­ding model gets dep­re­cated:

Full pre-em­bed­ding: Re-embed 1M docs × $0.01 = $10,000 + down­time Hot/cold tiers: Re-embed 200K docs × $0.01 = $2,000 + min­i­mal down­time On-the-fly: Change one line of code = $0 + zero down­time

Embed every­thing up­front. Store in vec­tor data­base. Search with ANN (approximate near­est neigh­bors).

Very high query vol­ume (more than 10K queries per day). Need un­der 50ms la­tency. Very sta­ble cor­pus (under 5% churn per month). Access pat­tern is broad (no long tail). You have ML team to man­age in­fra­struc­ture.

Pre-embedding (1M docs): - One-time em­bed­ding: 1M docs × 500 to­kens × $0.00002 = $10 - Storage: 1M × 1536 dims × 4 bytes = 6GB (~$10 – 30/month) - Search la­tency: un­der 50ms (blazing fast!) - Freshness: Only as fresh as last re-in­dex

Documents change fre­quently (more than 10% per week). You’re ex­per­i­ment­ing with em­bed­ding mod­els. Low query vol­ume (under 1K queries per day). You haven’t tried sim­pler ap­proaches first.

This is where full pre-em­bed­ding hurts the most. When you need to switch mod­els, you face down­time (your search is de­graded while re-em­bed­ding), com­pute cost (re-embedding mil­lions of doc­u­ments), test­ing bur­den (full re­gres­sion test suite on new em­bed­dings), chunk­ing reeval­u­a­tion (maybe new model works bet­ter with dif­fer­ent chunk sizes?), and risk (what if the new model is worse for your do­main?).

This is overkill for most sys­tems. I’ve seen teams spend months op­ti­miz­ing their vec­tor data­base setup when query rewrit­ing would have solved 90% of their prob­lems.

This is overkill for most sys­tems. I’ve seen teams spend months op­ti­miz­ing their vec­tor data­base setup when query rewrit­ing would have solved 90% of their prob­lems.

But if you’re Pinterest, Shopify, or han­dling mas­sive scale with a sta­ble cor­pus, this is where you end up.

We’ve been dis­cussing sin­gle-in­tent queries: How do I merge dataframes?”

But real users ask stuff like: How do I read a CSV file, clean miss­ing data, and plot the re­sults?”

That’s three sep­a­rate in­tents. Searching for this as one query is like try­ing to find a restau­rant that serves pizza, sushi, and tacos. Good luck.

Modern agen­tic RAG sys­tems (Perplexity, ChatGPT search) han­dle this el­e­gantly:

Break down the query.

Route each sub-query op­ti­mally

Combine re­sults into co­her­ent an­swer

Each sub-query is fo­cused and pre­cise, lead­ing to bet­ter re­trieval. Parallel ex­e­cu­tion means lower la­tency (max, not sum). Adaptive rout­ing re­sults in lower cost (only com­plex queries pay for LLM). Structured out­put pro­vides bet­ter UX.

Without de­com­po­si­tion

LLM rewrit­ing en­tire com­plex query: $0.005

LLM rewrit­ing en­tire com­plex query: $0.005

Embedding 50 docs: $0.025

Embedding 50 docs: $0.025

Total: $0.03

Total: $0.03

With de­com­po­si­tion

Decompose: $0.001

Decompose: $0.001

Sub-query 1 (simple): $0

Sub-query 1 (simple): $0

Sub-query 2 (simple): $0

Sub-query 2 (simple): $0

Sub-query 3 (complex): $0.001

Sub-query 3 (complex): $0.001

Total: $0.002

Total: $0.002

15x cheaper, bet­ter qual­ity.

This is where agen­tic re­trieval re­ally shines. The agent can in­tel­li­gently de­cide which sub-queries need ex­pen­sive pro­cess­ing (embeddings) and which can be han­dled with cheap meth­ods (simple pre­pro­cess­ing + BM25).

Okay, you’ve read this far. You just want to know: What should I build?”

Start here: Do you have search at all? If not, build BM25 first. Seriously. Stop read­ing and build it. If you do have search, con­tinue.

Measure your base­line. Run your cur­rent search for 2 – 4 weeks and col­lect user feed­back. Are users happy with the re­sults? If yes, stop. You’re done. Go ship fea­tures. If no, con­tinue.

What’s the main com­plaint?

If users say Can’t find docs that clearly ex­ist,” try query rewrit­ing first. At $0.001 per query with zero re-in­dex­ing, it’s worth test­ing. Run an A/B test for 2 weeks. If you see good im­prove­ment, keep it and you’re done. If it’s not enough, con­tinue.

If users say Results are okay but not great,” A/B test hy­brid search (sparse plus em­bed­ding rerank). Is the added la­tency worth it? If yes, de­cide on im­ple­men­ta­tion. If your data changes fre­quently, use on-the-fly em­bed­ding. If you have clear hot docs, use hot/​cold tiers. If you have a sta­ble cor­pus and high scale, use full pre-em­bed­ding. If the la­tency is­n’t worth it, op­ti­mize query rewrit­ing fur­ther in­stead.

Actually Queryable Executables

fzakaria.com

I was pleas­antly sur­prised and happy to see that my ar­ti­cle Your ex­e­cutable is a SQLite data­base’ res­onated with peo­ple. It is a for­mat I have been think­ing about for a while, and the idea seems to have struck a chord with oth­ers.

A quick re­cap: SELF, a for­mat where the pro­gram is a SQLite data­base. We can use binfmt_misc to trig­ger a cus­tom in­ter­preter that maps the rows in the seg­ments table and jumps to the en­try point, and a whole class of bi­nary tool­ing col­lapses into SQL.

What keeps sur­pris­ing me is how hav­ing the file for­mat be a SQLite data­base keeps col­laps­ing every­thing into SQL. One idea that was im­me­di­ately ev­i­dent to my­self and oth­ers through com­ments: If the ex­e­cutable is a data­base, and a data­base is some­thing you can write to, can the run­ning pro­gram use it to also store its state? 🤔

Yes! 🤯 We can col­lapse not only a com­plete dis­tri­b­u­tion but all the state for every ap­pli­ca­tion into a sin­gle file, al­le­vi­at­ing the need for /var/ or /tmp/ or /home/ or any other filesys­tem. The pro­gram can store its own state in the same file it is run­ning from, and it can do so trans­ac­tion­ally.

self-httpd is a proof-of-con­cept web­server that does ex­actly that. It is a sin­gle file pro­gram ex­e­cuted from a data­base. The file con­tains the pro­gram, the web­site, the routes and all the vis­i­tor logs. All state is up­dated in the same SQLite file as the pro­gram it­self.

# Our server is a sin­gle file, and it is a SQLite data­base $ file server server: SQLite 3.x data­base, ap­pli­ca­tion id 1397050438, …

$ ./server –journal wal 8080 self-httpd: serv­ing 3 routes out of /srv/self/server self-httpd: lis­ten­ing on http://​0.0.0.0:8080 with 4 work­ers

$ curl -s lo­cal­host:8080 | head -1 <!doctype html>

# no­body has pressed the but­ton on that page yet $ sqlite3 server SELECT count(*) FROM press­es’ 0

$ curl -s -X POST -d press lo­cal­host:8080/​api/​press {“presses”:1,“button”:“press”}

# the ap­pli­ca­tion data is in­side the same data­base $ sqlite3 server SELECT id, at, but­ton FROM press­es’ 1|2026 – 08-25 03:11:28|press

# so was the GET that fetched the page in the first place $ sqlite3 server SELECT count(*) AS n, path FROM vis­its GROUP BY path’ 1|/ 1|/api/press

This web-server is live at https://​selfdb.exe.xyz.11If the site is not work­ing for you, sorry. I de­ployed it on their small­est tier. I in­cluded a screen­shot of the site just in case for pos­ter­ity!  It is one file, a SQLite data­base, and it is also the server. It is the web­site, it is the pro­gram, and it is the vis­i­tor log and state.

§Everything is my de­mon muse

I have a lot of ad­mi­ra­tion for the work of Justine Tunney, whose prior art red­bean: a web­server in a sin­gle file, built as an Actually Portable Executable with a self-ex­tract­ing ZIP archive, in­spired the idea.

SELF is many ways is less bril­liant. It re­lies on sim­pler tools to achieve some­thing very sim­i­lar but I’m amazed how much col­lapses into a sin­gle do­main: SQL.

Whereas, red­bean needs to in­clude an archive for­mat (ZIP), the data­base it­self is the con­tainer. Redbean pro­vides Lua hooks to ma­nip­u­late the re­sponses, whereas the equiv­a­lent in SELF is a new row in a han­dlers table.

INSERT INTO han­dlers VALUES (‘/api/busiest’, SELECT path, count(*) FROM vis­its GROUP BY path ORDER BY 2 DESC LIMIT 5’);

If red­bean is an Actually Portable Executable, this is an Actually Queryable Executable. One of them runs any­where, the other one you can SELECT from.

§All you need is argv[0]

How does the process get ac­cess to it­self? 🤔

For now, you can­not use /proc/self/exe.22Funny enough, the VFS Linux main­tainer re­cently landed sup­port for trans­par­ent binfmt_misc in the ker­nel, which would make /proc/self/exe point to the orig­i­nal file. I wrote about it here.  When binfmt_misc matches, the ker­nel does not ex­ecve your file at all , it ex­ecs the in­ter­preter, and hands it the path:

self-exec passes argv + 1 through to the pro­gram, so the pro­gram’s argv[0] is the path to the ex­e­cutable it­self. The in­ter­preter also re­leases its SQLite con­nec­tion be­fore jump­ing to the en­try point, so the pro­gram can open its own file and query it.

int main(int argc, char **argv) { sqlite3 *db; /* the file the ker­nel just ex­e­cuted */ sqlite3_open(argv[0], &db); … }

This is pretty un­re­stricted and mag­i­cal. You can read your own seg­ment table or a new table next to it. The writes per­sist across in­vo­ca­tions. ✨

§self-httpd

The web-server for our ex­am­ple is three ta­bles: routes, vis­its and presses. We will record every vis­i­tor and every but­ton press.

– the con­tent, added to the ex­e­cutable — af­ter it is com­piled and linked CREATE TABLE routes (path TEXT PRIMARY KEY, mime TEXT, body BLOB); — what the site col­lects, writ­ten back — into the ex­e­cutable while it runs CREATE TABLE vis­its (id INTEGER PRIMARY KEY, at TEXT, ua TEXT, path TEXT); CREATE TABLE presses (id INTEGER PRIMARY KEY, at TEXT, but­ton TEXT);

Building the ap­pli­ca­tion feels very un­re­mark­able and fa­mil­iar. We ex­e­cute DDL to cre­ate the ap­pli­ca­tion schema and INSERT the web­site.

# an or­di­nary ELF for now $ cc -O2 server.c -o server.elf $(pkg-config –libs sqlite3) # the same pro­gram, as rows $ elf2­self server.elf server $ sqlite3 server < site/​schema.sql $ sqlite3 server INSERT INTO routes VALUES (‘/index.html’, text/html’, read­file(‘site/​in­dex.html’))”

The as­set pipeline looks like a normal web­server” un­til you re­al­ize it’s query­ing it­self with SQL for the con­tent. Oh, and itself” is a SQLite data­base.

The page at https://​selfdb.exe.xyz shows a lot of fun ad­di­tional in­for­ma­tion be­sides the vis­i­tor log and but­ton presses. I in­cluded seg­ments, sym­bols and re­lo­ca­tions. Those are not baked in at built time, they are queried from it­self while run­ning.

§Editing a live site is a trans­ac­tion

Once you have the ca­pa­bil­ity to do ACID trans­ac­tions, in­ter­est­ing things be­come pos­si­ble. The web­server can edit its own con­tent while it is run­ning, and the ed­its are trans­ac­tional. The UPDATE is com­mit­ted to the same file as the pro­gram, and a ROLLBACK un­does it.

# change the run­ning site. no restart, no re­load, no de­ploy $ sqlite3 server UPDATE routes SET body = read­file(‘new.html’) WHERE path = /index.html’” $ curl -s lo­cal­host:8080 <!doctype html><h1>edited in place</​h1>

Since the file for­mat is SQLite we can also take ad­van­tage of the cor­ni­co­pea of tool­ing that ex­ists. sqld­iff will tell you ex­actly what a deploy did”, this can let us au­dit and iden­tify changes be­tween two ver­sions of the same pro­gram.

$ sqld­iff –summary yes­ter­day.server server routes: 1 changes, 0 in­serts, 0 deletes, 2 un­changed seg­ments: 0 changes, 0 in­serts, 0 deletes, 13 un­changed sym­bols: 0 changes, 0 in­serts, 0 deletes, 174 un­changed re­lo­ca­tions: 0 changes, 0 in­serts, 0 deletes, 99 un­changed

What about full-text search? FTS5 is a CREATE VIRTUAL TABLE away, so a web­server can in­dex its own pages, in­side it­self, and still be a web­server af­ter­wards:

$ sqlite3 server CREATE VIRTUAL TABLE search USING fts5(path, body); INSERT INTO search SELECT path, body FROM routes WHERE mime LIKE text/%’”

$ sqlite3 server SELECT path, snip­pet(search, 1, [’, ]’, …’, 6) FROM search WHERE search MATCH transaction’” /index.html|…Editing is a [transaction].</h2>

# still runs. it just knows about it­self now $ ./server 8080

None of that is ma­chin­ery I wrote. It is ma­chin­ery SQLite al­ready has, that a pro­gram in­her­its for free by be­ing a data­base.

All the rage was sta­tic site gen­er­a­tors, but the fu­ture is an ac­tu­ally queryable ex­e­cutable.

§Deploying is scp of one file

I am re­ally en­joy­ing the sim­plic­ity that seems to be pop­u­lar and her­alded by prod­ucts like exe.dev. People of­ten yearn to go back to the good old days” of scp and ssh to de­ploy a sin­gle file, and SELF is a for­mat that makes that pos­si­ble again, but bet­ter! Rather than just ship­ping an archive of PHP, we ship the whole sys­tem or ap­pli­ca­tion clo­sure down to the libc.

How would we make a de­ploy­ment if the data and code is in­ter­twined?

We can think of a re­de­ploy as a data mi­gra­tion, and the mi­gra­tion is two INSERTSELECT, be­cause the pro­gram and its data are the same file!

– the run­ning de­ploy­ment ATTACH /srv/self/server’ AS old; INSERT INTO vis­its (at, ua, path) SELECT at, ua, path FROM old.vis­its; INSERT INTO presses (at, but­ton) SELECT at, but­ton FROM old.presses;

Swap the file, restart, and the vis­i­tor log sur­vives the new build. You can even do this for the pro­gram it­self in re­verse. The seg­ments table is just like any other table. 😈

§Go press the but­ton

https://​selfdb.exe.xyz has a but­ton on it. Pressing it is an INSERT into the ex­e­cutable that served you the page

The code is at fza­karia/​selfdb if you are cu­ri­ous. It is prob­a­bly a bit half-baked, and def­i­nitely AI as­sisted, but that’s OK with me. I wanted to ex­plore this idea and see if it was fea­si­ble and what might be pos­si­ble.

I think I only scratched the sur­face of some of the fun pos­si­bil­i­ties. I am cu­ri­ous to see what oth­ers might do with it, and I would love to see a few more ex­am­ples of actually queryable ex­e­cuta­bles” in the wild.33One idea a friend sug­gested was dis­cov­ery over mul­ti­case DNS to spread pro­gram up­dates via trans­ac­tions.

Turns out that when we re-en­vi­sion what we con­sid­ered to be sim­ply a byte lay­out spec­i­fi­ca­tion was ac­tu­ally bet­ter off be­ing a data­base, a lot of ma­chin­ery we have been us­ing for decades sim­ply stops be­ing nec­es­sary. The pro­gram is the data­base, and the data­base is the pro­gram.

Never, ever un­der­es­ti­mate the im­por­tance of hav­ing fun”

– Randy Pausch

Never, ever un­der­es­ti­mate the im­por­tance of hav­ing fun”

– Randy Pausch

Ma connexion internet | Arcep

cartefibre.arcep.fr

Dans les zones moins denses et dans les poches de basse den­sité des zones très denses, le point de mu­tu­al­i­sa­tion (PM) est le point où un opéra­teur dé­ploy­ant le réseau FttH sur un ter­ri­toire donné donne ac­cès à son réseau aux autres opéra­teurs pour que ceux-ci puis­sent y pro­poser des of­fres à des­ti­na­tion de leurs clients. On ap­pelle zone ar­rière de point de mu­tu­al­i­sa­tion (ZAPM) le ter­ri­toire dont les lo­caux ont vo­ca­tion à être desservi par le réseau situé en aval d’un point de mu­tu­al­i­sa­tion donné.

Ainsi, cette vue per­met de vi­su­aliser le taux de cou­ver­ture FttH ZAPM par ZAPM, au sein :

des poches de basse den­sité des zones très denses ;

des zones moins denses.

Par ex­cep­tion, dans les poches de haute den­sité des zones très denses, l’ac­cès aux réseaux se fait en pied d’im­meu­ble, générale­ment à l’in­térieur des im­meubles. Par souci de lis­i­bil­ité, dans ces zones la vue des « ZAPM » pro­pose égale­ment un niveau de lec­ture in­ter­mé­di­aire sur des zones in­fra-com­mu­nales : le taux de cou­ver­ture FttH y est ainsi af­fiché à la maille des îlots « IRIS » défi­nis par l’In­see.

Dans cette vue, les ZAPM sont af­fichées lorsque le PM as­so­cié a été mis à dis­po­si­tion des opéra­teurs tiers. Si au­cun lo­cal en aval du PM n’est rac­cord­able (taux de cou­ver­ture à 0%), le con­tour de la ZAPM est dess­iné en pointil­lés. Attention, cer­taines ZAPM peu­vent ne pas être af­fichées si les opéra­teurs con­cernés n’ont pas en­voyé les don­nées cor­re­spon­dantes à l’Ar­cep.

Dans cette vue, les ZAPM sont af­fichées lorsqu’elles ont été mise en con­sul­ta­tion qu’elles soient cibles ou co­hérentes po­ten­tielle. Si au­cun lo­cal en aval du PM n’est rac­cord­able (taux de cou­ver­ture à 0%), le con­tour du PM est dess­iné en pointillé. Attention, cer­taines ZAPM peu­vent ne pas être af­fichées si les opéra­teurs con­cernés n’ont pas en­voyé les don­nées cor­re­spon­dantes à l’Ar­cep.

Lorsque la carte est zoomée au niveau du quartier, les adresses réper­toriées par les dif­férents opéra­teurs d’in­fra­struc­ture sont af­fichées, représen­tées par une pastille, dont la couleur varie suiv­ant l’é­tat d’a­vance­ment du dé­ploiement.

Ces in­for­ma­tions à l’adresse ne sont disponibles qu’à par­tir du T2 2018.

On dis­tingue vi­suelle­ment cinq caté­gories :

Études réal­isées :

Programmé :

Raccordable sur de­mande :

Raccordable sur de­mande (en cours de dé­ploiement) : Adresse rac­cord­able sur de­mande pour laque­lle la pose du PBO a été de­mandée par un opéra­teur com­mer­cial

En cours de dé­ploiement :

Déployé (raccordable) :

Des in­for­ma­tions plus pré­cises quant à l’é­tat du dé­ploiement adresse par adresse (statut de l’im­meu­ble et statut du PM) sont présentes dans l’info-bulle as­so­ciée à chaque adresse, selon la ter­mi­nolo­gie définie par le groupe Interop’Fibre

Les points représen­tants des adresses et/​ou des bâ­ti­ments sont po­si­tion­nés d’après les co­or­don­nées in­diquées par les opéra­teurs dans leurs IPE

Les don­nées à la maille de l’im­meu­ble et de la com­mune con­sulta­bles ici sont pub­liées sur la page data.gouv.fr sous Licence Ouverte.

La carte est pub­liée sous la li­cence CC-BY-SA.

Merchants of Insecurity

blog.happyfellow.dev

One Happy Fellow - blog

Posts About Contact Subscribe RSS

25 Aug, 2026

First, a PSA: Do NOT use Omarchy if you care about se­cu­rity of your ma­chine even a lit­tle bit.

You can’t pol­ish a turd

Omarchy 4.0 shipped with a col­lec­tion of se­cu­rity is­sues which I can only de­scribe as re­gret­table (because I promised my mum I would swear less). There are bangers like video ti­tle bash in­jec­tion or all no­ti­fi­ca­tions be­ing able to run ar­bi­trary bash on your ma­chine.

All pro­jects have se­cu­rity is­sues but not all pro­jects have such pre­dictable se­cu­rity is­sues. We know how to deal with un­trusted in­puts. We know we should not use AI-generated bash scripts for pro­cess­ing un­trusted in­put, par­tic­u­larly with seem­ingly no re­view.

And you can’t get to a rea­son­ably se­cure sys­tem by start­ing with a pile of bash slop and hop­ing oth­ers will catch and fix the is­sues be­fore they are ex­ploited.

Simply put, Omarchy does­n’t treat se­cu­rity as im­por­tant. They do role-play tak­ing se­cu­rity se­ri­ously but their de­vel­op­ment prac­tices and the ease with which they speedrun decades of se­cu­rity is­sues and in­vent new ones tell us much more about the se­cu­rity of Omarchy than Security Team an­nounce­ments.

Lies or mar­ket­ing?

DHH loves to say he’s mak­ing the year of Linux on desk­top hap­pen. How Omarchy is the dis­tro peo­ple should use. Showing how pol­ished the ex­pe­ri­ence is. Essentially, he’s good at mar­ket­ing Omarchy.

On the se­cu­rity front, he high­lights the se­cu­rity team’s ef­forts in the re­cent point re­lease. A long list of re­solved se­cu­rity is­sues sure looks im­pres­sive if you start with Swiss cheese of an op­er­at­ing sys­tem.

I think the mar­ket­ing has turned into ly­ing, be­ing disin­gen­u­ous. An hon­est at­ti­tude would be to say that they care about it­er­at­ing on their dot­files much more than about ba­sic se­cu­rity of the sys­tem.

DHH si­lences his crit­ics with fake pos­i­tiv­ity, with a let’s fuck­ing do it” at­ti­tude. But the re­al­ity is that Omarchy is a pro­ject which does­n’t treat se­cu­rity se­ri­ously and I would­n’t be sur­prised if many com­pa­nies ban its use.

Just let peo­ple en­joy things, jeez

I’m not stop­ping any­one. I don’t like when the pub­lic per­cep­tion of how much risk some­one is tak­ing is dis­con­nected from the ac­tual risk of us­ing a pro­ject like Omarchy.

The team does­n’t look in­ter­ested in ac­cu­rately ex­plain­ing it to their users. I don’t like peo­ple be­ing de­ceived into hurt­ing them­selves, hence this post.

Peace ✌️

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.