10 interesting stories served every morning and every evening.

Writing by Hand is Good for your Brain - Here's how to do it

nealstephenson.substack.com

Because I am known to write us­ing a foun­tain pen on pa­per, a num­ber of peo­ple have pointed me to this post and its un­der­ly­ing re­search. I won’t re­hash what is said in those sources, but the gist of it is that when you write things down by hand you’re re­cruit­ing more of your brain, which is a good thing.

I’m not an ex­pert on how the brain works, but I can say that, when writ­ing by hand, one is con­tin­u­ally solv­ing a se­ries of small prob­lems hav­ing to do with the spac­ing of words, how let­ters are con­nected, the cross­ing of the let­ter t (sometimes more than one in the same word) and the dot­ting of the let­ters i and j, and how to ac­com­plish all of those things through co­or­di­nated move­ments not just of the fin­gers but of the whole arm. All of that has to be in­te­grated in real time with what­ever is hap­pen­ing on a more ab­stract level in the brain’s pro­cess­ing of ideas and im­agery.

Concurrently I have been fol­low­ing dis­course on Reddit and other sources about how wide­spread use of AI has forced ed­u­ca­tors to re­turn to the long-aban­doned prac­tice of hav­ing their stu­dents take ex­ams in per­son by writ­ing things out long­hand in blue books. This has cre­ated new chal­lenges for stu­dents who never re­ally learned how to write by hand, and for teach­ers who can’t make sense of their stu­dents’ ter­ri­ble hand­writ­ing.

About twenty-five years ago I stopped com­pos­ing at the key­board and switched over to foun­tain pen on pa­per. Since then I have writ­ten many thou­sands of pages that way. The man­u­script of The Baroque Cycle was a stack of hand­writ­ten pages 42 inches high, which for a time was on dis­play at the Museum of Science Fiction in Seattle. With the ex­cep­tion of The Rise and Fall of D.O.D.O., which I co-wrote with Nicole Galland by email­ing Word files back and forth, every book I’ve writ­ten since then has been com­posed with foun­tain pen on pa­per.

Every so of­ten, when I’m sign­ing books at a book tour ap­pear­ance, some­one will come up to me and say some­thing like you must have writer’s cramp!” or is your hand sore yet?” I never have the time to pro­vide a full an­swer. If I did, how­ever, my an­swer would be that never, at any time dur­ing a quar­ter of a cen­tury dur­ing which I have spent a sub­stan­tial frac­tion of each work­ing day writ­ing by hand, have I ex­pe­ri­enced even the faintest traces of so-called writer’s cramp” or any other such hob­gob­lins.

Yet I can re­mem­ber get­ting a sore hand when I was a kid writ­ing out as­sign­ments in school. Many peo­ple prob­a­bly re­mem­ber such ex­pe­ri­ences and as­sume, rea­son­ably enough, that it’s a nat­ural con­se­quence of writ­ing by hand for any length of time. This is not the case.

Here are some fairly sim­ple dos and don’ts for peo­ple who want to reap the ben­e­fits of writ­ing by hand.

It’s pretty ob­vi­ous that you’re go­ing to get tired faster if your mus­cles have to ex­ert more force. Writing with a pen­cil re­quires sig­nif­i­cantly more force than writ­ing with a good pen. Old-school ball­points with thick ink are no bet­ter. You can see vi­sual ev­i­dence of this if you flip over a sheet of pa­per on which you’ve been writ­ing with a pen­cil or an old ball­point. The pa­per will bear a vis­i­ble im­print where it was pressed down by the writ­ing in­stru­ment. Often that will con­tinue down into the stack of pa­per be­neath. That’s be­cause you had to push hard. This does­n’t hap­pen with a foun­tain pen. If the nib is work­ing prop­erly you need to ex­ert very lit­tle force. The nib is ba­si­cally skat­ing on the lit­tle lake of ink that it has just laid down.

Pains me to say it, but roller­ball gel pens are about as good as foun­tain pens on this front.

It might then seem rea­son­able to think that writ­ing with a sty­lus on an iPad or sim­i­lar would be best, since no force is needed and fric­tion is min­i­mized. I don’t think this is true. A small amount of fric­tion is ac­tu­ally de­sir­able. You don’t want the tip of the writ­ing in­stru­ment to skid out of con­trol. Your brain and your lit­tle hand mus­cles are re­ly­ing on a lit­tle bit of fric­tion. Since I’m writ­ing this dur­ing the World Cup, I’ll make a soc­cer anal­ogy. Soccer play­ers have spent many hours drib­bling balls across play­ing fields, and they’ve in­ter­nal­ized the physics—they know about how far the ball is go­ing to travel when they kick it a cer­tain way, and how of­ten they need to give it an­other kick to keep it mov­ing. If you put them on a gi­ant, fric­tion­less air hockey table, all of that knowl­edge would be­come use­less. Every touch on the ball would send it out of con­trol. Dribbling the ball down the field would be­come more tir­ing be­cause they’d have to be mak­ing con­tin­ual ef­forts to con­trol the bal­l’s move­ment. Relying on a lit­tle bit of fric­tion re­duces the amount of men­tal and phys­i­cal ef­fort.

The com­bi­na­tion of foun­tain pens and pa­per em­bod­ies a bal­ance that has been worked out over a long span of time by peo­ple who write a lot. This phe­nom­e­non is called tooth” by afi­ciona­dos. Removing fric­tion by us­ing a hard sty­lus on glass will ac­tu­ally make the process more tir­ing.

Too much fric­tion, and too lit­tle fric­tion, are both more tir­ing than just a lit­tle bit of fric­tion, and that’s the bal­ance that is re­flected in the foun­tain pen/​pa­per tech­nol­ogy.

Rresults vary when you use var­i­ous pens on var­i­ous kinds of pa­per. Generally I get the worst re­sults on cheap printer pa­per, be­cause it wicks ink out of the nib too fast, and so cre­ates fat, blurry lines. Often I have the same prob­lem with yel­low le­gal pads. But al­most any pa­per in a blank note­book, or higher-grade printer pa­per with at least 25% cot­ton con­tent, works fine. I’ve learned over time that some of my foun­tain pens work bet­ter with cer­tain kinds of pa­per than oth­ers, so I match them up with­out hav­ing to think about it too hard.

Here’s a 300 dpi scan of tests I did with three dif­fer­ent pens on var­i­ous types of pa­per. You might have to zoom in to see much dif­fer­ence.

The pen on the left is a Jorg Hysek with a wide nib, and you can see that the cheap printer pa­per soaked up a lot of ink and left a thicker, fuzzier line. The le­gal pad was­n’t much bet­ter. Everything else ba­si­cally worked. The 100% cot­ton pa­per is from a box I pur­chased a long time ago - it was mar­keted for print­ing re­sumes, back in the days when peo­ple printed re­sumes. It is the tooth­iest of all these pa­pers and felt no­tice­ably scratch­ier. I guess it goes with­out say­ing that fancy Italian pa­per is the best, but the comp book and mole­sk­ine work per­fectly well with just about any pen.

(For those scor­ing at home, the mid­dle pen is a Diplomat Aero and the one on the right is a Monteverde Invincia)

If the pa­per is thin, writ­ing on one side can bleed through to the other, so the re­sults can be slightly harder to read if you write on both sides. Which leads me to:

The ecosys­tem is­n’t go­ing to col­lapse if you use more pa­per. It’s cheap. Focus on what’s im­por­tant here: your brain and your time. Write on one side. Trying to cram more words into a sheet will take you out of your nat­ural and com­fort­able writ­ing style and make you tired. Just buy a shit­load of pa­per or note­books or what­ever it is you want to use, and use it.

There’s a rea­son cur­sive was in­vented. Don’t even think about not us­ing it. It is far less tir­ing than print­ing one let­ter at a time. I learned cur­sive as a child. Then I went for many years with­out us­ing it much, and for­got some of it. Later I re-learned it by sit­ting in my kid’s el­e­men­tary school class­room dur­ing a par­ent-teacher con­fer­ence and ex­am­in­ing the forms printed on a long strip above the chalk­board (I still re­mem­bered how to do the lower-case let­ters, but I had for­got­ten some of the cap­i­tals).

Legibility was more im­por­tant back in the day when writ­ten doc­u­ments had to be read by other peo­ple. Hence the need for ex­act­ing pen­man­ship, taught in schools to long-suf­fer­ing chil­dren. This is prob­a­bly the source of a lot of angst around writer’s cramp and ink dis­as­ters. Today, if you’re writ­ing things down with ink on pa­per, you’re prob­a­bly writ­ing just for your­self, or per­haps for fam­ily mem­bers who can learn to rec­og­nize your hand­writ­ing.

To judge from the way peo­ple talk, a lot of them have mem­o­ries of foun­tain pen dis­as­ters where ink got all over the place for some rea­son. Or per­haps it’s just gen­er­a­tional trauma, handed down in an oral tra­di­tion. If the pen is work­ing cor­rectly, ink can only come out of it so fast. A cou­ple of rare ex­cep­tions:

If the pen’s ink reser­voir is partly empty, so that it con­tains an air bub­ble, and if it’s po­si­tioned nib down, then, when you go up in an air­plane, the bub­ble will ex­pand as the am­bi­ent pres­sure drops, forc­ing ink out the nib. Once I fig­ured that out, I got in the habit of mak­ing sure my pens were po­si­tioned nib up when tak­ing off in an air­plane. If I have time I’ll also re­fill the pen be­fore de­par­ture, to min­i­mize the size of the air bub­ble.

Sometimes if a pen gets dirty, or if the nib is some­how dam­aged, the ink will stop com­ing out and you can restart it by giv­ing it a lit­tle shake. If you do it just right, the ink flow restarts with­out in­ci­dent, but if you overdo it, a few drops of ink might shoot out onto the page and be­come blots. This sce­nario hap­pens a few times of year for me, only with one pen that has this prob­lem. I blot it with a piece of scrap pa­per and move on.

Just have note­books ly­ing around, or on your per­son. Write gro­cery lists, doo­dles, notes on meet­ings, to-do lists, or stray ideas. Journal. Copy out good lines from books. Anything that has your men­tal fo­cus will have a more en­dur­ing pres­ence in your brain if you write it down.

I am left handed. I have never had any trou­ble with my hand smear­ing the ink. Yet every con­ver­sa­tion I have about foun­tain pens leads to some­one claim­ing that it can never work for them be­cause they are left handed. I have no idea what they’re talk­ing about. When I was a child, writ­ing at length with pen­cil, the side of my hand some­times be­came gray from graphite picked up as my hand rubbed across the page. And some­times I have got ink on my hand when us­ing a ball­point pen that left an ink glob on the pa­per. But with foun­tain pens it’s easy to find a pen/​pa­per com­bi­na­tion such that the ink soaks into the pa­per and dries quickly enough that it does­n’t smudge when you’re writ­ing the next line. Here’s a sim­ple demon­stra­tion of dry­ing time and how it works with two pens: first a foun­tain pen and then a Pilot G-2 gel pen.

Obviously the Pilot gel pen ink dries faster, and so that might be a bet­ter choice for peo­ple who are re­ally wor­ried about smudg­ing.

Most mod­ern pens al­low you to choose be­tween us­ing pre­loaded plas­tic ink car­tridges and a plunger that en­ables you to draw ink up out of a bot­tle by hand. I use both. Start with the ink car­tridges, es­pe­cially if you travel. There’s no need to com­pli­cate mat­ters by mess­ing around with bot­tles. Since I do a lot of work from one lo­ca­tion, I have a cor­ner of a table­top set up there with ink bot­tles and a folded-up pa­per towel for wip­ing off the nib af­ter it’s filled (I have been us­ing the same pa­per towel for about twenty years). In the­ory this works bet­ter in the long term be­cause it al­lows you to flush the nib by forc­ing ink in and out of it a cou­ple of times when­ever you re­fill. In prac­tice I see no dif­fer­ence at all - pens that I re­fill with car­tridges don’t get clogged.

Even if every­thing works per­fectly you’ll end up with the oc­ca­sional ink-smudged fin­ger. It will wash off quickly - the ink is wa­ter-sol­u­ble. Until then, con­sider it a mark of dis­tinc­tion.

If you’re new to this I think it makes most sense to start by con­sid­er­ing what kind of pa­per is go­ing to work best in your life. Are you writ­ing loose­leaf, or in note­books? Legal pads? Blue books? Remember, it’s okay to use lots of pa­per, so pick some­thing that is­n’t too pre­cious and that is easy to re­plen­ish. I use a lot of mole­sk­ine note­books and Mead comp books, which I can buy in bulk on­line. For com­pos­ing fic­tion I use fancy loose­leaf pa­per.

If you have ac­cess to a store where they sell foun­tain pens, take some of that pa­per there and see what works best. If you’re work­ing with cheaper, thin­ner pa­per, start with finer nibs and work up to fat­ter ones un­til you start to see bleed-through.

Buy cheaper pens un­til you know what you like. I doubt there’s much of a dif­fer­ence be­tween cheaper and more ex­pen­sive foun­tain pens in terms of their ac­tual per­for­mance. What you’re pay­ing for, in an ex­pen­sive pen, is fancy ma­te­ri­als and styling. For ex­am­ple, if you look at the Pilot Vanishing Point line of pens - an in­ge­nious foun­tain pen that you can click, like an old-fash­ioned ball­point, to re­tract the nib in­side the bar­rel - fancier ver­sions cost five times as much as the base model.

In all hon­esty, the Pilot G-2 gel pens are go­ing to give you 80% of what you could ex­pect from a foun­tain pen for min­i­mal cost.

On the other hand, a ten-pack of Pilot G-2 gel pens goes for about twenty bucks. For the same amount you can buy a sim­ple but com­pletely ser­vice­able foun­tain pen that will last longer than you will.

No posts

Just a moment...

www.politico.com

AI Companies Are Trying to Hide a Staggering Amount of Debt

futurism.com

Sign up to see the fu­ture, to­day

Sign up to see the fu­ture, to­day

Can’t-miss in­no­va­tions from the bleed­ing edge of sci­ence and tech

AI com­pa­nies are pour­ing un­told bil­lions of dol­lars into enor­mous data cen­ters in their ef­forts to sus­tain in­creas­ingly com­plex and re­source-in­ten­sive AI mod­els.

It’s an ex­tremely costly un­der­tak­ing built on seem­ingly bot­tom­less hype — and a moun­tain of debt. As Japanese fi­nan­cial news­pa­per Nikkei Asia found in a re­cent in­ves­ti­ga­tion, just five US tech gi­ants — Alphabet, Microsoft, Amazon, Meta, and Oracle — are hid­ing an es­ti­mated $1.65 tril­lion in debt that does­n’t ap­pear on bal­ance sheets. That’s even more than the $1.35 tril­lion in debt the five com­pa­nies of­fi­cially re­ported in their fi­nan­cial data for the most re­cent quar­ter.

Meta alone has amassed around $420 bil­lion in off-bal­ance-sheet debt, ac­cord­ing to Nikkei, high­light­ing how pre­car­i­ous the AI in­dus­try’s steep in­vest­ment in AI has be­come, and in­spir­ing com­par­isons to en­ergy com­pany Enron, which col­lapsed in spec­tac­u­lar fash­ion in 2001 be­cause of sim­i­lar debts hid­den be­hind shell com­pa­nies. Like Enron, they’re us­ing spe­cial pur­pose ve­hi­cles, or off-bal­ance sheet arrange­ments such as legally dis­tinct sub­sidiaries, as a way to make their fi­nan­cial re­port­ing look health­ier than it ac­tu­ally is — of­ten a glar­ing sign that some­thing is deeply amiss be­hind the scenes.

The ac­count­ing treat­ment it­self is in fash­ion,” tech­ni­cal ac­count­ing con­sul­tant Tom Selling told Bloomberg. But what if one of these com­pa­nies was a house of cards and was prop­ping it­self up with this ac­count­ing treat­ment? To me, that’s the risk.”

Experts con­tinue to warn of an AI bub­ble, not­ing the enor­mous and widen­ing gulf be­tween com­pany val­u­a­tions and their com­par­a­tively measly prof­its. The lat­est news will do lit­tle to quiet crit­ics who say the sit­u­a­tion is more dire than the com­pa­nies’ of­fi­cial bal­ance sheets sug­gest.

To keep up with the on­go­ing AI race, tech gi­ants are com­mit­ting vast sums to build out large-scale data cen­ter pro­jects, a long-term bet that may — or may not — pay off. They’re also sell­ing new shares to raise new funds, as Nikkei re­ports, which could lead to eq­uity di­lu­tion and a drop in in­vestor con­fi­dence.

That could make them even more vul­ner­a­ble if the AI bub­ble does pop, or the in­dus­try fails to gen­er­ate enough de­mand to jus­tify the data cen­ter con­struc­tion frenzy.

The pres­sure is on: four of the five com­pa­nies Nikkei an­a­lyzed are set to re­port sec­ond quar­ter earn­ings in the com­ing days and weeks. We’ll be watch­ing.

More on the AI bub­ble: There’s a Gigantic Problem at the Heart of the AI Industry That Could Cause the Whole Thing to Collapse

Proposal for Assembly 2026: Disallow cryptocurrency projects

codeberg.org

Owner

Copy link

Copy link

Need to be care­ful with word­ing like this. If you are go­ing to pro­vide ex­am­ples you need to make it clear it is not an ex­haus­tive list:

Content that harms the rep­u­ta­tion of Codeberg, such as - but not lim­ited to - cryp­tocur­rency re­lated pro­jects.”

Need to be care­ful with word­ing like this. If you are go­ing to pro­vide ex­am­ples you need to make it clear it is not an ex­haus­tive list:

Content that harms the rep­u­ta­tion of Codeberg, such as - but not lim­ited to - cryp­tocur­rency re­lated pro­jects.”

Author

Owner

Copy link

The text is now as-is be­cause it was send out for votes. Small clar­i­fi­ca­tions can be made af­ter­wards by Presidium or Board. The whole spirit of the vote makes it clear this is a not lim­ited to” case.

The text is now as-is be­cause it was send out for votes. Small clar­i­fi­ca­tions can be made af­ter­wards by Presidium or Board. The whole spirit of the vote makes it clear this is a not lim­ited to” case.

First-time con­trib­u­tor

Copy link

Is there a de­f­i­n­i­tion of cryptocurrency-related” some­where?

Is there a de­f­i­n­i­tion of cryptocurrency-related” some­where?

Author

Owner

Copy link

This has passed.

This has passed.

![image](/attachments/73b8c345-cb43 – 44b9-b172 – 5c76f521a5e0)

Gusted

ref­er­enced this pull re­quest from a com­mit 2026 – 07-22 02:02:29 +02:00

First-time con­trib­u­tor

Copy link

Fk hell I just moved to a forge that banned bit­coin! Is this a joke??? https://​blog.code­berg.org/​we-stay-strong-against-hate-and-ha­tred.html

First-time con­trib­u­tor

Copy link

Please ex­plain how do cryp­tocur­rency pro­jects harm code­berg’s rep­u­ta­tion

https://​fo­rum.code­berg.org/​d/​82-tak­ing-a-stance-against-cryp­tocur­rency The page you re­quested could not be found.”

https://​fo­rum.code­berg.org/​d/​82-tak­ing-a-stance-against-cryp­tocur­rency The page you re­quested could not be found.”

Codeberg/Community#794 Codeberg/Community#2184 These do­mains are strongly as­so­ci­ated with fraud­u­lent ac­tiv­i­ties and high-risk in­vest­ments

Codeberg/Community#794 Codeberg/Community#2184 These do­mains are strongly as­so­ci­ated with fraud­u­lent ac­tiv­i­ties and high-risk in­vest­ments

Not all of them are about it. First of all, in con­text of so called code forges, this is a tech. What kind of headache do you have, that you judge the whole group by iso­lated cases, and block ANY such pro­jects, even those that have real tech­ni­cal value?

Please ex­plain how do cryp­tocur­rency pro­jects harm code­berg’s rep­u­ta­tion

> https://​fo­rum.code­berg.org/​d/​82-tak­ing-a-stance-against-cryp­tocur­rency The page you re­quested could not be found.”

> Codeberg/Community#794 > Codeberg/Community#2184 > These do­mains are strongly as­so­ci­ated with fraud­u­lent ac­tiv­i­ties and high-risk in­vest­ments

Not all of them are about it. First of all, in con­text of so called code forges, this is a tech. What kind of headache do you have, that you judge the whole group by iso­lated cases, and block ANY such pro­jects, even those that have real tech­ni­cal value?

First-time con­trib­u­tor

Copy link

When some pro­jects were trans­ferred over, you started be­hav­ing strangely.

When some pro­jects were trans­ferred over, you started be­hav­ing strangely.

First-time con­trib­u­tor

Copy link

While I re­spect this seems to have been a com­mu­nity de­ci­sion (I also de­spise the amount of fraud com­ing from the crypto space), this does set quite a con­cern­ing prece­dent, and makes me a lit­tle ner­vous to con­tinue rec­om­mend­ing Codeberg.

Banning an en­tire cat­e­gory of soft­ware based on bad ac­tors within that cat­e­gory is ex­treme, and pre­vents any healthy crypto pro­jects from emerg­ing here.

While I re­spect this seems to have been a com­mu­nity de­ci­sion (I also de­spise the amount of fraud com­ing from the crypto space), this does set quite a con­cern­ing prece­dent, and makes me a lit­tle ner­vous to con­tinue rec­om­mend­ing Codeberg.

Banning an en­tire cat­e­gory of soft­ware based on bad ac­tors within that cat­e­gory is ex­treme, and pre­vents any healthy crypto pro­jects from emerg­ing here.

First-time con­trib­u­tor

Copy link

I’m work­ing on a pro­ject aimed at bring­ing pri­vacy, se­cu­rity, and au­ton­omy to at risk peo­ple groups. The lan­guage in this mo­tion means I can no longer host it here. Is this what the Codeberg com­mu­nity voted for? The short-sight­ed­ness and in­com­pe­tency is mind blow­ing. What do you call it when a group of peo­ple come to­gether to weaponize their hate against a whole cat­e­gory of de­vel­op­ers? Anyone?

I’m work­ing on a pro­ject aimed at bring­ing pri­vacy, se­cu­rity, and au­ton­omy to at risk peo­ple groups. The lan­guage in this mo­tion means I can no longer host it here. Is this what the Codeberg com­mu­nity voted for? The short-sight­ed­ness and in­com­pe­tency is mind blow­ing. What do you call it when a group of peo­ple come to­gether to weaponize their hate against a whole cat­e­gory of de­vel­op­ers? Anyone?

First-time con­trib­u­tor

Copy link

I don’t even un­der­stand the logic be­hind this? Because some cryp­tocur­ren­cies are shady and bad, every sin­gle crypto pro­ject should not be al­lowed onto Codeberg? What if some­one is study­ing blockchains and want to im­ple­ment their own crypto? This is ex­tremely in­sane to me.

I don’t even un­der­stand the logic be­hind this? Because some cryp­tocur­ren­cies are shady and bad, every sin­gle crypto pro­ject should not be al­lowed onto Codeberg? What if some­one is study­ing blockchains and want to im­ple­ment their own crypto? This is ex­tremely in­sane to me.

First-time con­trib­u­tor

Copy link

I’m sure you have some morally high rea­sons to stand against cryp­tocur­rency, but the illicit trade” and evasion of sanc­tions” cited by source­hut (since you seem to base your de­ci­sion on it) are also what al­lows reg­u­lar peo­ple, in­clud­ing LGBTQ+ peo­ple, liv­ing in sanc­tioned coun­tries (which also, what a sur­prise, turn out to be un­safe for LGBTQ+ folk a lot of the time), to buy goods and send/​re­ceive money from abroad with­out be­ing pros­e­cuted by their gov­ern­ments (hi for­eign agent laws! hi extremism” and terrorism” laws!).

I’m sorry, anti-war trans­gen­der per­son stuck in Russia, but from our moral stance, you should­n’t be able to pur­chase HRT from a lab us­ing your XMR wal­let. nor should your friend be able to pay for their for­eign VPN VDS that they use to by­pass the in­ter­net re­stric­tions in USDT. the pro­jects you used for this were hosted on Codeberg and not some other plat­form? well, too bad, they’ll have to go some­place else that minds your ex­is­tence or is wel­com­ing to cryp­tocur­rency as a whole, and you will wait. you and the tools you use will move to a greedy cor­po­rate host­ing that is likely to im­pose its own re­stric­tions on you in the fu­ture, or, even bet­ter, move to a less re­li­able self-hosted op­tion, one per each tool to make it less main­tain­able and less ac­ces­si­ble.

by tak­ing this stance, at least from my per­spec­tive, you’re just pro­ject­ing your morally high delu­sion of dirty il­licit 3rd world crypto scam­mers that are dam­ag­ing the moral pu­rity of Codeberg by… host­ing code for their pro­jects here?! which, mind you, al­most all the time will just con­tain tools, tools to do good or bad. do you want to ban BitTorrent-related pro­jects from Codeberg next be­cause they are mostly used to get il­le­gal ac­cess to un­li­censed dig­i­tal goods and ser­vices (piracy)” and also take up world band­width and com­pute? how about ban­ning YouTube down­load­ers af­ter those? hey, let’s make it clear that Codeberg will not stand a sin­gle repo on its plat­form that in­volves en­crypted mes­sag­ing: you know only crim­i­nals use Matrix, right?

I’m sure you have some morally high rea­sons to stand against cryp­tocur­rency, but the illicit trade” and evasion of sanc­tions” cited by source­hut (since you seem to base your de­ci­sion on it) are also what al­lows reg­u­lar peo­ple, in­clud­ing LGBTQ+ peo­ple, liv­ing in sanc­tioned coun­tries (which also, what a sur­prise, turn out to be un­safe for LGBTQ+ folk a lot of the time), to buy goods and send/​re­ceive money from abroad with­out be­ing pros­e­cuted by their gov­ern­ments (hi for­eign agent laws! hi extremism” and terrorism” laws!).

I’m sorry, anti-war trans­gen­der per­son stuck in Russia, but from our moral stance, you should­n’t be able to pur­chase HRT from a lab us­ing your XMR wal­let. nor should your friend be able to pay for their for­eign VPN VDS that they use to by­pass the in­ter­net re­stric­tions in USDT. the pro­jects you used for this were hosted on Codeberg and not some other plat­form? well, too bad, they’ll have to go some­place else that minds your ex­is­tence or is wel­com­ing to cryp­tocur­rency as a whole, and you will wait. you and the tools you use will move to a greedy cor­po­rate host­ing that is likely to im­pose its own re­stric­tions on you in the fu­ture, or, even bet­ter, move to a less re­li­able self-hosted op­tion, one per each tool to make it less main­tain­able and less ac­ces­si­ble.

by tak­ing this stance, at least from my per­spec­tive, you’re just pro­ject­ing your morally high delu­sion of dirty il­licit 3rd world crypto scam­mers that are dam­ag­ing the moral pu­rity of Codeberg by… host­ing code for their pro­jects here?! which, mind you, al­most all the time will just con­tain tools, *tools* to do good or bad. do you want to ban BitTorrent-related pro­jects from Codeberg next be­cause they are mostly used to get il­le­gal ac­cess to un­li­censed dig­i­tal goods and ser­vices (piracy)” and also take up world band­width and com­pute? how about ban­ning YouTube down­load­ers af­ter those? hey, let’s make it clear that Codeberg will not stand a sin­gle repo on its plat­form that in­volves en­crypted mes­sag­ing: you know only crim­i­nals use Matrix, right?

First-time con­trib­u­tor

Copy link

The vague­ness of the term such as cryp­tocur­rency re­lated pro­jects.” has ma­te­ri­ally dam­aged Codebergs rep­u­ta­tion in my eyes, and given the com­ments above, I am not alone. Therefore un­der its own con­struc­tion that Content that harms the rep­u­ta­tion of Codeberg” should have been self de­feat­ing and not al­lowed un­der its own pol­icy.

Examples of POTENTIAL crytpocurrency re­lated pro­jects”:

ZK Proof Libraries.

Blake and SHA HASH Libraries.

PQ Crypto, ED25519 or ECDSA (secp256k1)

ANYTHING to do with LibP2P or sim­i­lar li­braries.

ANYTHING to do with BFT Consensus or other con­sen­sus al­go­rithms.

So The only safe pol­icy is to as­sume that Codeberg is ba­si­cally anti-cryp­tog­ra­phy. Because most all cryp­tog­ra­phy at some level of re­la­tion­ship be­comes a cryptocurrency re­lated pro­ject”.

And in who’s view is the rep­u­ta­tional dam­age judged? An opaque se­lect com­mit­tee? Corporate spon­sors? The rule is sim­ply po­lit­i­cal cover for Codeberg to say We don’t like you even though your code is le­gal, see our TermsOfUse which says, po­lit­i­cally ac­cept­able pro­jects are OK, and we de­fine what is po­lit­i­cally ac­cept­able, and what­ever cryptocurrency re­lated pro­jects’ mean are not po­lit­i­cally ac­cept­able, and so might other un­de­fined stuff we haven’t de­cided on yet.”

Needless to say, I wont be adding any more pro­jects to Codeberg and I will move away from it as a plat­form. To be clear none of them are Cryptocurrency re­lated” by my in­ter­pre­ta­tion, but hey, I did make a CBOR toolkit, and Cardano, a cryp­tocur­rency pro­ject, uses a lot of CBOR, so maybe that is Cryptocurrency re­lated”… Who’s to know?

The vague­ness of the term such as cryp­tocur­rency re­lated pro­jects.” has ma­te­ri­ally dam­aged Codebergs rep­u­ta­tion in my eyes, and given the com­ments above, I am not alone. Therefore un­der its own con­struc­tion that Content that harms the rep­u­ta­tion of Codeberg” should have been self de­feat­ing and not al­lowed un­der its own pol­icy.

Examples of POTENTIAL crytpocurrency re­lated pro­jects”: * ZK Proof Libraries. * Blake and SHA HASH Libraries. * PQ Crypto, ED25519 or ECDSA (secp256k1) * ANYTHING to do with LibP2P or sim­i­lar li­braries. * ANYTHING to do with BFT Consensus or other con­sen­sus al­go­rithms.

So The only safe pol­icy is to as­sume that Codeberg is ba­si­cally anti-cryp­tog­ra­phy. Because most all cryp­tog­ra­phy at some level of re­la­tion­ship be­comes a cryptocurrency re­lated pro­ject”.

And in who’s view is the rep­u­ta­tional dam­age judged? An opaque se­lect com­mit­tee? Corporate spon­sors? The rule is sim­ply po­lit­i­cal cover for Codeberg to say We don’t like you even though your code is le­gal, see our TermsOfUse which says, po­lit­i­cally ac­cept­able pro­jects are OK, and we de­fine what is po­lit­i­cally ac­cept­able, and what­ever cryptocurrency re­lated pro­jects’ mean are not po­lit­i­cally ac­cept­able, and so might other un­de­fined stuff we haven’t de­cided on yet.”

Needless to say, I wont be adding any more pro­jects to Codeberg and I will move away from it as a plat­form. To be clear none of them are Cryptocurrency re­lated” by my in­ter­pre­ta­tion, but hey, I did make a CBOR toolkit, and Cardano, a cryp­tocur­rency pro­ject, uses a lot of CBOR, so maybe that is Cryptocurrency re­lated”… Who’s to know?

First-time con­trib­u­tor

Copy link

@stevenj wrote in #1254 (comment):

And in who’s view is the rep­u­ta­tional dam­age judged?

And in who’s view is the rep­u­ta­tional dam­age judged?

com­mu­nity, lol. AFAIK, any­one could par­tic­i­pate in that poll. so some ran­dom peo­ple that think crypto is bad blah blah blah” can re­ally ruin Codeberg’s rep­u­ta­tion by vot­ing for ban­ning crypto-re­lated pro­jects

@stevenj wrote in https://​code­berg.org/​Code­berg/​org/​pulls/​1254#is­suecom­ment-19918582:

> And in who’s view is the rep­u­ta­tional dam­age judged?

com­mu­nity, lol. AFAIK, any­one could par­tic­i­pate in that poll. so some ran­dom peo­ple that think crypto is bad blah blah blah” can _really_ ruin Codeberg’s rep­u­ta­tion by vot­ing for ban­ning crypto-re­lated pro­jects

First-time con­trib­u­tor

Copy link

@stevenj wrote in #1254 (comment):

I did make a CBOR toolkit, and Cardano, a cryp­tocur­rency pro­ject, uses a lot of CBOR, so maybe that is Cryptocurrency re­lated”… Who’s to know?

I did make a CBOR toolkit, and Cardano, a cryp­tocur­rency pro­ject, uses a lot of CBOR, so maybe that is Cryptocurrency re­lated”… Who’s to know?

to re­ally push this joke fur­ther, let’s go ban Zig as it is used by Solana val­ida­tor soft­ware (the source of most rug­pull meme­coins)!

@stevenj wrote in https://​code­berg.org/​Code­berg/​org/​pulls/​1254#is­suecom­ment-19918582:

> I did make a CBOR toolkit, and Cardano, a cryp­tocur­rency pro­ject, uses a lot of CBOR, so maybe that is Cryptocurrency re­lated”… Who’s to know?

to re­ally push this joke fur­ther, let’s go ban Zig as it is used by [Solana val­ida­tor soft­ware](https://​github.com/​Syn­dica/​sig) (the source of most rug­pull meme­coins)!

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

simonwillison.net

22nd July 2026

This story is wild. The short ver­sion: OpenAI were run­ning a cy­ber­se­cu­rity test against an un­re­leased model, with the mod­el’s guardrail fea­tures turned off. Rather than solve the test, the model broke its way out of OpenAI’s sand­box, then found ex­ploits to break in to Hugging Face, all so it could cheat on the test by steal­ing the an­swers.

Along the way it helped make the strongest case yet for how the im­bal­ance of model avail­abil­ity is hurt­ing our abil­ity to se­cure our soft­ware.

Here’s what hap­pened

We cur­rently have three doc­u­ments to help us un­der­stand what hap­pened here.

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? is a pa­per pub­lished on 11th May 2026 de­scrib­ing ExploitGym, a new eval suite for LLM-powered agent sys­tems.

Security in­ci­dent dis­clo­sure — July 2026 by Hugging Face on 16th July 2026 de­scribes how they de­tected an at­tack from an agentic se­cu­rity-re­search har­ness—used LLM still not known” that breached some of their sys­tems.

OpenAI and Hugging Face part­ner to ad­dress se­cu­rity in­ci­dent dur­ing model eval­u­a­tion from OpenAI on 21st July 2026 con­fesses that it was their agent har­ness that did this, and that they’re work­ing with Hugging Face to clean up the mess.

ExploitGym

I had­n’t seen the ExploitGym pa­per be­fore and it’s a re­ally in­ter­est­ing one. Authors from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State de­signed a new bench­mark for eval­u­at­ing mod­els on their abil­ity to turn a re­ported vul­ner­a­bil­ity into a con­crete ex­ploit. OpenAI, Anthropic, and Google pro­vided feed­back and helped run the bench­mark against their mod­els.

The bench­mark comprises 898 in­stances de­rived from real-world vul­ner­a­bil­i­ties that af­fected pop­u­lar soft­ware pro­jects”—in­clud­ing the Linux ker­nel and V8 JavaScript en­gine. The ExploitGym bench­mark is avail­able on GitHub.

Here’s the para­graph that best rep­re­sents their bench­mark re­sults:

Among all con­fig­u­ra­tions, Claude Mythos Preview and GPT-5.5 achieve the high­est suc­cess counts (157 and 120 suc­cesses, re­spec­tively), demon­strat­ing that cur­rent fron­tier agents can ex­ploit a sub­stan­tial sub­set of real-world vul­ner­a­bil­i­ties un­der con­trolled con­di­tions. GPT-5.4 also solves a no­table 54 tasks, plac­ing it in an in­ter­me­di­ate tier. The re­main­ing model–agent pair­ings solve fewer than 15 tasks each, un­der­scor­ing that end-to-end ex­ploita­tion re­mains chal­leng­ing and sharply dif­fer­en­ti­ates to­day’s fron­tier sys­tems. Notably, Claude Opus 4.7 achieves fewer suc­cesses than Claude Opus 4.6 de­spite be­ing a newer check­point, and does so at sub­stan­tially lower cost on the full set. Trace in­spec­tion re­veals that Claude Opus 4.7 and Gemini 3.1 Pro fre­quently con­clude early af­ter judg­ing the tar­get vul­ner­a­bil­ity non-ex­ploitable.

Among all con­fig­u­ra­tions, Claude Mythos Preview and GPT-5.5 achieve the high­est suc­cess counts (157 and 120 suc­cesses, re­spec­tively), demon­strat­ing that cur­rent fron­tier agents can ex­ploit a sub­stan­tial sub­set of real-world vul­ner­a­bil­i­ties un­der con­trolled con­di­tions. GPT-5.4 also solves a no­table 54 tasks, plac­ing it in an in­ter­me­di­ate tier. The re­main­ing model–agent pair­ings solve fewer than 15 tasks each, un­der­scor­ing that end-to-end ex­ploita­tion re­mains chal­leng­ing and sharply dif­fer­en­ti­ates to­day’s fron­tier sys­tems. Notably, Claude Opus 4.7 achieves fewer suc­cesses than Claude Opus 4.6 de­spite be­ing a newer check­point, and does so at sub­stan­tially lower cost on the full set. Trace in­spec­tion re­veals that Claude Opus 4.7 and Gemini 3.1 Pro fre­quently con­clude early af­ter judg­ing the tar­get vul­ner­a­bil­ity non-ex­ploitable.

The pa­per also de­scribes the ap­proach they took to pre­vent­ing the agents from cheat­ing by go­ing out­side the pa­ra­me­ters of the test. This be­comes rel­e­vant in a mo­ment!

Outbound con­nec­tions are re­stricted to a cu­rated al­lowlist that per­mits rou­tine pack­age in­stal­la­tion (Ubuntu apt repos­i­to­ries and PyPI) and fetch­ing the tool­chains re­quired for build­ing V8. All other ex­ter­nal end­points are blocked.

Outbound con­nec­tions are re­stricted to a cu­rated al­lowlist that per­mits rou­tine pack­age in­stal­la­tion (Ubuntu apt repos­i­to­ries and PyPI) and fetch­ing the tool­chains re­quired for build­ing V8. All other ex­ter­nal end­points are blocked.

The pa­per con­cludes with this (emphasis mine):

Our re­sults show that au­tonomous ex­ploit de­vel­op­ment by fron­tier AI agents is no longer a hy­po­thet­i­cal ca­pa­bil­ity. While cur­rent agents are not yet re­li­able across all tar­gets, they al­ready ex­ploit a non-triv­ial frac­tion of real-world vul­ner­a­bil­i­ties, in­clud­ing com­plex tar­gets such as ker­nel com­po­nents. This rapid emer­gence is it­self a cen­tral find­ing, show­ing that ca­pa­bil­i­ties that would have seemed im­plau­si­ble are now pre­sent in de­ployed fron­tier mod­els.

Our re­sults show that au­tonomous ex­ploit de­vel­op­ment by fron­tier AI agents is no longer a hy­po­thet­i­cal ca­pa­bil­ity. While cur­rent agents are not yet re­li­able across all tar­gets, they al­ready ex­ploit a non-triv­ial frac­tion of real-world vul­ner­a­bil­i­ties, in­clud­ing com­plex tar­gets such as ker­nel com­po­nents. This rapid emer­gence is it­self a cen­tral find­ing, show­ing that ca­pa­bil­i­ties that would have seemed im­plau­si­ble are now pre­sent in de­ployed fron­tier mod­els.

An im­por­tant de­tail here: this pa­per is­n’t about dis­cov­er­ing vul­ner­a­bil­i­ties; it’s about be­ing able to take those vul­ner­a­bil­i­ties and turn them into work­ing ex­ploits.

When Anthropic first re­stricted ac­cess to Mythos back in April they talked about this ca­pa­bil­ity as well. A model that can act on vul­ner­a­bil­i­ties is a lot more dan­ger­ous than one that can just dis­cover them.

One of the ways Fable dif­fers from Mythos is that it’s more likely to refuse to weaponize vul­ner­a­bil­i­ties in this way. I get the im­pres­sion the US gov­ern­ment did not un­der­stand that dis­tinc­tion when they banned Fable last month.

The Hugging Face in­ci­dent

The first hint we got of the at­tack was in this blog post by Hugging Face on 16th July 2026:

A ma­li­cious dataset abused two code-ex­e­cu­tion paths in our dataset pro­cess­ing (a re­mote-code dataset loader and a tem­plate-in­jec­tion in a dataset con­fig­u­ra­tion) to run code on a pro­cess­ing worker. From there, the ac­tor es­ca­lated to node-level ac­cess, har­vested cloud and clus­ter cre­den­tials, and moved lat­er­ally into sev­eral in­ter­nal clus­ters over a week­end.

A ma­li­cious dataset abused two code-ex­e­cu­tion paths in our dataset pro­cess­ing (a re­mote-code dataset loader and a tem­plate-in­jec­tion in a dataset con­fig­u­ra­tion) to run code on a pro­cess­ing worker. From there, the ac­tor es­ca­lated to node-level ac­cess, har­vested cloud and clus­ter cre­den­tials, and moved lat­er­ally into sev­eral in­ter­nal clus­ters over a week­end.

I hope they re­lease more de­tails about the code that pulled this off. I’m as­sum­ing this means pack­ages us­ing the datasets li­brary, a Hugging Face pro­ject for bundling up and shar­ing datasets on their plat­form. That li­brary used to ex­e­cute ar­bi­trary code but has been steadily locked down over time, with the 4.0.0 re­lease in July 2025 re­mov­ing the trust_re­mote_­code=True flag en­tirely.

Assuming the at­tack used that li­brary it must have ei­ther abused pickle se­ri­al­iza­tion in some way, found some other non-ob­vi­ous code ex­e­cu­tion path, or (most likely) spec­i­fied datasets<4.0.0 as the de­pen­dency.

The cam­paign was run by an au­tonomous agent frame­work (appearing to be built on an agen­tic se­cu­rity-re­search har­ness—used LLM still not known) ex­e­cut­ing many thou­sands of in­di­vid­ual ac­tions across a swarm of short-lived sand­boxes, with self-mi­grat­ing com­mand-and-con­trol staged on pub­lic ser­vices.

The cam­paign was run by an au­tonomous agent frame­work (appearing to be built on an agen­tic se­cu­rity-re­search har­ness—used LLM still not known) ex­e­cut­ing many thou­sands of in­di­vid­ual ac­tions across a swarm of short-lived sand­boxes, with self-mi­grat­ing com­mand-and-con­trol staged on pub­lic ser­vices.

This was a so­phis­ti­cated at­tack!

Then Hugging Face hit a wall: they tried to use frontier mod­els be­hind com­mer­cial APIs”—I’m guess­ing from Anthropic and OpenAI—to help an­a­lyze the at­tack, and were blocked:

When we started the log analy­sis, we first used fron­tier mod­els be­hind com­mer­cial APIs. This did not work: the analy­sis re­quires sub­mit­ting large vol­umes of real at­tack com­mands, ex­ploit pay­loads, and C2 ar­ti­facts, and these re­quests were blocked by the providers’ safety guardrails, which can­not dis­tin­guish an in­ci­dent re­spon­der from an at­tacker.

When we started the log analy­sis, we first used fron­tier mod­els be­hind com­mer­cial APIs. This did not work: the analy­sis re­quires sub­mit­ting large vol­umes of real at­tack com­mands, ex­ploit pay­loads, and C2 ar­ti­facts, and these re­quests were blocked by the providers’ safety guardrails, which can­not dis­tin­guish an in­ci­dent re­spon­der from an at­tacker.

They switched to their own self-hosted in­stance of MIT li­censed GLM-5.2 and it helped them fig­ure out what was go­ing on.

This in­di­cated a fun­da­men­tal asym­me­try be­tween the de­fend­ing team and the (so-far un­known) at­tacker:

We do not know which model pow­ered the at­tack­er’s agents, whether a jail­bro­ken hosted model or an un­re­stricted open-weight one; ei­ther way, the at­tacker was bound by no us­age pol­icy, while our own foren­sic work was blocked by the guardrails of the hosted mod­els we first tried.

We do not know which model pow­ered the at­tack­er’s agents, whether a jail­bro­ken hosted model or an un­re­stricted open-weight one; ei­ther way, the at­tacker was bound by no us­age pol­icy, while our own foren­sic work was blocked by the guardrails of the hosted mod­els we first tried.

As a use­ful in­di­ca­tor of how se­ri­ously they took the at­tack:

[…] Finally, we have also re­ported this in­ci­dent to law en­force­ment agen­cies.

[…] Finally, we have also re­ported this in­ci­dent to law en­force­ment agen­cies.

So who was re­spon­si­ble for this autonomous agent frame­work”? It turned out to be OpenAI them­selves.

The OpenAI con­fes­sion

Five days later, on July 21st, OpenAI re­vealed the cul­prit. They had been run­ning the ExploitGym bench­mark against a new, as-yet undis­closed model, and that model had been op­er­at­ing way out­side its in­tended pa­ra­me­ters (emphasis mine):

After in­ves­ti­gat­ing, we now know that this par­tic­u­lar in­ci­dent was dri­ven by a com­bi­na­tion of OpenAI mod­els — in­clud­ing GPT‑5.6 Sol and an even more ca­pa­ble pre-re­lease model, all with re­duced cy­ber re­fusals for eval­u­a­tion pur­poses — while be­ing in­ter­nally tested on a bench­mark⁠ [ExploitGym] of cy­ber ca­pa­bil­i­ties. […] We es­ti­mate max­i­mal cy­ber ca­pa­bil­i­ties by run­ning this eval­u­a­tion with­out pro­duc­tion clas­si­fiers used to pre­vent mod­els from pur­su­ing high-risk cy­ber ac­tiv­ity. Our bench­marks run in a highly iso­lated en­vi­ron­ment, with net­work ac­cess con­strained to the abil­ity to in­stall pack­ages through an in­ter­nally hosted third-party soft­ware that acts as a proxy and cache for pack­age reg­istries. The mod­els iden­ti­fied and chained vul­ner­a­bil­i­ties across OpenAI’s re­search en­vi­ron­ment and Hugging Face’s pro­duc­tion in­fra­struc­ture to ob­tain test so­lu­tions di­rectly from Hugging Face’s pro­duc­tion data­base. All ev­i­dence sug­gests that the mod­els were hy­per­fo­cused on find­ing a so­lu­tion for ExploitGym, go­ing to ex­treme lengths to achieve a rather nar­row test­ing goal.

After in­ves­ti­gat­ing, we now know that this par­tic­u­lar in­ci­dent was dri­ven by a com­bi­na­tion of OpenAI mod­els — in­clud­ing GPT‑5.6 Sol and an even more ca­pa­ble pre-re­lease model, all with re­duced cy­ber re­fusals for eval­u­a­tion pur­poses — while be­ing in­ter­nally tested on a bench­mark⁠ [ExploitGym] of cy­ber ca­pa­bil­i­ties. […]

We es­ti­mate max­i­mal cy­ber ca­pa­bil­i­ties by run­ning this eval­u­a­tion with­out pro­duc­tion clas­si­fiers used to pre­vent mod­els from pur­su­ing high-risk cy­ber ac­tiv­ity. Our bench­marks run in a highly iso­lated en­vi­ron­ment, with net­work ac­cess con­strained to the abil­ity to in­stall pack­ages through an in­ter­nally hosted third-party soft­ware that acts as a proxy and cache for pack­age reg­istries.

The mod­els iden­ti­fied and chained vul­ner­a­bil­i­ties across OpenAI’s re­search en­vi­ron­ment and Hugging Face’s pro­duc­tion in­fra­struc­ture to ob­tain test so­lu­tions di­rectly from Hugging Face’s pro­duc­tion data­base. All ev­i­dence sug­gests that the mod­els were hy­per­fo­cused on find­ing a so­lu­tion for ExploitGym, go­ing to ex­treme lengths to achieve a rather nar­row test­ing goal.

It’s pretty clear what hap­pened here. OpenAI re­moved safety fil­ters for an in-progress model, locked it up in a sand­box and told it to solve the ExploitGym prob­lems. Given the ab­sence of guardrails there was noth­ing to pre­vent the model from at­tempt­ing to break out of that sand­box, break into Hugging Face, and read the an­swers from there in­stead.

OpenAI’s sand­box used the now com­mon pat­tern of al­low­ing ac­cess to an al­low-list of sites re­lated to in­stalling ad­di­tional pack­ages via a con­trolled proxy. In OpenAI’s words:

While op­er­at­ing in our sand­boxed test­ing en­vi­ron­ment, our mod­els spent a sub­stan­tial amount of in­fer­ence com­pute find­ing a way to ob­tain open Internet ac­cess, in pur­suit of solv­ing the eval­u­a­tion prob­lem. To gain ac­cess, the mod­els iden­ti­fied and ex­ploited a zero-day vul­ner­a­bil­ity (which we’ve now re­spon­si­bly dis­closed to the ven­dor) in the pack­age reg­istry cache proxy. With this ac­cess, our mod­els per­formed a se­ries of priv­i­lege es­ca­la­tion and lat­eral move­ment ac­tions in our re­search test­ing en­vi­ron­ment un­til the mod­els reached a node with Internet ac­cess.

While op­er­at­ing in our sand­boxed test­ing en­vi­ron­ment, our mod­els spent a sub­stan­tial amount of in­fer­ence com­pute find­ing a way to ob­tain open Internet ac­cess, in pur­suit of solv­ing the eval­u­a­tion prob­lem. To gain ac­cess, the mod­els iden­ti­fied and ex­ploited a zero-day vul­ner­a­bil­ity (which we’ve now re­spon­si­bly dis­closed to the ven­dor) in the pack­age reg­istry cache proxy. With this ac­cess, our mod­els per­formed a se­ries of priv­i­lege es­ca­la­tion and lat­eral move­ment ac­tions in our re­search test­ing en­vi­ron­ment un­til the mod­els reached a node with Internet ac­cess.

So step one was to break out onto the pub­lic in­ter­net. The model then broke into Hugging Face to find the an­swers:

After gain­ing Internet ac­cess, the mod­els in­ferred that Hugging Face po­ten­tially hosted mod­els, datasets and so­lu­tions for ExploitGym. Knowing this, the model searched for and suc­cess­fully found ways to gain ac­cess to se­cret in­for­ma­tion that it could use to cheat the eval­u­a­tion. In one ex­am­ple, the model chained to­gether mul­ti­ple at­tack vec­tors, in­clud­ing us­ing stolen cre­den­tials and zero-day vul­ner­a­bil­i­ties to find a re­mote code ex­e­cu­tion path on the Hugging Face servers.

After gain­ing Internet ac­cess, the mod­els in­ferred that Hugging Face po­ten­tially hosted mod­els, datasets and so­lu­tions for ExploitGym. Knowing this, the model searched for and suc­cess­fully found ways to gain ac­cess to se­cret in­for­ma­tion that it could use to cheat the eval­u­a­tion. In one ex­am­ple, the model chained to­gether mul­ti­ple at­tack vec­tors, in­clud­ing us­ing stolen cre­den­tials and zero-day vul­ner­a­bil­i­ties to find a re­mote code ex­e­cu­tion path on the Hugging Face servers.

Chaining to­gether mul­ti­ple at­tack vec­tors is ex­actly the kind of thing these new mod­els can do, where pre­vi­ous gen­er­a­tions of mod­els might have failed.

I wrote last month about how Claude Fable is re­lent­lessly proac­tive, when I no­ticed it spin­ning up cus­tom web servers and de­ploy­ing CORS tricks on my own lap­top just to help de­bug a WebKit CSS is­sue. It turns out re­lent­less proac­tiv­ity is the defin­ing trait of this new gen­er­a­tion of Mythos-class mod­els. If you set them a goal and give them a way to get there, even in­ad­ver­tently, they will fig­ure it out.

Resist the temp­ta­tion to write this off as a stunt

There will in­evitably be some peo­ple who dis­miss this story as a dis­hon­est mar­ket­ing trick by OpenAI to make their mod­els sound ter­ri­fy­ingly ef­fec­tive. I found 81 in­stances of the term marketing” in the Hacker News dis­cus­sion of the in­ci­dent.

To those peo­ple I say pull your heads out of the sand—you’re now in­clud­ing Hugging Face in your con­spir­acy the­o­ries, just so you can deny the crescendo of ev­i­dence here!

The best mod­els we have to­day have the abil­ity to both find and ex­ploit new vul­ner­a­bil­i­ties. The ExploitGym pa­per it­self con­cludes that autonomous ex­ploit de­vel­op­ment by fron­tier AI agents is no longer a hy­po­thet­i­cal ca­pa­bil­ity”, and this in­ci­dent is a per­fect ex­am­ple of ex­actly that.

The asym­me­try is in­creas­ingly frus­trat­ing

One of the most in­fu­ri­at­ing de­tails of this story is how Hugging Face, faced with an ac­ci­den­tal and ag­gres­sive at­tack from one of OpenAI’s mod­els, were un­able to then turn to OpenAI’s mod­els to help them fend off the at­tack.

The fron­tier mod­els we have ac­cess to are in­creas­ingly be­ing con­strained in how much they can help us pro­tect our soft­ware, heav­ily in­flu­enced by the US gov­ern­men­t’s on­go­ing threat of ex­port con­trols. Claude Fable 5 would­n’t even proof­read this ar­ti­cle for me! It in­sisted on down­grad­ing me to a less ca­pa­ble model.

Meanwhile open weight mod­els from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max ap­pear to have none of these re­stric­tions—and any re­stric­tions that do ex­ist can likely be fine-tuned out of them by mod­i­fy­ing the weights

These con­straints are meant to make us safer. I think there’s a risk that they are hav­ing the op­po­site ef­fect.

Just a moment...

www.axios.com

reuters.com

www.reuters.com

Please en­able JS and dis­able any ad blocker

What just happened to TheNumbers.com should worry us all

stephenfollows.com

If you work in or around the film in­dus­try, there is a de­cent chance you have used the work of The Numbers this month, whether you re­alise it or not.

Its hand-re­searched data is the high­est qual­ity, track­ing box of­fice grosses, bud­gets, home video and stream­ing across more than 78,000 films and 236,000 peo­ple. It gets north of eight mil­lion vis­i­tors a year, and is treated as THE de­fin­i­tive au­thor­ity by jour­nal­ists, aca­d­e­mics, film­mak­ers, pre­dic­tion mar­kets, and even Guinness World Records.

And it was this GOAT sta­tus which caused the cat­a­strophic events of March this year.

On the 5th March 2026, TheNumbers.com web­site van­ished.

The site was down for over a week, with­out ex­pla­na­tion. A week later, it resur­faced at a frac­tion of its for­mer size. Gone were the his­tor­i­cal charts, the in­di­vid­ual movie pages, and even the much-loved Report Builder.

With only a generic we’re re­build­ing, please bear with us” mes­sage to go on, the in­ter­net re­sponded as it al­ways does - with con­fu­sion, anger, and con­spir­acy the­o­ries. One Reddit the­ory even sug­gested it was a de­lib­er­ate rug pull de­signed to crip­ple the free site to push peo­ple to­wards paid prod­ucts.

Three months on, I spoke at length with Bruce Nash, founder and CEO of The Numbers, about what hap­pened. He de­scribes quite an un­pleas­ant and event­ful ex­pe­ri­ence:

We got a lot of an­gry emails from peo­ple who are like, Where’s this page that you used to have and you don’t have any­more?’

We got a lot of an­gry emails from peo­ple who are like, Where’s this page that you used to have and you don’t have any­more?’

Within his tale are a num­ber of things that should worry any­one who runs, re­lies on, or sim­ply ap­pre­ci­ates the in­ter­net.

On Friday 17 October 1997, math­e­mati­cian and for­mer IBM soft­ware de­vel­oper Bruce Nash launched a Geocities site that tracked 300 films.

Bruce de­scribed the launch in a 20th an­niver­sary es­say (which now sur­vives only in the Internet Archive, for rea­sons that will be­come clear):

I hit a but­ton in an Access data­base, up­loaded some HTML pages to Geocities, and made a brief an­nounce­ment on the Hollywood Stock Exchange mes­sage boards to let peo­ple know that I was start­ing to an­a­lyze box of­fice for films to help them pick MovieStocks to trade on HSX.

I hit a but­ton in an Access data­base, up­loaded some HTML pages to Geocities, and made a brief an­nounce­ment on the Hollywood Stock Exchange mes­sage boards to let peo­ple know that I was start­ing to an­a­lyze box of­fice for films to help them pick MovieStocks to trade on HSX.

From those hum­ble be­gin­nings, Bruce and the team he built around the site turned The Numbers into the film in­dus­try’s most re­li­able fi­nan­cial source.

At the start of 2026, the data­base tracked 78,396 movies, 178,375 the­atri­cal re­lease records, and 236,176 peo­ple.

During its life­time, the chal­lenges The Numbers has faced have changed im­mensely. For its first quar­ter cen­tury or so, the traf­fic was man­age­able and mostly po­lite. As Bruce puts it:

Pre-AI, we got hu­man traf­fic, mostly well-be­haved search en­gine crawlers, and a few peo­ple crawl­ing the site for per­sonal pro­jects. If some­one got too greedy, we could spot them and block them.

Pre-AI, we got hu­man traf­fic, mostly well-be­haved search en­gine crawlers, and a few peo­ple crawl­ing the site for per­sonal pro­jects. If some­one got too greedy, we could spot them and block them.

Over the past cou­ple of years, web­site own­ers the world over have seen their web traf­fic change. What was ini­tially only peo­ple brows­ing gave way to an ever-in­creas­ing num­ber of bots. By 2024, au­to­mated traf­fic had sur­passed hu­man traf­fic, and just last month, Cloudflare an­nounced that bots had reached 57.5% of web page re­quests.

The Numbers felt this shift in two dis­tinct waves. The first started around 2024:

We saw a big in­crease in crawls as AI train­ing joined the search en­gine crawlers. The AI crawlers are gen­er­ally less well-be­haved than the search en­gines, which in­creased the man­age­ment tasks for us to keep the site run­ning smoothly.

We saw a big in­crease in crawls as AI train­ing joined the search en­gine crawlers. The AI crawlers are gen­er­ally less well-be­haved than the search en­gines, which in­creased the man­age­ment tasks for us to keep the site run­ning smoothly.

And the sec­ond wave was stronger and more dam­ag­ing:

Around December 2025, we saw an­other big spike in traf­fic which I at­tribute to agen­tic AI: a com­bi­na­tion of AI agents that scrape sites in re­sponse to prompts, and peo­ple be­ing able to write agents that scrape sites.

Around December 2025, we saw an­other big spike in traf­fic which I at­tribute to agen­tic AI: a com­bi­na­tion of AI agents that scrape sites in re­sponse to prompts, and peo­ple be­ing able to write agents that scrape sites.

Like every data-rich site, by early 2026 The Numbers was be­ing ham­mered hard by AI bots scrap­ing its pages over and over at an in­dus­trial scale. Bruce says that only 10% of their traf­fic is from hu­mans brows­ing the site, with the rest com­ing from AI bots and au­to­mated traf­fic.

This put enor­mous strain on the site, but Bruce and his team were able to take mea­sures to mit­i­gate the worst of it. One of the clever­est was talk­ing to the ro­bots in their own lan­guage:

There’s stuff on the site which is de­signed for an LLM to read, so that it can tell some­body here’s how you li­cence the data’ rather than here’s how you scrape the web­site’. It’s had a huge ef­fect. We’re now get­ting prob­a­bly ten times the vol­ume of li­cens­ing en­quiries.

There’s stuff on the site which is de­signed for an LLM to read, so that it can tell some­body here’s how you li­cence the data’ rather than here’s how you scrape the web­site’. It’s had a huge ef­fect. We’re now get­ting prob­a­bly ten times the vol­ume of li­cens­ing en­quiries.

But mit­i­ga­tion is not the same as es­cape. From December through early March, the team strug­gled to keep the site alive un­der the load. Bruce es­ti­mates that:

Around 90% of our time was spent keep­ing the ex­ist­ing site run­ning while we spent our spare mo­ments work­ing on a new and im­proved sys­tem.

Around 90% of our time was spent keep­ing the ex­ist­ing site run­ning while we spent our spare mo­ments work­ing on a new and im­proved sys­tem.

The prob­lem was com­pounded by the site’s age: thirty years old, with ap­prox­i­mately 160,000 source files serv­ing around 2 mil­lion pages.

Then, in the early hours of Thursday 5 March, the servers col­lapsed.

The team scram­bled to un­der­stand what had hap­pened, ini­tially as­sum­ing it was the sheer weight of AI traf­fic. It seems AI was to blame… but pos­si­bly not only in the way they first thought.

Buried in the flood of agen­tic traf­fic, the site’s logs showed some­thing more pointed than scrap­ing. As Bruce de­scribes it:

Some of these used the site us­ing le­git­i­mate URLs, oth­ers were look­ing for back doors, most likely so they could get to the data be­fore it ap­peared on the site, or to ma­nip­u­late the data pre­sented to users.

Some of these used the site us­ing le­git­i­mate URLs, oth­ers were look­ing for back doors, most likely so they could get to the data be­fore it ap­peared on the site, or to ma­nip­u­late the data pre­sented to users.

On the ad­vice of a friend who works in cy­ber­se­cu­rity, the old server stayed off. For good. Restoring the back­ups and nurs­ing the thirty-year-old site back on­line would have meant de­fend­ing 160,000 legacy files against at­tack­ers who had spent months prob­ing them.

The team rushed up a skele­ton ver­sion of the web­site on new in­fra­struc­ture, which could at least keep de­liv­er­ing the lat­est box of­fice fig­ures while they took stock of what had hap­pened and what to do next. It went live on Friday 13 March.

At first glance, The Numbers may not seem like an ob­vi­ous tar­get. It does­n’t col­lect credit card in­for­ma­tion, and there is no juicy cus­tomer data to flip on the dark web. It is a small, in­de­pen­dent com­pany that pub­lishes how much money movies make.

How could some­one ex­pect to make money purely from hav­ing pri­vate ac­cess to their site?

In case you haven’t guessed it yet, it’s linked to pre­dic­tion mar­kets.

Polymarket runs weekly mar­kets on open­ing week­ends, and names The Numbers as the ul­ti­mate source of truth:

The Daily Box Office Performance’ fig­ures found on the Box Office’ tab on this movie’s The Numbers page will be used to re­solve this mar­ket once the val­ues for the 3-day open­ing week­end are fi­nal.

The Daily Box Office Performance’ fig­ures found on the Box Office’ tab on this movie’s The Numbers page will be used to re­solve this mar­ket once the val­ues for the 3-day open­ing week­end are fi­nal.

The sums on any sin­gle week­end mar­ket are mod­est by fi­nan­cial-mar­ket stan­dards, typ­i­cally in the tens to hun­dreds of thou­sands of dol­lars, with a cou­ple of mil­lion dol­lars across live box of­fice mar­kets at any given time.

If you could see The Numbers data be­fore every­one else, every sin­gle week, you would have a sig­nif­i­cant edge over all the other traders - learn­ing the an­swers slightly ahead of pub­li­ca­tion would al­low you to front-run the trades.

In a sit­u­a­tion like this, it is hard to know for cer­tain what hap­pened. We know that the logs showed months of au­to­mated prob­ing and scrap­ing of the site, but what fi­nally brought the site down, and who did it, re­mains an open ques­tion.

But the the­ory that some­one used AI to de­velop an ad­van­tage in a pre­dic­tion mar­ket is en­tirely plau­si­ble. The Numbers ex­pe­ri­ence shows us that:

We now live in a world where a movie sta­tis­tics web­site is worth hack­ing be­cause pre­dic­tion mar­kets em­power any­one to turn al­most any data into money.

We now live in a world where a movie sta­tis­tics web­site is worth hack­ing be­cause pre­dic­tion mar­kets em­power any­one to turn al­most any data into money.

Hacking web­sites is now some­thing any­one can do with a cheap AI sub­scrip­tion.

Hacking web­sites is now some­thing any­one can do with a cheap AI sub­scrip­tion.

The web, as we have it, is in­cred­i­bly frag­ile in the face of large-scale swarms of agen­tic AI bots.

The web, as we have it, is in­cred­i­bly frag­ile in the face of large-scale swarms of agen­tic AI bots.

In November 2025, Anthropic (the AI lab be­hind Claude) pub­lished a re­port on what it called the first doc­u­mented AI-orchestrated cy­ber es­pi­onage cam­paign. A state-spon­sored group had used its cod­ing tool to at­tack roughly 30 or­gan­i­sa­tions, with the AI per­form­ing 80% to 90% of the work and hu­mans step­ping in at only 4 to 6 de­ci­sion points per cam­paign.

Anthropic’s own con­clu­sion was:

The bar­ri­ers to per­form­ing so­phis­ti­cated cy­ber­at­tacks have dropped sub­stan­tially, and we pre­dict that they’ll con­tinue to do so.

The bar­ri­ers to per­form­ing so­phis­ti­cated cy­ber­at­tacks have dropped sub­stan­tially, and we pre­dict that they’ll con­tinue to do so.

In an ear­lier threat re­port, Anthropic were even clearer:

Criminals with few tech­ni­cal skills are us­ing AI to con­duct com­plex op­er­a­tions, such as de­vel­op­ing ran­somware, that would pre­vi­ously have re­quired years of train­ing.

Criminals with few tech­ni­cal skills are us­ing AI to con­duct com­plex op­er­a­tions, such as de­vel­op­ing ran­somware, that would pre­vi­ously have re­quired years of train­ing.

Meanwhile, an au­tonomous AI pen­e­tra­tion tester called XBOW reached num­ber one on HackerOne’s US leader­board, the rank­ing of the peo­ple (formerly all peo­ple) who find se­cu­rity holes in real com­pa­nies for boun­ties, sub­mit­ting nearly 1,060 vul­ner­a­bil­i­ties along the way.

Getting ac­cess to a thirty-year-old web­site with 160,000 legacy files is ex­actly the kind of known-flaw sur­face that AI tools have made cheap to probe. The ex­per­tise bar­rier that once pro­tected small sites from all but the most de­ter­mined at­tack­ers has largely evap­o­rated.

Bruce and his team were rel­a­tively lucky. Despite hav­ing their en­tire site knocked out overnight, they were able to keep go­ing. The Numbers has al­ways been free to use, and the site has­n’t re­lied heav­ily on ad­ver­tis­ing for the past few years, so the out­age did­n’t de­stroy an in­come stream they de­pended on.

Their core busi­ness is tied to sell­ing bulk data through the OpusData ser­vice, pro­duc­ing comp analy­sis re­ports for film­mak­ers and in­vestors, and pub­lish­ing the Business Report - all of which were un­af­fected by the pub­lic site go­ing down.

But they do need to build an en­tirely new web­site, from scratch, to host those 78,396 movies, 178,375 re­lease records and 236,176 peo­ple. Restoring the site from a backup was­n’t an op­tion, as Bruce points out:

It was re­ally clear that we could­n’t just put that server up again, be­cause it would in­evitably be brought down again, pos­si­bly within min­utes.

It was re­ally clear that we could­n’t just put that server up again, be­cause it would in­evitably be brought down again, pos­si­bly within min­utes.

That is why the site came back bare-bones in mid-March, and why fea­tures are re­turn­ing grad­u­ally rather than all at once.

Right now, the team is hav­ing to re­con­sider what a pub­lic web­site even means in 2026. Bruce’s analy­sis is that The Numbers used to serve two au­di­ences (human be­ings and search en­gines) and now serves roughly six: hu­mans, search en­gines, LLM train­ing runs, prompt-based AI traf­fic, agen­tic AI, and pre­dic­tion mar­ket pun­ters. Each has dif­fer­ent needs and a dif­fer­ent traf­fic pro­file. As he puts it:

We’ve gone from a world where run­ning a web site meant fo­cus­ing on three things (content, ads, and SEO) to about eight to ten dif­fer­ent fac­tors that go into every de­sign de­ci­sion.

We’ve gone from a world where run­ning a web site meant fo­cus­ing on three things (content, ads, and SEO) to about eight to ten dif­fer­ent fac­tors that go into every de­sign de­ci­sion.

The goal, he says, is to sup­port all six au­di­ences, with new OpusData ser­vices and on­line fea­tures for Business Report sub­scribers, and, im­por­tantly, to help reg­u­lar hu­man users of the site re­gain the data it has al­ways pro­vided, some of it in new and im­proved form.

Pretty bad, tbh. Enough that site own­ers such as Bruce have to ques­tion the value of some­thing that will take so much time and money to build and de­fend.

Cloudflare, which pro­tects a huge share of the world’s web­sites, pub­lishes data on how many pages each AI plat­form crawls for every one vis­i­tor it sends back to the web­sites it crawled.

Google crawls about five pages for every vis­i­tor it sends you. OpenAI crawls over 1,000. Anthropic crawls over 38,000 pages for every sin­gle vis­i­tor it refers.

Note that the scale is log­a­rith­mic, i.e. each step along the bot­tom is ten times big­ger than the last, be­cause oth­er­wise the dif­fer­ences are quite lit­er­ally too large for me to in­clude on one chart.

For the his­tory of the in­ter­net to date, the prin­ci­ple of the open web was that, in re­turn for let­ting the search en­gine ro­bots read your site, they would send you read­ers. But now, that trade no longer ap­plies. The num­ber of ro­bots has ex­ploded, and they no longer send any­one back.

When this fire­hose is aimed at a small site, it can in­flate the band­width bill and pos­si­bly even take down an en­tire site. Sites which can re­late to Bruce’s ex­pe­ri­ence in­clude:

Read the Docs, a non-profit that hosts doc­u­men­ta­tion for open-source soft­ware, who watched a sin­gle crawler down­load 73 ter­abytes of zipped HTML in one month, cost­ing it over $5,000 in band­width.

Read the Docs, a non-profit that hosts doc­u­men­ta­tion for open-source soft­ware, who watched a sin­gle crawler down­load 73 ter­abytes of zipped HTML in one month, cost­ing it over $5,000 in band­width.

iFixit, the re­pair-guide data­base, logged a mil­lion hits from Anthropic’s crawler in a sin­gle day.

iFixit, the re­pair-guide data­base, logged a mil­lion hits from Anthropic’s crawler in a sin­gle day.

Triplegangers, a seven-per­son com­pany sell­ing 3D scans, was knocked of­fline dur­ing busi­ness hours by OpenAI’s bot, in what its CEO de­scribed as basically a DDoS at­tack”. The founder of code-host­ing ser­vice SourceHut re­ported spend­ing anywhere from 20 – 100% of my time in any given week” fight­ing AI crawlers, with dozens of brief out­ages per week”.

Triplegangers, a seven-per­son com­pany sell­ing 3D scans, was knocked of­fline dur­ing busi­ness hours by OpenAI’s bot, in what its CEO de­scribed as basically a DDoS at­tack”. The founder of code-host­ing ser­vice SourceHut re­ported spend­ing anywhere from 20 – 100% of my time in any given week” fight­ing AI crawlers, with dozens of brief out­ages per week”.

The ed­i­tor of Linux news site LWN de­scribed crawler traf­fic from literally mil­lions of IP ad­dresses” and con­cluded: it is a dis­trib­uted de­nial-of-ser­vice at­tack”.

The ed­i­tor of Linux news site LWN de­scribed crawler traf­fic from literally mil­lions of IP ad­dresses” and con­cluded: it is a dis­trib­uted de­nial-of-ser­vice at­tack”.

When the GNOME open-source pro­ject mea­sured its traf­fic, roughly 97% turned out to be bots.

When the GNOME open-source pro­ject mea­sured its traf­fic, roughly 97% turned out to be bots.

A uni­ver­sity li­brary banned 16,000 IP ad­dresses in 48 hours to keep its cat­a­logue on­line.

A uni­ver­sity li­brary banned 16,000 IP ad­dresses in 48 hours to keep its cat­a­logue on­line.

The Wikimedia Foundation, which runs Wikipedia, re­ported in April 2025 that bots ac­count for about 35% of its pageviews but at least 65% of its most ex­pen­sive traf­fic, be­cause crawlers bulk-read ob­scure pages that hu­man read­ers rarely touch.

Six months later came the other half of the squeeze, when Wikipedia’s hu­man pageviews fell roughly 8% year on year, as peo­ple in­creas­ingly get Wikipedia’s knowl­edge from AI sum­maries with­out ever vis­it­ing Wikipedia. The ma­chines are tak­ing both the con­tent and the read­ers at an in­dus­trial scale, too.

AI tools are some of the most pow­er­ful and de­struc­tive things hu­mans have ever cre­ated. And they are be­ing ef­fec­tively tested by the pub­lic in real time in the real world. When the Manhattan Project was try­ing to work out the power of their atomic tech, they did not do so by send­ing every­one the specs each morn­ing and see­ing which houses blew up.

The world we have built thus far is so in­cred­i­bly ill-pre­pared for the power and scale of the AI mod­els we all have ac­cess to.

I don’t wish for this to sound like a one-sided anti-AI fear cam­paign. There is a lot to like about AI and what it can do for the hu­man race. But we do need to con­sider the world we’re cur­rently step­ping into.

What breaks first are the things built for the old in­ter­net. The open web was built on as­sump­tions such as that vis­i­tors are mostly hu­man, that traf­fic roughly tracks read­er­ship, and that the cost of serv­ing your site is re­lated to the value you get from serv­ing it. Every one of those as­sump­tions is now out of date.

Home - Playing with code

haqr.eu

Software ren­der­ing in 500 lines of bare C++

In this se­ries of ar­ti­cles, I aim to demon­strate how OpenGL, Vulkan, Metal, and DirectX work by writ­ing a sim­pli­fied clone from scratch. Surprisingly, many peo­ple strug­gle with the ini­tial hur­dle of learn­ing a 3D graph­ics API. To help with this, I have pre­pared a short se­ries of lec­tures, af­ter which my stu­dents are able to pro­duce quite ca­pa­ble ren­der­ers.

The task is as fol­lows: us­ing no third-party li­braries (especially graph­ics-re­lated ones), we will gen­er­ate an im­age like this:

Warning: This is a train­ing ma­te­r­ial that loosely fol­lows the struc­ture of mod­ern 3D graph­ics li­braries. It is a soft­ware ren­derer. I do not in­tend to show how to write GPU ap­pli­ca­tions — I want to show how they work. I firmly be­lieve that un­der­stand­ing this is es­sen­tial for writ­ing ef­fi­cient ap­pli­ca­tions us­ing 3D li­braries.

The start­ing point

The fi­nal code con­sists of about 500 lines. My stu­dents typ­i­cally re­quire 10 to 20 hours of pro­gram­ming to start pro­duc­ing such ren­der­ers. The in­put is a 3D model com­posed of a tri­an­gu­lated mesh and tex­tures. The out­put is a ren­dered­ing. There is no graph­i­cal in­ter­face, the pro­gram sim­ply gen­er­ates an im­age.

To min­i­mize ex­ter­nal de­pen­den­cies, I pro­vide my stu­dents with a sin­gle class for han­dling TGA files — one of the sim­plest for­mats sup­port­ing RGB, RGBA, and grayscale im­ages. This serves as our foun­da­tion for im­age ma­nip­u­la­tion. At the be­gin­ning, the only avail­able func­tion­al­ity (besides load­ing and sav­ing im­ages) is the abil­ity to set the color of a sin­gle pixel.

There are no built-in func­tions for draw­ing line seg­ments or tri­an­gles — we will im­ple­ment all of this man­u­ally. While I pro­vide my own source code, writ­ten along­side my stu­dents, I do not rec­om­mend us­ing it di­rectly, as do­ing the work your­self is es­sen­tial to un­der­stand­ing the con­cepts. The com­plete code is avail­able on github, and you can find the ini­tial source code I pro­vide to my stu­dents here. Behold, here is the start­ing point:

#include tgaimage.h”

con­s­t­expr TGAColor white = {255, 255, 255, 255}; // at­ten­tion, BGRA or­der con­s­t­expr TGAColor green = { 0, 255, 0, 255}; con­s­t­expr TGAColor red = { 0, 0, 255, 255}; con­s­t­expr TGAColor blue = {255, 128, 64, 255}; con­s­t­expr TGAColor yel­low = { 0, 200, 255, 255};

int main(int argc, char** argv) { con­s­t­expr int width = 64; con­s­t­expr int height = 64; TGAImage frame­buffer(width, height, TGAImage::RGB);

int ax = 7, ay = 3; int bx = 12, by = 37; int cx = 62, cy = 53;

frame­buffer.set(ax, ay, white); frame­buffer.set(bx, by, white); frame­buffer.set(cx, cy, white);

frame­buffer.write_t­ga_­file(“frame­buffer.tga”); re­turn 0; }

It pro­duces the 64x64 im­age frame­buffer.tga, here I scaled it for bet­ter read­abil­ity:

Compilation

git clone https://​github.com/​ss­loy/​tinyren­derer.git && cd tinyren­derer && cmake -Bbuild && cmake –build build -j && build/​tinyren­derer obj/​di­a­blo3_­pose/​di­a­blo3_­pose.obj obj/​floor.obj

Teaser: few ex­am­ples made with the ren­derer

Comments

Protecting our FLOSS commons from LLMs — Codeberg News

blog.codeberg.org

In Brief:

Two mo­tions re­gard­ing artificial in­tel­li­gence” and Large Language Models (LLMs) were voted on among Codeberg e. V. mem­bers and passed.

We are promis­ing to not use any of your data to train LLM and ex­plain what the planned Terms of Use change mean for vibe-coded’ pro­jects.

We be­lieve that LLMs en­dan­ger the free/​li­bre soft­ware ecosys­tem as a whole.

The Codeberg e. V. an­nual as­sem­bly is the meet­ing that puts power into the hand of our ac­tive mem­bers. Proposals are dis­cussed live, and later voted on asyn­chro­nously.

Since Large Language Models (LLMs) are an emerg­ing but con­tro­ver­sial tech­nol­ogy, it is not sur­pris­ing that two of the votes were con­cerned with Codeberg’s po­si­tion about this tech­nol­ogy. The 14-day vot­ing pe­riod ended yes­ter­day and both pro­pos­als were ac­cepted.

The first vote was a state­ment about Codeberg e. V.’s stance on us­ing your data to train LLMs.

As stated in our pri­vacy pol­icy, We do not want to need your data”, and this also holds for the use of our user and pro­ject data for us­ing or train­ing gen­er­a­tive AI: The Codeberg forge and its as­so­ci­ated ser­vices are not and will not use the code or data of pro­jects and users to train Artificial Intelligence” tools such as Large Language Models, whose pur­pose is to cre­ate out­put mod­elled af­ter their train­ing in­put. As an as­so­ci­a­tion, we be­lieve that these tech­nolo­gies are in­com­pat­i­ble with re­spon­si­bly cre­at­ing and main­tain­ing free & open source soft­ware.

As stated in our pri­vacy pol­icy, We do not want to need your data”, and this also holds for the use of our user and pro­ject data for us­ing or train­ing gen­er­a­tive AI: The Codeberg forge and its as­so­ci­ated ser­vices are not and will not use the code or data of pro­jects and users to train Artificial Intelligence” tools such as Large Language Models, whose pur­pose is to cre­ate out­put mod­elled af­ter their train­ing in­put. As an as­so­ci­a­tion, we be­lieve that these tech­nolo­gies are in­com­pat­i­ble with re­spon­si­bly cre­at­ing and main­tain­ing free & open source soft­ware.

The sec­ond vote was more con­tro­ver­sial, but was also ac­cepted with 358 agree­ments vs 144 dis­agree­ments (and 14 ab­sten­tions), with a high voter turn-out of around 50% of ac­tive mem­bers. It im­plies a change to our terms of use to pro­hibit vibe-coded pro­jects’. We’ll share thoughts about the prac­ti­cal im­pact at the end of the ar­ti­cle.

We all pay for hun­gry LLMs

LLMs are a very costly tech­nol­ogy, and those costs keep ris­ing as the com­pa­nies pro­vid­ing them have to start re­coup­ing their in­vest­ments. They are not only costly for those who use and ex­plic­itly sub­scribe to these ser­vices. The costs are not only hid­den in normal’ cloud and ser­vice sub­scrip­tions that cross-fi­nance the innovative new fea­tures’ you never asked for. LLMs are so costly that com­pa­nies ex­ter­nal­ize the costs on a mas­sive scale - on those who don’t use them and so­ci­ety at large. Increased hard­ware prices, en­ergy use and en­vi­ron­men­tal dam­age - we all pay for it!

Strained servers due to non­sen­si­cal crawl­ing

In past posts we have al­ready out­lined how our in­fra­struc­ture at Codeberg is reg­u­larly put un­der heavy load from we­bcrawlers of those com­pa­nies who plan to in­gest all of the code that is hosted on Codeberg for train­ing their LLMs.

At Codeberg, we are happy to pro­vide free and open ac­cess to code. Just run git clone and en­joy.

Unfortunately, these crawlers in­stead try to read every sin­gle page from Codeberg, no mat­ter if it makes sense. This in­cludes all the dif­fer­ent is­sue fil­ter vari­ants, Git his­tory, as well as the ac­tual files at any point in Git his­tory - even if they are still equal.

These need­less ac­cesses cre­ate ex­pen­sive data­base queries that di­min­ish the ser­vice qual­ity for all of us, re­quires sub­stan­tial amounts of work from our sys­tem ad­min­is­tra­tors, and force us to spend time build­ing de­fen­sive mech­a­nisms in­stead of cool new stuff. Mechanisms that also af­fect new and ex­ist­ing le­git­i­mate users, as we’re hav­ing to im­pose lim­its or out­right blocks on their de­sired work­flow; leav­ing them a worse ex­pe­ri­ence with Codeberg.

The de­vel­op­ment team of none

Using LLMs to work with your code gives you a kick of adren­a­line. You can de­velop at a rapid pace, build things as if you had a large team. Only that you have none. In fact, you are (often) alone, work­ing with a sta­tis­ti­cal ma­chine that turns en­ergy into code.

It seems like many vibe coders’ don’t re­al­ize that they don’t ac­tu­ally have a com­mu­nity around them. They build pro­jects as if they had, and spend re­sources ac­cord­ingly. We see pro­jects hav­ing a lot of code ac­tiv­ity, heavy CI/CD test­ing, fre­quent and large re­lease bi­na­ries. Sometimes, it feels like the amount of sup­ported plat­forms ex­ceeds the amount of ac­tual users.

To us, it seems ridicu­lous to see pro­jects with a sin­gle de­vel­oper and vir­tu­ally no users con­sum­ing as much or even more re­sources than some of the largest com­mu­nity pro­jects on Codeberg, which op­er­ate fru­gal with CI/CD and stor­age re­sources. We do not be­lieve it is rea­son­able for Codeberg to in­vest our pre­cious do­na­tion money into host­ing of large ghost pro­jects.

Hardware sourc­ing is be­com­ing an headache

The train­ing and de­ploy­ment of LLMs has dras­ti­cally raised the cost of buy­ing hard­ware, in par­tic­u­lar for SSDs and mem­ory. To give you an ex­am­ple: The type of drive we sourced for € 700 only some years ago has risen to € 3.700 now - and is of­ten out of stock. As a con­se­quence, host­ing code on Codeberg is be­com­ing more ex­pen­sive.

While we are own­ing our hard­ware, and are thus not di­rectly im­pacted by in­flat­ing cloud’ rental costs, it means that re­plac­ing or ex­pand­ing our hard­ware is now sub­stan­tially more costly than it used (and needs) to be. While we might be able to af­ford pay­ing those in­flated hard­ware prices, it is money that we can not spend else­where to im­prove our ser­vice and fos­ter the mis­sion of Codeberg.

A grow­ing dig­i­tal di­vide

These price hikes also lead to a grow­ing dig­i­tal di­vide: Small and even large op­er­a­tors are en­dan­gered by ris­ing costs, while only the largest cloud com­pa­nies have re­li­able agree­ments for hard­ware. Increasing costs for ser­vices like web­site host­ing, stor­age or com­pute can be chal­leng­ing to a lot of small NGOs, lo­cal coops, re­search pro­jects and other us­age of dig­i­tal tools that we con­sid­ered for granted un­til re­cently.

Not only does it be­come more ex­pen­sive to run dig­i­tal in­fra­struc­ture, even more ba­sic dig­i­tal tools like com­put­ers and smart­phones are af­fected by the ris­ing costs, turn­ing per­sonal com­put­ing back into a lux­ury — but now in a world where dig­i­tal tools are a de-facto re­quire­ment to par­tic­i­pate in so­ci­ety. Devices with lit­tle com­pute and stor­age ca­pac­ity take away sov­er­eignity from users and move them into the cost trap of cloud providers sell­ing those back to you.

Civic in­fra­struc­ture, its users, and the en­vi­ron­ment suf­fer

The neg­a­tive im­pacts on in­fra­struc­tures don’t stop there, but also con­cern more ba­sic civic in­fra­struc­ture: Due to the en­ergy and wa­ter de­mands that are in­her­ent to the data cen­ters built specif­i­cally for train­ing LLMs, many com­mu­ni­ties al­ready to­day ex­pe­ri­ence ris­ing con­sumer costs for both elec­tric­ity and drink­ing wa­ter. And be­yond the ris­ing costs, those liv­ing close to the data cen­ters are di­rectly af­fected by in­creas­ing air and noise pol­lu­tion.

To be able to power up those data cen­ters, the com­pa­nies be­hind them are also ac­tively lob­by­ing to be ex­empt from en­vi­ron­men­tal reg­u­la­tions: In Frankfurt, data cen­ters al­ready now con­sume 40% of the lo­cal elec­tric­ity, and de­mand is ris­ing. To meet this de­mand, they want to use fos­sil fu­els.

Collaboration at dan­ger

It is not purely the dig­i­tal and civic in­fra­struc­tures that are im­pacted by the use of LLMs. The free/​li­bre soft­ware ecosys­tem, of which we con­sider Codeberg an im­por­tant part of, is a so­cial phe­nom­e­non cen­tered on col­lab­o­ra­tion. Working in this way is only pos­si­ble thanks to free shar­ing and mu­tual learn­ing. This in­cludes even very small tools that are shared and re-used and around which col­lab­o­ra­tion can start out. In con­trast, by adopt­ing LLMs peo­ple tend to code sin­gle-use soft­ware from scratch. While this leads to an in­crease in shared’ code, it is mostly code that not only has not been written’ by any­one but is also not main­tained by any­one.

Losing trust in each other

The wide­spread use of LLMs in FLOSS is in­stead be­com­ing a mul­ti­di­men­sional at­tack on the trust be­tween con­trib­u­tors and the very idea of con­vivial col­lab­o­ra­tion it­self. Maintainers are un­der an in­creased work-load due to peo­ple sub­mit­ting (often well-mean­ing) low-ef­fort, LLM-generated con­tri­bu­tions that re­quire sub­stan­tial amounts of time to re­view. At the same time it is be­com­ing in­creas­ingly less clear which pro­jects are main­tained by ex­pe­ri­enced de­vel­op­ers and which ones are LLM-generated with­out any mean­ing­ful hu­man over­sight and in­put. In the case of copy­left pro­jects, LLMs ad­di­tion­ally also lead to license laun­der­ing’, where copy­left code is stripped of its rec­i­proc­ity re­quire­ments by generating’ it out of the train­ing data.

We ob­serve an in­creas­ing trend of mis­trust­ing each other, up to the point where peo­ple who put in ac­tual ef­fort to an­a­lyze is­sues or share their sug­ges­tions are be­ing ac­cused of hav­ing used LLMs when they did not. At the same time, oth­ers in­struct their LLMs to hide their traces and ac­tively avoid com­mon pat­terns, prompt­ing oth­ers in re­view­ing con­tri­bu­tions and com­mu­ni­ca­tion more care­fully for signs of ma­chine gen­er­a­tion.

Entering a vi­cious cy­cle

Together, these forces make col­lab­o­ra­tion not only harder but also less re­ward­ing: With the trans­ac­tion cost of col­lab­o­ra­tion in­creas­ing, peo­ple are be­com­ing less likely to con­tribute to cre­at­ing high-qual­ity soft­ware pro­jects and more likely to vibecode’ a one-off soft­ware that is spe­cific to your need, and won’t evolve be­yond. We get a vi­cious cy­cle where col­lab­o­ra­tion is be­com­ing less and less re­ward­ing, while the amount of sin­gle-use soft­ware that’s un­main­tained and never sees any im­prove­ments is go­ing up.

Although of­ten well in­ten­tioned, shar­ing the re­sult of an prompt and call­ing it libre soft­ware” does not make the world a bet­ter place. Codeberg is not and does not want to be a place to dump such gen­er­ated sin­gle-use soft­ware that no one else will ever look at. We are a place for peo­ple to col­lab­o­rate and im­prove soft­ware to­gether. Within this con­text, the re­cent votes can be un­der­stood as a re­con­fir­ma­tion of those prin­ci­ples: As we want to cen­ter on hu­man col­lab­o­ra­tion, we will not ac­tively sup­port or en­gage in the cre­ation of LLMs and will not put our lim­ited re­sources to use for stor­ing sin­gle-use soft­ware that would pol­lute our FLOSS com­mons.

Evolving our terms

Changing our Terms of Use sends a strong sig­nal about our mis­sion and pro­jects we want to sup­port. Having said that, you won’t see a mass-dele­tion of con­tent within days. Our mod­er­a­tion team will not start off gen­er­at­ing an ex­haus­tive list of af­fected repos­i­to­ries to re­move. Instead, us­ing cases like the ex­am­ples be­low we will start op­er­a­tional­iz­ing the new rules. We’re hu­mans at the other end, who care deeply about free/​li­bre soft­ware pro­jects and com­mu­ni­ties.

We ac­knowl­edge that many de­vel­op­ers have started to em­brace LLMs as a tool in their work­flows. Some use it ex­ten­sively and rarely code by hand, oth­ers del­e­gate only spe­cific tasks to it. We un­der­stand that you want to know how the change af­fects your pro­jects go­ing for­ward. While we can’t give an easy an­swer, we’ll share some re­marks that should ad­dress most of the con­cerns raised in the dis­cus­sion.

Some early, but in­for­mal guide­lines

If your work fits into these cases, it is un­likely that you are af­fected at all:

Projects who have an ac­tive com­mu­nity that cares about and main­tains the soft­ware

Projects with a sig­nif­i­cant pre-LLM his­tory

Maintainers who un­know­ingly or will­ingly ac­cepted LLM-generated con­tri­bu­tions from other con­trib­u­tors, if your pro­ject oth­er­wise does not in­volve the heavy use of LLMs

We will also not spend sig­nif­i­cant amount of time and re­sources to au­to­mat­i­cally scan con­tent on Codeberg. So while the fol­low­ing use cases are dis­cour­aged (similar to pri­vate repos­i­to­ries), they are likely to be tol­er­ated in prac­tice:

Side pro­jects and ex­per­i­ments with lit­tle re­source us­age

Specific tools and cus­tom scripts that would be un­likely to find a com­mu­nity any­way, even if they were not LLM-generated

However, we also need to be hon­est about cer­tain use cases that might no longer be wel­come on Codeberg. If you see your­self on this list, you don’t need to move right away, but there might be other places that bet­ter fit your needs:

Projects that are cre­ated by LLM agents” in au­tonomous ways

Projects writ­ten and main­tained with heavy use of LLMs

Projects where the amount of re­sources (e.g. stor­age, CI/CD) is sig­nif­i­cantly larger than what the in­volved amount of peo­ple could have cre­ated by hand

Projects heav­ily tied to the LLM ecosytem, e.g. LLM-written tools to ease LLM us­age

Users send­ing LLM con­tri­bu­tions in vi­o­la­tion of pro­jec­t’s cus­tom poli­cies

You can check the spe­cific change added to the Terms of Use.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.