10 interesting stories served every morning and every evening.

Oxiis Intelligent Bike Booster|Bike Booster|ASUS Global

www.asus.com

ASUS Oxiis E250G1

Ride Easy. Explore More.

Oxiis is a uni­ver­sal, fric­tion-drive mo­tor sys­tem that trans­forms any con­ven­tional bi­cy­cle into a smart e-bike. Its name fuses the Greek Oxis (agility) and Axis (pivot)—perfectly re­flect­ing how its high-per­for­mance drive core de­liv­ers ul­ti­mate ac­cel­er­a­tion and in­stant, nim­ble re­spon­sive­ness right to your ride.

Focus on es­sen­tials

Adaptive boost tech­nol­ogy

Precisely de­tects in­clines, pro­vid­ing seam­less as­sis­tance for ef­fort­less climbs.

Peak power 500W

Instant burst, ef­fort­lessly con­quer­ing chal­leng­ing ter­rains, mak­ing every ride eas­ier.

Wireless ca­dence sen­sor

Ditch the wires. Simple to in­stall, smart to ride.

Smart brake-de­tect­ing tail­light

Enhances night vis­i­bil­ity and safety, safe­guard­ing your jour­ney.

Design

Simple de­sign. Impressive ben­e­fits.

The ASUS Oxiis E250G1 de­signed with the user in mind: min­i­mum struc­ture, max­i­mum util­ity. Every de­tail is metic­u­lously crafted to el­e­vate your daily ex­pe­ri­ence, de­liv­er­ing seam­less in­ter­ac­tion and a func­tional mas­ter­piece of pre­mium qual­ity and beauty.

Premium alu­minum con­struc­tion

Built to last with high-grade, rugged ma­te­ri­als.

Anti-slip tech­nol­ogy

Dynamically pres­sure au­to­mat­i­cally grips the tire for ef­fi­cient, slip-free power trans­fer.

Easy in­stal­la­tion

Quick trans­for­ma­tion. Zero mod­i­fi­ca­tions to gears or brakes.

Efficient heat dis­si­pa­tion

Reduced risk of over­heat­ing.

Battery

Removable mod­u­lar bat­tery

Easy to charge and swap, of­fer­ing ul­ti­mate free­dom for your jour­ney.

100W USB-C® PD

2 hours fast charge

158 Wh

bat­tery ca­pac­ity

Flight-safe for carry on lug­gage (Requires air­line ap­proval pri­or­ity to fly­ing)

Compatibility

Universal de­sign, per­fect fit

ASUS Oxiis E250G1 works with a wide range of bike frame types, in­clud­ing city, road, gravel, and fold­ing bikes, as well as hy­brid or hard­tail moun­tain bikes.

Accommodates

Find your fit

To guar­an­tee op­ti­mal per­for­mance, seam­less in­stal­la­tion, and max­i­mum safety, please dou­ble-check that your bike’s spe­cific tire width, wheel size, and seat post mea­sure­ments fully meet our cri­te­ria be­fore your first setup.

Tire width

Supports tire widths up to 60 mm.

Tire size

Compatible with tire sizes from 16 to 29 inches, plus 700C.

Seat post

Fits 25.4 – 34.9 mm posts (spacers in­cluded).

Installation Guide

Easy in­stal­la­tion

No com­plex tools are needed to in­stall ASUS Oxiis E250G1.

App

ASUS Oxiis app

Three modes. One seam­less ex­pe­ri­ence.

Switch modes ef­fort­lessly us­ing the in­te­grated but­ton or ded­i­cated app.

Eco mode: Efficient sup­port for flat ter­rain and tail­winds.

Normal mode: Smooth, bal­anced power for every­day ad­ven­tures.

Sport mode: Instant boost for steep climbs and high-speed per­for­mance.

Spec

Product Specifications

Motor Power

250W (Rated) / 500W (Peak)

Range

50 km (31 miles) in eco mode

Battery

158.4 Wh / 36 V

Weight

3.7 kg (including bat­tery)

Dimensions

400 x 84 x 128 mm (L x W x H)

Max Speed

32 km/​h (Limited to 25 km/​h in spe­cific re­gions)

Charging Time

2 hours (100W PD Charger re­quired)

Waterproof Rating

IPX4

App Compatibility

Android & iOS

Get lo­cal ASUS Zenbook 14 avail­abil­ity alerts - and more from ASUS/ROG

Client Challenge

support.mozilla.org

A re­quired part of this site could­n’t load. This may be due to a browser ex­ten­sion, net­work is­sues, or browser set­tings. Please check your con­nec­tion, dis­able any ad block­ers, or try us­ing a dif­fer­ent browser.

System Prompts

platform.claude.com

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Loading

Sorry...

scholar.google.com

We’re sorry…

… but your com­puter or net­work may be send­ing au­to­mated queries. To pro­tect our users, we can’t process your re­quest right now.

A Third World Embedded Engineer Responds to "RISC-V: They Should Have Known Better"

rvembedded.com

Dmitry Grinberg pub­lished a long piece ex­plain­ing his dis­taste for RISC-V, you can read his ar­ti­cle here: RISC-V: They Should Have Known Better - Dmitry.GR. It went to the front page of Hacker News and it started a good ar­gu­ment on Lobsters. It is the most sub­stan­tial crit­i­cism the ar­chi­tec­ture has had in a while and though I switched my en­tire stack away from STM32 and ARM to RISC-V and did a video on it about a year ago Good­bye STM32 ARM — Meet the CH32 RISC-V Chips That Replaced It! , part of me is in­fu­ri­ated be­cause so much of what he said seems like a bi­ased per­spec­tive.

Look, I am not go­ing to de­fend the ISA com­mit­tee, RISC-V in­ter­na­tional de­nied me mem­ber­ship to their golden tower. On the ar­chi­tec­ture it­self, the com­pressed store off­sets re­ally are strange, Zicsr re­ally should not be a sep­a­rate thing you have to re­mem­ber to ask for, I have hit every one of these and I have writ­ten a book thats about 80% com­plete about hit­ting them on the CH32V003, which is one of the very RV32E type chip he men­tions.

Maybe I should say where I am writ­ing from, be­cause it changes which parts of this ar­gu­ment look im­por­tant from my per­spec­tive.

I work out of Trinidad and Tobago, a small is­land na­tion off the coast of Venezuela. When I want a de­vel­op­ment board I am not click­ing through to next day de­liv­ery, I am check­ing whether the seller ships here at all, what cus­toms will do to it (if I get it at all), and what the to­tal lands at in TT dol­lars. Free Shipping” from Digikey, Mouser or any US or European man­u­factuer dosen’t ap­ply to me. I pay any­where from US $60 to US $200 to ship one dol­lar chips that peo­ple every­where else get free ship­ping on. In fact a well known PCB com­pany who reached out to me con­sid­er­ing spon­sor­ship turned me down so­ley based on ship­ping to my lo­ca­tion. Have a look here:

The stu­dents I want to teach are in the same po­si­tion, and so are the ones in Nigeria and Bangladesh and every­where else the peo­ple in the in­dus­try does not think about when it writes its blog posts. From that po­si­tion, the dif­fer­ence be­tween a ten cent part and a one dol­lar part is not a round­ing er­ror and it is not a de­tail you get to wave past on the way to the in­ter­est­ing dis­cus­sion about en­cod­ings. It is the dif­fer­ence be­tween a class of thirty stu­dents each hav­ing their own chip and a class of thirty stu­dents watch­ing one demo board if any at all. Instruction set el­e­gance is a thing you can af­ford to care about once the hard­ware is al­ready on your desk. Whether the hard­ware can get to your desk at all comes first. That is why the para­graph most peo­ple scrolled past is, to me, the most im­por­tant one in the ar­ti­cle.

Grinberg missed that part that RISC-V cre­ates a space for the other 99% out­side of the world” (which in this space world” is mainly the US and Europe) and it has noth­ing to do with ar­chi­tec­ture.

He Derives the Requirements and Lands on RV32EC

Before the in­ter­rupt arith­metic, be­fore the en­cod­ing com­plaints, he does some­thing care­ful. He asks what a cheap mi­cro­con­troller core is ac­tu­ally for. His an­swer is that it sits in­side a larger chip prod­ding reg­is­ters and con­fig­ur­ing hard­ware blocks, in an MP3 player, an SD card, a USB stick”. The real work is done by cus­tom sil­i­con around it. From that he de­rives what such a core needs. Low in­ter­rupt la­tency a small die area and good code den­sity, be­cause the code lives in ROM or SRAM and both are ex­pen­sive per byte. No hard­ware di­vider, pos­si­bly not even a mul­ti­plier, since you are not do­ing much arith­metic. No priv­i­lege sep­a­ra­tion, be­cause noth­ing un­trusted ever runs there.

Then he writes the line him­self:

But,” you might say, you just de­scribed RV32IC (or RV32EC)!”

But,” you might say, you just de­scribed RV32IC (or RV32EC)!”

And ear­lier, plainly:

I am 100% sure that RISC-V will own the cheap-as-dirt sin­gle-use mi­cro­con­troller space even­tu­ally.

I am 100% sure that RISC-V will own the cheap-as-dirt sin­gle-use mi­cro­con­troller space even­tu­ally.

So the most cred­i­ble RISC-V critic of the month sat down, worked out from first prin­ci­ples what a cheap mi­cro­con­troller core should be, ar­rived at the in­struc­tion set a ten cent chip im­ple­ments, and stated that this seg­ment is go­ing to be RISC-V’s.

He de­rives the case for the chip and then spends the rest of the ar­ti­cle an­noyed that the chip ex­ists.

This is al­most satir­i­cal.

His quar­rel is with whether that out­come was earned. That is a real ques­tion and I un­der­stand why it both­ers him. It is not, how­ever, a ques­tion that af­fects any­body de­cid­ing what to learn on, be­cause the chip is on the shelf ei­ther way.

Where I Actually Disagree, Strongly.

His cen­tral claim is the first one in the ar­ti­cle, and it is big­ger than any of the en­cod­ing com­plaints:

Simply put, the things a high-end CPU needs are di­a­met­ri­cally op­posed to the things a small cost-sav­ing mi­cro­con­troller core needs.

Simply put, the things a high-end CPU needs are di­a­met­ri­cally op­posed to the things a small cost-sav­ing mi­cro­con­troller core needs.

The con­clu­sion he draws is that no sin­gle ISA can serve both ends, and that RISC-V fans are fool­ing them­selves, in the­ory the premise is true. The con­clu­sion does not fol­low, and I can show you why from three parts sit­ting on my desk as we speak.

CH32V003. This is the cheap RV32EC with six­teen reg­is­ters, no mul­ti­plier, no di­vider, ma­chine mode only, 2KB of SRAM, 16KB of flash, ten cents, it’s EXACTLY the core he spec­i­fied. I shipped two prod­ucts with these, one is a bin mon­i­tor that has a ToF sen­sor, an LED and an air tag. The other is an agri­cul­tural prod­uct for a client that opens and closes a door at a cer­tain time. It also makes a good throw away part, as I show case in my whis­tle switch Clap Switch Is Dead. Here’s the RISC-V Powered Whistle Switch! and which in my view is the BEST part to re­place the over­priced, out­dated Arduino Did Arduino Q Ruin Arduino? - Here’s how to Switch to RISC-V with the CH32V003.

CH32H417. A dual core MCU that is un­matched in per­for­mance to price point and is at the higher end of the MCU line of things. It has a QingKe V5F at 400 MHz along­side a V3F at 144 MHz, 896KB of SRAM, 960KB of flash. USB 3.2 Gen1 with an in­te­grated 5 Gbps trans­ceiver, 100M Ethernet MAC and PHY, a SerDes iso­lated trans­ceiver, a 500 MB/s high speed in­ter­face, SDMMC, a cam­era in­ter­face, a dis­play con­troller, a graph­ics ac­cel­er­a­tor etc etc. I got a web browser run­ning on this thing I Built a Web Browser on a RISC-V Microcontroller (No Linux) Quantum en­tropy based GAN cat gen­er­a­tion Schrödinger’s De/Motivational Quantum Cat: GAN Image Generation on CH32 RISC-V Microcontroller and real-time fa­cial recog­ni­tion Real Time Facial Recognition on The Edge With CH32H417 RISC-V MCU in un­der 150KB of ram. I got a host of other pro­jects run­ning but those are just some I got time to record and put up.

Baochip. A VexRISC-V with an MMU built around a stack thats open from sil­i­con to os Baochip-1x: A Mostly-Open, 22nm SoC for High Assurance Applications « bun­nie’s blog, that runs Xous be­trusted-io/​xous-core: The Xous mi­cro­ker­nel de­signed by leg­endary hard­ware hacker bunnie” Huang , a Rust mi­cro­ker­nel with real process iso­la­tion. Privilege sep­a­ra­tion, the ex­act thing he says the cheap end does not need and there­fore does not get. In ad­di­tion to Xous it also sup­ports op­er­at­ing sys­tems like SEL4 vk2seb/bao1x-seL4: seL4 port to baochip-1x and Linux pkoscik/​baochip-linux: An at­tempt to boot main­line Linux on a stock Dabao board. I wrote the bare metal C SDK for the chip Arm­strong­Subero/​dabao-sdk: Bare metal C SDK for the Baochip-1x RISC-V SoC and it was of course the chip in­side the badge of DEFCON 34 The New Defcon Badges Pack a Unique Open Source Chip That Doubles as a Security Key | WIRED this year.

I can also point to the NES em­u­la­tor I wrote for the $1 ESP32C3 RISC-V based chip NES Emulator on $1 ESP32-C3 RISC-V Microcontroller, or ex­per­i­ment­ing with Linux on the Orange Pi RV2 OrangePi RV2 5 Minute Unboxing and Setup | RISC-V Ubuntu Linux that takes 5 min­utes to setup and has been run­ning since the day I boot it up.

Point is I could go on and on about how di­verse and ac­ces­si­ble cur­rently ship­ping RISC-V parts are, but then we’ll be stray­ing too much from the topic at hand.

I linked all those to say this, that all these parts all have the same base in­struc­tion set and I gained ex­per­tise in all in un­der a year and un­der US $100 across the en­tire stack, from dis­posi­ble sil­i­con to PC level, of course mi­nus data cen­ter com­pute.

For un­der US $100 in­clud­ing ship­ping I was able to ex­plore an en­tire ver­ti­cal stack us­ing one ar­chi­tec­ture. Due to the AI race the OrangePi RV2 has now gone up in price but at re­lease it cost $30 and shipped free. For about 7 dol­lars I got 50 CH32V003s with a de­bug­ger, the CH32H417 board is $20 on ana­log lamb and uses the same cheap (and of­fi­cial) de­bug­ger for the CH32V003 and the Baochip Dabao board (which I wrote a book about by the way check it out here (The Dabao Book - Payhip) was $9.50 on crowd sup­ply when I bought it, two with ship­ping from crowd sup­ply cost me $35, un­der $100 in to­tal. A de­bug­ger for an ARM part alone a Segger J-Link costs about $600, though I guess for that $100, and add an­other $100 to ship,so about $200 I could get an EDU edi­tion J-link and no chips or boards. Yaay.

Back to RISC-V, across all these parts, the base set is the same. So that means the same reg­is­ter model, same call­ing con­ven­tion, same tool­chain. Yes the ex­ten­sions dif­fer, but the thing is what I learned writ­ing as­sem­bly on the ten cent CH32V003 part did not stop be­ing true on any of the oth­ers. A dual core MCU, an SBC run­ning Linux or an ad­vanced cus­tom se­cu­rity chip run­ning a novel op­er­at­ing sys­tem. My skills were trans­ferrable to the point that in each case within a few hours I had tool­chains setup, could fo­cus on my ap­pli­ca­tions and when de­bug­ging I felt at home. All I need to work with them is the ISA man­ual and a C com­piler.

Now price the same jour­ney on the other side, for­get x86 – 64 and that du­op­oly, patent mine­field, with multi-thou­sand dol­lar de­bug probes; we’ll take a look at ARM.

The equiv­a­lent to the CH32V003 is the Cortex-M0 is ARMv6-M so some­thing like an STM32F030, step it up we have a Cortex-M7 which is ARMv7-M, to get an MMU in a part for Linux or SEL4 and Xous, you’re look­ing at an ap­pli­ca­tion proces­sor like the ARMv8-A.  These are dif­fer­ent Arm pro­files with sig­nif­i­cantly dif­fer­ent priv­i­lege, ex­cep­tion, and sys­tem mod­els, so mov­ing up the stack in­volves sub­stan­tially more re­learn­ing than sim­ply en­abling an­other RISC-V ex­ten­sion. Trust me I’ve used them all.

And at the top of that range the gap is not even about learn­ing curves. There is no Cortex-M mi­cro­con­troller with an in­te­grated USB 3.0 SuperSpeed PHY. The near­est dual core Arm part is an STM32H747, which is a fine chip and does not have one. If you need USB 3.0 you leave the mi­cro­con­troller class en­tirely: an i.MX 8 or an RK3xxx, which means Cortex-A. You want an MMU, Linux, DDR, a PMIC, and a board you are not lay­ing out in two lay­ers. Or you keep the M7 and add an ex­ter­nal bridge chip.

The H417 eval­u­a­tion board is around twenty dol­lars. The H747 in TFBGA240 car­ries a twenty week man­u­fac­turer lead time, chip only, costs about the same, be­fore you have any­thing to plug in, and Mouser asks for ID be­fore you can or­der, Digikey has also been known to deny peo­ple parts de­pend­ing on where they are and their name as Hussein Ali, well known Youtuber from NorthridgeFix de­scribes Star­link Repair - Digi-key re­fused my or­der.. Oh and it’s about US $60 – 100+ to ship to my lo­ca­tion. I can pick up H417s on the of­fi­cial WCH store on Aliexpress with free ship­ping and no ver­i­fi­ca­tion hul­la­balu. We haven’t even started talk­ing about the Cortex-A parts that have MMUs or thier de­bug­ging tools and ecosys­tem frag­men­ta­tion.

The Boundary Is Not Technical

Here is the part that un­der­cuts his fram­ing most di­rectly, and it has noth­ing to do with en­cod­ings. He treats the gap be­tween a small core and a large one as an ar­chi­tec­tural fact, some­thing that falls out of op­posed re­quire­ments. On ARM chips it is not an ar­chi­tec­tural fact. It is a PRODUCT bound­ary, and it is en­forced by li­cens­ing. Has any­one tried adding an MMU to a Cortex-M? The phys­i­cal trade­offs are real, the dif­fer­ence is that with RISC-V, the ISA owner does not de­cide for you where that bound­ary must be drawn. If you want vir­tual mem­ory on ARM you li­cense a Cortex-A in­stead, which is a dif­fer­ent core fam­ily, a dif­fer­ent pro­file, a dif­fer­ent ne­go­ti­a­tion, and a dif­fer­ent roy­alty. There is no in­cre­men­tal path. there is a wall, with a sales team on the other side of it.

Compare what hap­pened with Baochip. The RISC-V priv­i­leged spec­i­fi­ca­tion de­fines su­per­vi­sor mode and Sv32 pag­ing as op­tional things an im­ple­men­ta­tion may pro­vide. VexRISC-V is an open core, some­body added an MMU to it. bun­nie built a chip around it and runs a mi­cro­ker­nel with real process iso­la­tion on it that me in Trinidad a coun­try who’s name does not even come up in ISA cir­cles can ex­per­i­ment with at low cost and teach to other peo­ple in the re­gion.

That’s what free­dom looks like.

Nobody asked per­mis­sion, no­body signed any­thing, no­body pays a roy­alty per unit shipped and any­body can learn down to the RTL the sil­i­con is built on.  So when Grinberg in his ar­ti­cle lists priv­i­lege sep­a­ra­tion among the things the cheap end does not need and there­fore does not get, it is de­scrib­ing a prop­erty of ARMs prod­uct seg­men­ta­tion and at­tribut­ing it to in­struc­tion set de­sign. On RISC-V it is a check­box in the priv­i­leged spec, you leave it off in a ten cent part be­cause it costs area you do not want to spend, and you turn it on when you do, and the in­struc­tion set un­der­neath is the same ei­ther way.

That is the real dif­fer­ence be­tween the two ecosys­tems, and it is why one ISA can­not serve both ends” reads dif­fer­ently de­pend­ing on which side you are stand­ing on. On one side the ends are sep­a­rated by physics and cost, on the other they are sep­a­rated by physics, cost, and a con­tract.

The Thing He Calls Fragmentation

Before I close I want to ad­dress his stance on frag­men­ta­tion. He is not wrong that the ex­ten­sion mech­a­nism frag­ments the stan­dard. Zcb split­ting off from C is an­noy­ing and Zicsr not be­ing im­plied by the base is an­noy­ing. Vendors adding pro­pri­etary in­ter­rupt hard­ware does frag­ment things fur­ther, I learned first hand port­ing NuttX to the CH32V307 Porting Apache NuttX RTOS to the WCH CH32V307: A Deep Dive into the PFIC and Everything That Went Wrong.

But that mech­a­nism is the an­swer to his own open­ing ques­tion. The rea­son one in­struc­tion set can sit in a ten cent part with six­teen reg­is­ters and also in a chip run­ning a pro­tected multi-process op­er­at­ing sys­tem is pre­cisely that the small part is not car­ry­ing the large part’s bag­gage. There is no com­pro­mise core in the mid­dle serv­ing both badly, which is what diametrically op­posed re­quire­ments” would nor­mally force. Fragmentation and scal­a­bil­ity are the same prop­erty, you do not get one with­out the other and whether the trade­off was worth it is a fair ar­gu­ment and I do not think it has an ob­vi­ous an­swer.

What I do think is that he is right about the im­por­tant part, and right in a way that favours the thing he is crit­i­cis­ing. RISC-V is not go­ing to take the cheap mi­cro­con­troller space be­cause its en­cod­ing is el­e­gant. It is go­ing to take it be­cause the part costs ten cents, and be­cause the lad­der above it is the same in­struc­tion set all the way up. It is go­ing there be­cause an em­bed­ded en­gi­neer in a 3rd world coun­try can shine a cheap LED and see the tran­sis­tors in the sil­i­con, In­fra-Red, In Situ (IRIS) Inspection of Silicon « bun­nie’s blog and get 50 chips with a de­bug­ger and free de­vel­op­ment tools for the price of a cup of cof­fee and shipped free. It also means that world class en­gi­neers can de­sign MMUs onto chips that the gate keep­ers will never give a li­cense for.

He writes that this will hap­pen not due to its ISA de­sign, but de­spite it,” and he means it as a mild in­dict­ment. Read it from here and it is not one. Winning on price and avail­abil­ity is not a lesser way to win. It de­cides who is in the room. An ar­chi­tec­ture that ar­rives in my coun­try at ten cents a part, with an open tool­chain and no li­cense to ne­go­ti­ate, puts em­bed­ded sys­tems within reach of peo­ple who were pre­vi­ously go­ing to watch some­body else’s demo board and con­sume thier prod­ucts with­out ever be­ing able to match what they have ac­cess to. That’s the power of free­dom, open­ness and is democ­racy in it’s truest sense.

That is a bet­ter rea­son than el­e­gance. and I want to tell Mr Grinberg, that the word priv­iledge he tosses around in his ar­ti­cle also ex­tends be­yond the ISA de­pend­ing on where you are in the world.

Nuff said.

Armstrong Subero is an em­bed­ded sys­tems en­gi­neer and pub­lished au­thor with Apress/Springer. He builds the Rovari RISC-V ed­u­ca­tion plat­form from Trinidad and Tobago.

Software Engineering fundamentals matter more than ever

rhonabwy.com

The man­i­fes­ta­tion of my im­poster syn­drome, for me and to­day, is what does it mean to be a soft­ware en­gi­neer. There’s a lot more noise than sig­nal on the Internet about agen­tic en­gi­neer­ing, what can be ac­com­plished, and its im­pli­ca­tions for the fu­ture. The ti­tle I chose rather gives it away; it’s about choos­ing — care­fully — all the things you need to choose when you’re solv­ing the puz­zles of soft­ware and sys­tems de­vel­op­ment.

Beyond the hype and junkie-like mar­ket­ing fer­vor of major model providers”, I found a re­ally in­ter­est­ing power tool with the com­bi­na­tion of har­ness and mod­els. I’ve been fol­low­ing how friends have been us­ing these tools, and learn­ing a ton. As usual, the folks do­ing some of the most amaz­ing things aren’t the ones crow­ing about it, or post­ing nar­ra­tive blurbs in so­cial me­dia about the end of this pro­fes­sion. They found a big damn stick”, they’re ex­plor­ing the ful­crum points, and they’re rep­re­sent­ing good ole Archimedes to lean into that lever, mov­ing the world.

In the past year, agent har­nesses crossed the can it be done” ru­bi­con. (yep, jump­ing for­ward to Roman ref­er­ences). I would not have wished for the world’s knowl­edge to taken with­out per­mis­sion and re­gard, or the lu­natics to delve into eco­nomic self-deal­ing that’s peanut but­ter­ing over the oth­er­wise tank­ing US econ­omy. The eco­nomic mod­els for the large mod­els aren’t vi­able from any re­port that I’ve seen, but the ca­pa­bil­ity is­n’t go­ing away. Instead it’s shrink­ing (fast!). Open weight mod­els are mak­ing (beefy) per­sonal com­put­ers quite ca­pa­ble of do­ing the same. They’re not quite as ef­fec­tive, but the delta in time and ca­pa­bil­ity is­n’t large.

Can it be done” is only the start, not even close to the ma­jor­ity a soft­ware or sys­tem en­gi­neer’s pro­fes­sion. It’s like when I learned to weld in my 20’s — I quickly cre­ated things that I could­n’t lift or even get out the door of the shop. (thank good­ness for acety­lene torches). What I learned then is I think the same les­son, dif­fer­ent medium: How some­thing goes to­gether is what makes all the dif­fer­ence.

If you use agen­tic har­nesses to de­velop with a bit of fore­sight, you can get not only it works”, but also it’s testable” (I heav­ily lean into the prompt develop with red/​green TDD). But it’s not very solid much above that. The seams — how your code works, it’s API, and how it fits with other soft­ware — are as much art as sci­ence. It is made up of sub­jec­tive mea­sures that rely on your view­point (and ex­pe­ri­ence, as well as your guesses) for both what you’re solv­ing now, and how to live with that soft­ware over a long pe­riod of time.

Making soft­ware de­bug­gable, main­tain­able, lay­ered, and com­pos­able — that’s still quite a trick. Quite a lot of that work re­quires ex­ten­sive, thought­ful rea­son­ing. And that’s where the LLMs to­day, even the lead­ing edge of the capability” from fron­tier mod­els, fall short.

It helps to know that LLMs don’t reason”. They pre­dict, and the mod­els them­selves are ef­fec­tively writ­ten hu­man knowl­edge com­pressed. So if it’s in hu­man knowl­edge that was en­coded, it can echo out the hu­man rea­son­ing. For agents fo­cused on soft­ware de­vel­op­ment, those rea­son­ing traces are the pre­cious data for the mod­els. There’s a very ap­proach­able re­search pa­per on just how bad LLMS are at rea­son­ing called The Illusion of Thinking. There is some re­search I’m fol­low­ing that in­cludes pre­dic­tion of re­sults of ac­tions, but that’s not what we have to­day with cod­ing agents. It’s a pretty dif­fer­ent — and fas­ci­nat­ing — area of re­search. If you want to ex­plore, go dig­ging on how JEPA mod­els” work, LeWorld Model, and re­cent talks by Yann LeCun.

While you’re work­ing with LLMs though, there’s still a ton of ways to make them more ef­fec­tive. I think there’s a lot of ad­vances that we haven’t even re­ally be­gun to eek out. Most of the wins I’m see­ing to­day in­volve pro­vid­ing it good, con­cise data to work from, at the right time, and pro­vid­ing de­ter­min­is­tic val­i­da­tion tool­ing with nat­ural lan­guage feed­back that the LLM can use to cor­rect it­self. The amaz­ing thing to me is­n’t that it can pre­dict what to write, but that it is ef­fec­tive at tool call­ing and fol­low­ing in­struc­tions.

Another down­side of this in­struc­tion fol­low­ing is what Simon Willison coined as the lethal tri­fecta. Basically — LLM mod­els can’t dis­tin­guish be­tween good ad­vice and bad. They’re foun­da­tion­ally in­ca­pable of al­ways and con­sis­tently pre­vent­ing prompt in­jec­tion at­tacks. Alignment work”, safety har­nesses, and sand­boxes all help to add bar­ri­ers against the worst, but there are fun­da­men­tal gaps. And frankly, some­thing that tire­lessly fol­lows in­struc­tions with­out hav­ing good rea­son­ing is night­mare fuel to me.

Another down­side of this in­struc­tion fol­low­ing is what Simon Willison coined as the lethal tri­fecta. Basically — LLM mod­els can’t dis­tin­guish be­tween good ad­vice and bad. They’re foun­da­tion­ally in­ca­pable of al­ways and con­sis­tently pre­vent­ing prompt in­jec­tion at­tacks. Alignment work”, safety har­nesses, and sand­boxes all help to add bar­ri­ers against the worst, but there are fun­da­men­tal gaps. And frankly, some­thing that tire­lessly fol­lows in­struc­tions with­out hav­ing good rea­son­ing is night­mare fuel to me.

I hope there will be near-term nad­vances in how mod­els are trained to in­clude the equiv­a­lent of rea­son­ing traces for post-train­ing (RLHF). In my ideal fu­ture, these in­clude more of what it means to build soft­ware with clean in­ter­faces, that’s de­bug­gable, and and that’s main­tain­able as a key part of the re­in­forced eval­u­a­tions. Carefully re­view­ing, plan­ning, and fix­ing the seams of soft­ware (and sys­tems) is one of the crit­i­cal skills we both can, and need to, em­ploy when de­vel­op­ing soft­ware — with or with­out agen­tic as­sis­tants. And as I see the wave of Oh, that’s easy to im­ple­ment…” and peo­ple reach­ing for clankers to get it done, I think it’s more im­por­tant than ever.

It’s a great time to be fol­low­ing folks who write, talk, and share about the craft of soft­ware, and how we can be bet­ter ar­ti­sans. Hopefully it’s ob­vi­ous, but there’s never a sin­gle an­swer — a panacea. It’s al­ways about trade­offs, choos­ing what makes sense for the prob­lem at hand. With the help of a lot of great minds shar­ing their thoughts — both now and go­ing back decades — we have a great tool chest for this work. It’s about pick­ing, or re­work­ing to move to a bet­ter choice, the right ab­strac­tions. It’s core is man­ag­ing the cog­ni­tive load, learn­ing which pieces we need to be sta­ble, and where we want our work to flex and bend (and how).

And yes, I wrote the damn em-dashes my­self. I’m too in love with a re­cur­sive par­en­thet­i­cal in my writ­ing, and I like a break from com­mas and paren­the­ses.

And yes, I wrote the damn em-dashes my­self. I’m too in love with a re­cur­sive par­en­thet­i­cal in my writ­ing, and I like a break from com­mas and paren­the­ses.

Asynchronous I/O in DuckDB: Work, Thread, Work

duckdb.org

Pedro Holanda

2026 – 07-31

· 21 min

TL;DR: Starting with v2.0, sched­uled for fall 2026, DuckDB will sup­port asyn­chro­nous reads of Parquet and CSV files. This can sig­nif­i­cantly speed up queries when syn­chro­nous I/O does not sat­u­rate the avail­able band­width, as is typ­i­cal in EC2/S3 com­pute-stor­age se­tups.

It does­n’t mat­ter how fast query op­er­a­tors are in a data­base sys­tem if we can’t pull in the data quickly. For most of DuckDB’s his­tory, how­ever, this prob­lem was largely avoided by prun­ing data early. By push­ing down fil­ters and pro­jec­tions, we could en­sure that we only read what we ac­tu­ally needed.

This worked par­tic­u­larly well be­cause DuckDB pri­mar­ily ran lo­cally, with its main use case be­ing as a quick-draw data­base en­gine for query­ing data di­rectly from your ma­chine’s SSD. We could split the data into sev­eral par­ti­tions, such as row groups for Parquet files or fixed-size buffers for CSV files, and load them with low la­tency and high band­width. As a re­sult, the main bot­tle­necks were else­where: sub­queries, joins, ag­gre­ga­tions, and so on. The ac­tual data ac­cess path re­ceived less at­ten­tion be­cause syn­chro­nous ac­cess was per­fectly suit­able for this use case.

As usual, things changed. We re­al­ized that DuckDB’s ar­chi­tec­ture was a great fit for query­ing re­motely stored large-scale datasets, such as data lakes (e.g., DuckLake). Since May this year, we can even run DuckDB as a server us­ing the Quack pro­to­col. The orig­i­nal ex­pec­ta­tion of data files sit­ting on a lo­cal SSD there­fore no longer al­ways holds.

The prac­ti­cal im­pli­ca­tion of these changes is that many cur­rent DuckDB se­tups need to trans­fer files from re­mote stor­age to the ma­chine that will ac­tu­ally process them. For data lakes, for ex­am­ple, a typ­i­cal setup is to store the data in blob stor­age, such as S3, and process it on an EC2 ma­chine in the same re­gion. In this setup, la­tency and band­width play a much more sig­nif­i­cant role. If we can­not is­sue enough con­cur­rent re­quests to use the avail­able net­work band­width, per­for­mance can suf­fer dras­ti­cally, with threads spend­ing a large amount of their time wait­ing for re­mote reads in­stead of pro­cess­ing data.

As an ex­am­ple, let’s con­sider a sim­ple query over a re­mote Parquet file. For sim­plic­ity, let’s as­sume we only have a sin­gle thread ex­e­cut­ing.

FROM read­_­par­quet(‘s3://​bucket/​file.par­quet’);

A Parquet scan is par­ti­tioned into row-group-based jobs, with each job con­tain­ing one or more fetch tasks that is­sue byte-range re­quests. With syn­chro­nous I/O, the worker thread will be blocked, wait­ing for the data to ar­rive at the ma­chine be­fore per­form­ing ac­tual work, such as de­cod­ing, ag­gre­gat­ing, and so on. You can see a vi­sual de­pic­tion in the fig­ure be­low, where the thread is blocked from do­ing any work while it waits for the read to fin­ish.

Synchronous read

To ad­dress this, we have been im­ple­ment­ing asyn­chro­nous I/O pipelines in DuckDB. They are cur­rently im­ple­mented for Parquet and for un­com­pressed, seek­able UTF-8 CSV files, with sup­port for other for­mats, such as DuckDB’s na­tive for­mat and JSON, still to come. In the re­main­der of this blog post, we will give a sim­ple ex­pla­na­tion of how asyn­chro­nous I/O is im­ple­mented in DuckDB and pro­vide bench­marks for both Parquet and CSV files.

If you would like to try asyn­chro­nous I/O now, you can do so by us­ing DuckDB’s v2.0.0-dev pre­view builds. Asynchronous I/O will be used by de­fault from the next ma­jor DuckDB ver­sion, v2.0, re­leased in the fall.

If you would like to try asyn­chro­nous I/O now, you can do so by us­ing DuckDB’s v2.0.0-dev pre­view builds. Asynchronous I/O will be used by de­fault from the next ma­jor DuckDB ver­sion, v2.0, re­leased in the fall.

Asynchronous I/O

The con­cep­tual idea of asyn­chro­nous I/O is rather sim­ple: we should be able to start an I/O op­er­a­tion with­out block­ing the worker thread that re­quested it. Applied to our Parquet ex­am­ple, the same pic­ture would look like the fol­low­ing:

Asynchronous read

In this ex­am­ple, we have two ASYNC threads and one reg­u­lar worker thread. The ASYNC threads keep fetch tasks in flight while the worker thread de­codes data. During the ini­tial warm-up, the scan task parks, leav­ing the worker thread free to run other pipeline tasks. Once the first job is ready, fetch­ing and de­cod­ing can over­lap.

In DuckDB, we im­ple­mented some­thing sim­i­lar. We have two sep­a­rate thread pools:

REGULAR — This pool con­tains our worker threads (by de­fault: one for each avail­able CPU thread). These are the ones that do real work, like de­cod­ing, joins, and ag­gre­ga­tions. They pri­or­i­tize reg­u­lar work but can also per­form I/O tasks when idle.

ASYNC — A pool of threads in­tended for asyn­chro­nous tasks, pri­mar­ily block­ing I/O.

The main rea­son we have these two dif­fer­ent pools is that, for re­mote I/O, these threads can spend al­most all their time blocked, wait­ing for an HTTP re­sponse, for ex­am­ple, and hence have very lit­tle CPU uti­liza­tion. Because of that, we have many more ASYNC work­ers than sys­tem threads, with the de­fault set­ting be­ing 4 * sys­tem threads and the to­tal be­ing capped at 256.

It’s of ut­most im­por­tance to keep as many of our ASYNC threads busy as pos­si­ble. To en­sure that, we im­ple­ment a read-ahead strat­egy in­stead of is­su­ing reads on de­mand. This means sched­ul­ing fetch tasks ahead of what our reg­u­lar worker threads cur­rently need.

One thing we need to be at­ten­tive to is that read-ahead buys through­put by hold­ing mem­ory. If de­cod­ing is slow and the net­work is fast, prefetched data can ac­cu­mu­late and lead to out-of-mem­ory is­sues. To mit­i­gate this, we also im­ple­mented asyn­chro­nous mem­ory gov­er­nance. Both read-ahead and mem­ory gov­er­nance will be ex­plained in more de­tail in the fol­low­ing sec­tions.

Read-Ahead Queue

The idea of read-ahead is also straight­for­ward. Instead of start­ing a read at the ex­act mo­ment a reg­u­lar worker needs the data, we sched­ule fetch tasks for work that lies fur­ther ahead. While a reg­u­lar worker de­codes the cur­rent job, the ASYNC threads are al­ready pulling in data for the next jobs. The goal is to keep enough fetch tasks in flight to hide the la­tency of re­mote stor­age.

The jobs are units of work that can be sched­uled and processed in­de­pen­dently, and they can be dif­fer­ent de­pend­ing on the un­der­ly­ing file for­mat. For a Parquet file, a job is one row group of one file. For a CSV file, a job is a scan bound­ary that gen­er­ally cov­ers a fixed byte range within the file.

A Parquet job might be bro­ken down into mul­ti­ple fetch tasks de­pend­ing on the query pro­jec­tions, fil­ter push­downs, phys­i­cal col­umn lo­ca­tions, and which nearby byte ranges can be com­bined. The two fetch tasks in the fig­ure be­low are il­lus­tra­tive, as their ex­act group­ing and sizes de­pend on the file and query.

For CSV files, we don’t have the same gran­u­lar­ity of in­for­ma­tion as we do for Parquet files. A job’s fetch tasks load its start­ing buffer if it is not al­ready in mem­ory and, when the scan bound­ary reaches the end of that buffer, the fol­low­ing buffer as well (e.g., to han­dle lines that are split across two buffers).

Jobs

Filling the queue re­quires no ded­i­cated pro­ducer thread. Any reg­u­lar worker that comes look­ing for scan work first tops up the queue as far as it is al­lowed to. The limit is ei­ther given by a user-spec­i­fied num­ber of slots or by a mem­ory bud­get. If there is space, a job and its fetch tasks are cre­ated. The fetch tasks are sched­uled im­me­di­ately on the ASYNC pool, while the job is ad­mit­ted to the read-ahead queue in batch or­der.

ASYNC threads ex­e­cute in­di­vid­ual fetch tasks in­de­pen­dently of the job queue’s claim or­der. Fetch tasks from the same job can run con­cur­rently, al­though no par­tic­u­lar as­sign­ment to ASYNC threads is guar­an­teed. All fetch tasks of a job share a count­down, and the fetch task that brings it to zero com­pletes the job’s I/O.

A worker thread claims the old­est job in the queue and checks that count­down. If I/O is done, the worker starts de­cod­ing the job. If not, it parks the scan task and is free to run other pipeline tasks. The last fetch task then un­blocks the scan task, which may re­sume on any reg­u­lar worker.

Claiming the job also im­me­di­ately frees a queue slot, al­low­ing any reg­u­lar worker look­ing for scan work to pro­duce a re­place­ment job at the back of the queue. The fig­ure be­low de­picts this cy­cle:

Read-ahead cy­cle

Memory Management

Keeping more fetch tasks in flight con­sumes more mem­ory. To de­ter­mine a bud­get and avoid out-of-mem­ory is­sues, we in­tro­duced the read­_a­head­_depth con­fig­u­ra­tion op­tion. It can have three types of val­ues:

-1 (default): un­lim­ited depth, bounded by mem­ory.

N > 0: at most N jobs ahead, with no mem­ory bud­get.

0: read-ahead is off, each scan task sched­ules I/O only for its own job.

To con­fig­ure it, use the SET clause, e.g.:

SET read­_a­head­_depth = 5;

In the de­fault mode, the bud­get is ne­go­ti­ated with the tem­po­rary mem­ory man­ager, which is the same man­ager that splits mem­ory be­tween con­cur­rent joins, sorts, and win­dow op­er­a­tors. When there is a lot of mem­ory pres­sure, for ex­am­ple, be­cause an op­er­a­tor is us­ing a large amount of mem­ory, queue reser­va­tions might in­stantly be over bud­get. In prac­tice, this means that the queue will only al­low one job at a time, and the scan will be­have close to a syn­chro­nous scan.

When the mem­ory-heavy op­er­a­tor fin­ishes, the mem­ory man­ager has more bud­get to give, and the queue fills back up.

Benchmarks

Asynchronous I/O should have the largest ef­fect when the la­tency of syn­chro­nous re­quests pre­vents us from us­ing the avail­able re­mote band­width. To mea­sure this ef­fect, we ran TPC-H Query 6 at SF100, with the data sit­ting on S3, and com­pared the re­sults against DuckDB v1.5.5, our lat­est sta­ble re­lease. The SF100 dataset was writ­ten as a sin­gle file per table for both the Parquet and CSV bench­marks, with the lineitem table con­tain­ing 600,037,902 rows.

For com­pute, we used an EC2 r7i.16xlarge ma­chine (64 vC­PUs and 512 GB of RAM), with both the ma­chine and the S3 bucket with the data lo­cated in the same re­gion. We ex­e­cuted the query five times and re­port the mean ex­e­cu­tion time. The files were never cached (i.e., SET en­able_ex­ter­nal_­file_­cache = false;), mean­ing that every ex­e­cu­tion read the data straight from S3.

Parquet

The Parquet file is ap­prox­i­mately 22 GB and has around 4,880 row groups, with each row group con­tain­ing ap­prox­i­mately 122,880 rows. With asyn­chro­nous I/O, the mean run­time drops from 8.230 sec­onds to 2.844 sec­onds, mak­ing the query al­most faster.

Below we also show the net­work through­put over the course of the query:

Network through­put

In it, we run DuckDB v1.5.5 and two vari­a­tions of DuckDB v2.0.0-dev. One with the read-ahead depth de­ter­mined by the mem­ory gov­er­nor, and one tuned for this ma­chine, where we cap the read-ahead at 64 in-flight jobs and ad­just the I/O set­tings (SET async_threads = 48; SET http_re­tries = 8; SET http_retry_wait­_ms = 50; SET http_retry_back­off = 2). We can see that v2.0.0-dev uses the avail­able band­width much more ef­fec­tively, ap­proach­ing the net­work limit and reach­ing it at sev­eral points. The tuned ver­sion goes fur­ther. With fewer, hot­ter con­nec­tions and cheap re­tries, the through­put vari­ance drops to a min­i­mum and the 25 Gbit/s net­work stays al­most fully sat­u­rated. Its query time was 2.227 sec­onds, re­duc­ing the run­time of the un­tuned v2.0.0-dev run by 21.7% and mak­ing it about 3.7× faster than DuckDB v1.5.5. In com­par­i­son, v1.5.5 stays around 5 Gbit/s be­cause its syn­chro­nous reads do not keep enough re­quests in flight to sat­u­rate the net­work.

One other de­tail worth not­ing is that, in all ex­per­i­ments, a few hun­dred mil­lisec­onds pass be­fore the first bump in net­work traf­fic, fol­lowed by an­other few hun­dred mil­lisec­onds be­fore the main data trans­fer be­gins. The first gap is the time needed to open a DuckDB con­nec­tion, per­form the first TLS hand­shake, and open the file. The bump cor­re­sponds to down­load­ing the file footer, while the sec­ond gap comes from pro­cess­ing the in­for­ma­tion in the footer be­fore ex­e­cut­ing the query. We be­lieve this is an area we can in­ves­ti­gate and op­ti­mize fur­ther be­fore the v2.0 re­lease.

We sam­pled the NICs re­ceived-byte counter every 50 ms and cal­cu­lated the through­put from the change in bytes be­tween sam­ples. We in­de­pen­dently con­firmed that the ma­chine can ac­cess the net­work at 25 Gbit/s with both a DuckDB full-file read and the s5cmd tool.

We sam­pled the NICs re­ceived-byte counter every 50 ms and cal­cu­lated the through­put from the change in bytes be­tween sam­ples. We in­de­pen­dently con­firmed that the ma­chine can ac­cess the net­work at 25 Gbit/s with both a DuckDB full-file read and the s5cmd tool.

Local Disk

Remote stor­age is the main tar­get for asyn­chro­nous I/O, but cold lo­cal reads give us a use­ful con­trast. To mea­sure them, we ran TPC-H Query 6 over the SF100 Parquet file, this time with the file sit­ting on the lo­cal disk of a MacBook Pro (Apple M4 Max, 14 cores, and 36 GB of RAM). Since the ben­e­fit of asyn­chro­nous I/O on lo­cal disks comes from cold reads, we cleared the OS caches (with the ma­cOS purge com­mand) be­tween runs, mak­ing sure every ex­e­cu­tion ac­tu­ally read the file from disk.

We can see that, for cold runs, asyn­chro­nous I/O is ap­prox­i­mately 1.5× faster, re­duc­ing run­time by about 33%. The per­for­mance dif­fer­ence is much smaller than in the cases pre­sented above, due to the SSD hav­ing much lower la­tency and much higher band­width than the EC2/S3 net­work. For hot runs, the dif­fer­ence is neg­li­gi­ble, as there is no disk ac­cess hap­pen­ing if data is prop­erly cached.

Small Files

Partitioned datasets are a par­tic­u­larly rel­e­vant use case here, as par­ti­tion­ing can eas­ily spread the data across many small files. To see how asyn­chro­nous I/O be­haves in this setup, we also per­formed a Parquet run us­ing the same TPC-H SF100 dataset. Instead of us­ing one file, we gen­er­ated 976 files with five row groups each. Each file con­tains ap­prox­i­mately 615,000 rows and is around 22 MB in size.

We can see that v2.0.0-dev de­liv­ers a sim­i­lar per­for­mance im­prove­ment here as it does for the sin­gle-file bench­mark, run­ning around faster. This shows that read-ahead can also par­al­lelize across mul­ti­ple files with­out be­com­ing bot­tle­necked by open­ing files or fetch­ing their foot­ers.

Large Row Groups

We also wanted to see what hap­pens at the other ex­treme, when a Parquet file has only a few very large row groups. For this ex­per­i­ment, we gen­er­ated six ver­sions of the same TPC-H SF100 lineitem table as a sin­gle file, chang­ing only the re­quested row group (RG) size, and ran Q6 us­ing DuckDB v2.0.0-dev. The table be­low re­ports the run­time for each ver­sion.

At first, larger row groups de­crease query times. As we in­crease the row-group size, re­quest la­tency is amor­tized over much larger trans­fers. However, be­yond a cer­tain point, the avail­able par­al­lelism starts to fall. A row group is DuckDB’s unit of Parquet scan par­al­lelism, so ide­ally a scan should ex­pose at least one row group per sys­tem thread. On this 64-vCPU ma­chine, the ver­sion with 64 row groups pro­vides ex­actly that and fin­ishes in 2.27 sec­onds, while the fastest run comes from the ver­sion with 306 row groups, at 2.11 sec­onds.

However, with fewer row groups than threads, we lose par­al­lelism and can no longer sat­u­rate the net­work. For Q6, the pro­jec­tions and the phys­i­cal lo­ca­tion of the columns re­sult in two fetch re­quests per row group. Four row groups there­fore ex­pose only about eight con­cur­rent S3 streams, rais­ing the run­time to 8.01 sec­onds. For the largest con­fig­u­ra­tion, the file con­tains a sin­gle row group, and its I/O is ef­fec­tively re­duced to two gi­ant streams, push­ing the run­time to 25.26 sec­onds. This hap­pens even though bet­ter com­pres­sion makes the file a lit­tle over half the size of the ver­sion with 4,886 row groups. In this case, the ex­tra band­width re­quired by smaller row groups is cheaper than the par­al­lelism lost with ex­tremely large ones.

Concurrent Queries

The ef­fect be­comes even clearer when sev­eral queries run at the same time. For this ex­per­i­ment, we ran TPC-H queries 1, 6, 9, and 18 con­cur­rently against the same SF100 Parquet dataset on S3, us­ing a sin­gle DuckDB in­stance. We picked these queries be­cause they cover a mix of scans, ag­gre­ga­tions, and joins, with dif­fer­ent CPU and mem­ory re­quire­ments. We re­peated the ex­per­i­ment with the de­fault mem­ory con­fig­u­ra­tion and with mem­ory lim­its of 16 GB and 8 GB. The to­tal run­time is the wall-clock time un­til all four queries fin­ish. We re­port the av­er­age and peak CPU uti­liza­tion (number of cores uti­lized), the peak band­width (bw.) and the peak res­i­dent set size (RSS).

With the de­fault mem­ory con­fig­u­ra­tion, DuckDB v1.5.5 keeps an av­er­age of only about 6 of the 64 cores busy. In other words, around 90% of the ma­chine sits idle wait­ing for syn­chro­nous S3 reads. DuckDB v2.0.0-dev, on the other hand, av­er­ages 48 busy cores, reaches all 64 at its peak, and sat­u­rates the 25 Gbit/s net­work. As a re­sult, all four queries fin­ish in less than half the time.

The mem­ory re­sults are also in­ter­est­ing. As we lower the limit, the mem­ory gov­er­nor re­duces the read-ahead back­log, while mem­ory-heavy op­er­a­tors such as those in Q18 can spill to disk. This low­ers the peak phys­i­cal mem­ory used by the DuckDB process (i.e., RSS) of DuckDB v2.0.0-dev from 20.1 GB with the de­fault con­fig­u­ra­tion to 15.7 GB with a 16 GB limit and 11.5 GB with an 8 GB limit. The ad­di­tional spilling and re­duced read-ahead also lower av­er­age CPU uti­liza­tion and in­crease the run­time, but v2.0.0-dev con­tin­ues to sat­u­rate the net­work and re­mains sub­stan­tially faster than v1.5.5 in both cases.

One might no­tice that the 8 GB re­sult still peaks at 11.5 GB of RSS. This is be­cause je­mal­loc keeps re­cently freed pages res­i­dent for about one sec­ond so they can be reused. This mem­ory is no longer counted by DuckDB’s mem­ory man­ager, and v1.5.5 shows the same al­lo­ca­tor be­hav­ior.

One might no­tice that the 8 GB re­sult still peaks at 11.5 GB of RSS. This is be­cause je­mal­loc keeps re­cently freed pages res­i­dent for about one sec­ond so they can be reused. This mem­ory is no longer counted by DuckDB’s mem­ory man­ager, and v1.5.5 shows the same al­lo­ca­tor be­hav­ior.

CSV

The ef­fect is larger on CSV files. The CSV file is 80.89 GB, and asyn­chro­nous I/O re­duces the mean run­time from 878 sec­onds to just 45 sec­onds, mak­ing the query al­most 20× faster. CSV is row-ori­ented, so the scan trans­fers sub­stan­tially more data and per­forms fixed-size buffer reads, mak­ing con­cur­rent re­mote reads es­pe­cially valu­able.

As in the other ex­per­i­ments, we used the de­fault mem­ory-gov­erned read-ahead depth, so this run was not tuned to keep the 25 Gbit/s net­work sat­u­rated on av­er­age.

Conclusion

In this blog post, we pre­sented the re­cent work on asyn­chro­nous I/O for Parquet and CSV files. Most of its ben­e­fit comes from ac­cess­ing re­mote data, but cold lo­cal reads can also ben­e­fit, al­beit less. Next, we plan to add async reads for JSON and DuckDB-native files, as these are the two other for­mats most rel­e­vant to DuckDB core. Formats that live in out-of-tree ex­ten­sions are not on the roadmap yet. We will also in­ves­ti­gate io_ur­ing, Linux’s asyn­chro­nous I/O in­ter­face, which could re­duce sys­tem-call over­head and the num­ber of threads blocked on I/O. If it proves ben­e­fi­cial in prac­tice, we will in­te­grate it into DuckDB. One im­por­tant thing to no­tice is that any of the data lake so­lu­tions sup­ported in DuckDB can al­ready ben­e­fit from asyn­chro­nous I/O au­to­mat­i­cally as long as the un­der­ly­ing data for­mat is Parquet (or CSV, if you are brave enough).

Recent Posts

Thank You for 40 000 Stars on GitHub

The DuckDB team

Announcing DuckDB 1.5.5

The DuckDB team

Announcing DuckDB 1.4.5 LTS (Andium)

The DuckDB team

Language Models Under Pedagogically-Controlled Knowledge Exposure

littlelearner-ll.github.io

Talk to LittleLearner

The hosted 5B model, live in your browser. Open in a new tab ↗ if the chat does­n’t load be­low.

A con­trolled sand­box for study­ing how mod­els ac­quire knowl­edge

Modern LMs are trained on every­thing at once, so it is hard to tell whether a new skill was learned or merely elicited. We con­strain the train­ing dis­tri­b­u­tion it­self: an 88B-token cor­pus fil­tered to the U.S. el­e­men­tary-school cur­ricu­lum, with mod­els trained from scratch on it and matched un­fil­tered con­trols.

Dataset

LittleCurriculum

An 88B-token cor­pus dis­tilled from FineWeb-Edu through a five-stage fil­ter­ing pipeline aligned with Common Core stan­dards (K–5). Concepts, facts, and vo­cab­u­lary taught above Grade 5 are ex­plic­itly ex­cluded.

Models

LittleLearner

Three scales (0.6B / 1.3B / 5B) trained from scratch on LittleCurriculum: chat­table mod­els with an in­ter­pretable knowl­edge bound­ary. Each ships with a matched Unfiltered con­trol for clean com­par­i­son.

Findings

Elicitation, not ac­qui­si­tion

In our ex­per­i­ments, scal­ing, SFT+GRPO post-train­ing, and in-con­text learn­ing am­plify what the cur­ricu­lum taught, but none mean­ing­fully im­proves out-of-scope per­for­mance, in­di­cat­ing that the pre­train­ing fil­ter sets the ef­fec­tive ca­pa­bil­ity ceil­ing.

Model check­points

LittleLearner at three scales (0.6B / 1.3B / 5B), each with a matched Unfiltered con­trol shar­ing its ar­chi­tec­ture, to­kens, and recipe.

Base: the pre­trained model. GRPO: math spe­cial­ists post-trained on MathCAMPS; re­sponses may ex­hibit a ten­dency to­ward math-ori­ented out­put. Chatty: vari­ants tuned for gen­eral chat be­hav­ior.

Capability stays in­side the cur­ricu­lum

Can stan­dard in­ter­ven­tions push a model past what its pre­train­ing data taught it? With the bound­ary un­der ex­per­i­men­tal con­trol, we can ask cleanly. In our ex­per­i­ments, each in­ter­ven­tion am­pli­fies in-scope abil­ity; none of them mean­ing­fully im­proves out-of-scope per­for­mance.

Scaling

Scaling model size im­proves per­for­mance within the mod­el’s con­trolled knowl­edge ex­po­sure and ex­tends mod­estly to prob­lems along the same learn­ing tra­jec­tory, but yields lit­tle im­prove­ment on prob­lems re­quir­ing more ad­vanced ca­pa­bil­i­ties out­side the ex­po­sure.

MathCAMPS ac­cu­racy by grade, across model size

Post-training

Post-training through GRPO sig­nif­i­cantly boosts in-scope K–5 ca­pa­bil­i­ties, but fails to re­cover out-of-scope be­yond-K–5 ca­pa­bil­i­ties, even when train­ing with out-of-scope data.

Post-training am­pli­fies K–5, not the be­yond-K–5 gap

In-context learn­ing

In-context learn­ing with the prompts we test does not un­lock new rea­son­ing ca­pa­bil­i­ties in be­yond-K–5 for our trained 5B LittleLearner.

Accuracy by prompt­ing con­di­tion

What will you teach it?

Because LittleLearner’s train­ing ex­po­sure is ex­plic­itly spec­i­fied, be­hav­ioral and rep­re­sen­ta­tional changes can be re­lated di­rectly to the con­cepts you in­tro­duce. Three di­rec­tions we’re ex­cited about:

01

RL & dis­cov­ery

Can RL cre­ate ca­pa­bil­ity?

The prior is re­stricted to K–5, so ca­pa­bil­i­ties that emerge un­der RL can be at­trib­uted to the RL process it­self. A tractable proxy for re­ward-dri­ven dis­cov­ery.

02

Continual learn­ing

Watch a con­cept be­ing learned

Introduce neg­a­tive num­bers and mea­sure sam­ple ef­fi­ciency, re­ten­tion, and in­ter­fer­ence. Or probe be­hav­ior near the bound­ary: does it an­swer, ab­stain, or hal­lu­ci­nate?

03

Educational sci­ence

Machine vs. child learn­ers

Specified ex­po­sure en­ables con­trolled hu­man-model com­par­i­son. Do mod­els and chil­dren need sim­i­lar ex­po­sure to learn frac­tions, or make sim­i­lar er­rors on word prob­lems?

Your turn

Bring your own ques­tion

A known bound­ary turns your idea into a clean ex­per­i­ment!

If you find this work use­ful

Please cite our pa­per:

@misc{littlelearner2026, ti­tle={Lit­tle­Learner: Language Models Under Pedagogically-Controlled Knowledge Exposure}, au­thor={Fan­fei Li and Jana Zeller and Manuel Prada-Corral and Thaddäus Wiedemer and Prasanna Mayilvahanan and Ryan Cotterell and Wieland Brendel}, year={2026}, eprint={2608.13545}, archivePre­fix={arXiv}, pri­ma­ryClass={cs.CL}, url={https://​arxiv.org/​abs/​2608.13545} }

Models Are Getting Dumber on Purpose

w4g1.dev

Reasoning scores keep climb­ing while per-to­ken com­pute keeps drop­ping. GLM-5.2 scores 99.2% on AIME 2026 with about 40 bil­lion pa­ra­me­ters ac­tive per to­ken. Qwen3.5 scores 91.3% with 17 bil­lion ac­tive. DeepSeek V4-Flash runs 13 bil­lion ac­tive. For scale, GPT-4 was ru­mored to run around 280 bil­lion ac­tive pa­ra­me­ters in 2023, and it could barely solve an AIME prob­lem. At the small end, Qwen3.5 9B fits in 6GB of VRAM quan­tized and roughly dou­bles the score of the next best model un­der 10B pa­ra­me­ters on Artificial Analysis’s in­tel­li­gence in­dex. If you only looked at math and code bench­marks, you’d con­clude that mod­els are get­ting smarter per pa­ra­me­ter at an ab­surd rate.

They are, on those bench­marks. Ask the same mod­els a plain fac­tual ques­tion and the pic­ture flips. On SimpleQA, a bench­mark of fac­tual re­call with no tools al­lowed, the cur­rent leader is Gemini 2.5 Pro at 53%, so the best re­call money can buy still misses half the ques­tions. The small mod­els barely reg­is­ter. Artificial Analysis mea­sures Qwen3.5 4B and 9B at hal­lu­ci­na­tion rates of 80 to 82% on its knowl­edge bench­mark, which means that when they don’t know a fact, which is most of the time, they make one up. Ask the 9B for the birth year of a mi­nor 19th-century math­e­mati­cian and you get a con­fi­dent, plau­si­ble, wrong an­swer. The pa­ra­me­ter count did­n’t drop for free. Labs are trad­ing world knowl­edge for rea­son­ing skill, and the trade is de­lib­er­ate.

What the pa­ra­me­ters were for

Facts take space. Research on knowl­edge ca­pac­ity (the Physics of Language Models” se­ries has the clean­est mea­sure­ments) puts it on the or­der of two bits of fac­tual knowl­edge per pa­ra­me­ter. If you want a model that knows the birth year of every mi­nor Wikipedia fig­ure, the pop­u­la­tion of every Dutch mu­nic­i­pal­ity, and the ar­gu­ment or­der of every func­tion in every npm pack­age, you pay for that in weights, and it’s a big part of why fron­tier mod­els grew to tril­lions of pa­ra­me­ters.

Reasoning com­presses much bet­ter than facts do, be­cause it’s a rel­a­tively small set of pro­ce­dures ap­plied over and over: break the prob­lem into parts, track in­ter­me­di­ate state, check your own work, back­track when a step fails. Distillation and re­in­force­ment learn­ing on ver­i­fi­able tasks turn out to trans­fer those pro­ce­dures into small mod­els re­mark­ably well. Phi-4 is 14 bil­lion pa­ra­me­ters, trained heav­ily on syn­thetic text­book-style data, and it’s good at math and bad at trivia, which tells you ex­actly what its train­ing data con­tained. That mix used to look like a lim­i­ta­tion of the syn­thetic-data ap­proach. It now looks like the de­sign goal.

The knowl­edge that sur­vives the trade has a shape. These mod­els are gen­er­al­ists: they know a lit­tle about nearly every­thing and al­most noth­ing in depth. Ask one about PostgreSQL and it knows what it is, what it’s good at, and roughly how MVCC works, but ask which ver­sion added a spe­cific plan­ner fea­ture and you’re back to in­vented facts. That’s the right layer to keep in weights, be­cause breadth is what lets a model un­der­stand what a ques­tion is about, know what to look up, and judge whether a source is plau­si­ble. The depth is cheap to re­trieve and ex­pen­sive to store, so it’s the part that goes.

Facts rot, pro­ce­dures don’t

A fron­tier train­ing run takes months and costs hun­dreds of mil­lions of dol­lars, and the mo­ment it fin­ishes, the facts in­side it start go­ing stale. Library APIs change, prices change, peo­ple change jobs, and half of what a 2024 model be­lieved about the JavaScript ecosys­tem was out­dated be­fore the model shipped. Every fact you bake into weights has a shelf life, and the only way to re­fresh it is an­other train­ing run.

The pro­ce­dures don’t rot. Algebra worked the same way in 1970 as it does now, and so does break­ing a prob­lem down or spot­ting a con­tra­dic­tion be­tween two sources. A model that’s mostly pro­ce­dure and only lightly loaded with facts does­n’t age the way a knowl­edge-heavy model does. Its train­ing cut­off mat­ters much less, be­cause the cur­rent state of the world was never sup­posed to live in the weights in the first place. I think this is the best ar­gu­ment for the whole ap­proach: it de­cou­ples the ex­pen­sive, slow ar­ti­fact (the trained model) from the thing that changes daily (what’s true).

The har­ness car­ries the knowl­edge

If the model does­n’t know things, some­thing else has to, and that some­thing is the har­ness: re­trieval over a knowl­edge base, tool calls, web search, a filesys­tem full of docs. I wrote ear­lier that Rust is a har­ness for agents, a source of cheap ma­chine-check­able feed­back. This is the same shape from the other side. The model con­tributes rea­son­ing, and every­thing it rea­sons about gets sup­plied at run­time.

You can al­ready watch agents work this way. A cod­ing agent does­n’t need to have mem­o­rized your de­pen­den­cy’s API sur­face, be­cause it greps node_­mod­ules or reads the docs be­fore call­ing any­thing, and its an­swer is grounded in the ver­sion you ac­tu­ally have in­stalled rather than whichever ver­sion dom­i­nated the train­ing data. The re­call that used to be a fixed cost in every for­ward pass be­came an on-de­mand lookup.

A fron­tier model on your GPU

Follow the trend a cou­ple of years out and I think we get a model with fron­tier-qual­ity rea­son­ing, Fable-quality, that runs on a sin­gle con­sumer GPU. The com­pute half is nearly there. DeepSeek V4-Flash rea­sons with about 13 bil­lion ac­tive pa­ra­me­ters per to­ken, well within con­sumer-GPU range. What does­n’t fit is the other 271 bil­lion pa­ra­me­ters sit­ting in its ex­perts, and ex­pert lay­ers are mostly fact stor­age. That’s the part this whole trade makes op­tional. Strip the knowl­edge out and to­tal size shrinks to­ward ac­tive size, and a 20 to 40B model at 4-bit quan­ti­za­tion fits on the 24GB card that’s been sit­ting in gam­ing PCs since 2022.

The catch is that it won’t know much. Ask it a bare fac­tual ques­tion with no tools at­tached and the right be­hav­ior is to say it does­n’t know and go look it up. Paired with a de­cent har­ness, that’s most of what I use a fron­tier model for to­day, run­ning lo­cally with no per-to­ken bill and no data leav­ing the ma­chine.

This mostly solves hal­lu­ci­na­tion

The part I find most promis­ing is what this does to hal­lu­ci­na­tion. When a fact lives in weights, a wrong fact is un­find­able and un­fix­able. You can’t grep the weights, you can’t diff them against last month, and cor­rect­ing one er­ror means a fine-tune that might break who knows what else. The model states the wrong fact with the same flu­ent con­fi­dence as a right one, and there’s no ar­ti­fact to check it against.

When the fact lives out­side the model, a wrong an­swer has an ad­dress. The model cites a doc­u­ment, so you can open the doc­u­ment. If the doc­u­ment is wrong, you edit the doc­u­ment, and every fu­ture query gets the cor­rec­tion, which beats wait­ing for the next train­ing run by roughly a year. Retrieval does­n’t get you to zero, since a model can still mis­read a source or stitch two of them to­gether wrong, but a claim with a source is check­able and a claim from weights is­n’t. A wrong fact in a knowl­edge base is an or­di­nary data bug, the kind we al­ready know how to trace, fix, and write a re­gres­sion test for.

There’s a ver­sion of this fu­ture where the model card stops list­ing a knowl­edge cut­off at all, be­cause what’s left in the weights goes stale on a scale of years in­stead of weeks. The model just gets handed the world’s cur­rent state at run­time, the same way a CPU gets handed a pro­gram.

Who Are the Token Brokers?

vectoral.com

threat-re­search llm-se­cu­rity

August 10, 2026 Matt Lenhard 5 min read

Share

Where This Started

This is a fol­low-up ar­ti­cle to a piece I re­cently wrote about the to­ken re­lay mar­ket. Noticeably ab­sent from that piece was a men­tion of the rise of token bro­kers” — peo­ple who buy un­used cred­its from star­tups and then re­sell them.

I first heard about to­ken bro­kers while chat­ting with a good friend of mine who was re­ceiv­ing of­fers for Anthropic to­kens at steep dis­counts.

It was­n’t just him, though. As I started talk­ing to more founders about what I was build­ing, they said the same thing: they were get­ting a lot of in­bound email from peo­ple look­ing to buy or sell off-mar­ket in­fer­ence.

Startups swap­ping cred­its is noth­ing new, and I knew this was hap­pen­ing in sev­eral startup fo­rums and groups, but this was when I re­al­ized that the mar­ket was be­ing com­mer­cial­ized.

So I did what any nor­mal per­son would do. I got the bro­kers’ email ad­dresses and started email­ing them to learn more.

Before my own out­reach, it’s worth see­ing what founders are ac­tu­ally re­ceiv­ing. Both of these were for­warded to me by friends.

I started by sourc­ing a few email ad­dresses from friends. The first two emails I sent bounced, but the third was a hit. Here’s a screen­shot of that con­ver­sa­tion:

What’s in­ter­est­ing is the amount of sup­ply. The seller was of­fer­ing $100k in spend per day.

They aren’t hand­ing out the provider keys di­rectly; in­stead, they act as a proxy that prob­a­bly picks from a pool of keys and for­wards the re­quest.

The Listings

Credit Marketplaces

There are a few web­sites pro­mot­ing credit bro­ker­ing as well. One of them, AI Credits, bills it­self as a credit mar­ket­place. For an­other fla­vor of the pure-play credit re­seller mar­ket­places, take a look at AICreditMart.

These sites of­fer cred­its at most of the ma­jor cloud and in­fer­ence providers.

AI Credits’ on­board­ing process is pretty straight­for­ward, and you can even se­lect your pre­ferred de­liv­ery method as the seller.

I went ahead and listed my cred­its, which are still pend­ing ap­proval.

Bulk Discounts

Another site that I found through a friend was CheapCredits. This site po­si­tions it­self as a router that is able to achieve its dis­counts through bulk pric­ing.”

I no­ticed that this was a trend with a num­ber of sites that I be­lieve are act­ing as credit bro­kers. They pre­sent them­selves as be­ing able to of­fer dis­counts based on bulk pur­chases. Some other ex­am­ples in­clude Tokvana and Neokens.

Having spent time in the in­dus­try, I’d say that a 40% dis­count is very un­likely un­less you are one of the provider’s top cus­tomers. My hunch is that CheapCredits is ac­quir­ing the sup­ply in other ways.

CheapCredits even has a Data Processing Agreement for any­one look­ing to stay GDPR com­pli­ant.

The Message Boards

I checked where you’d ex­pect to find un­der­ground mar­ket­places.

Telegram had a few chan­nels, with one be­ing rel­a­tively ac­tive.

There are also spo­radic Reddit posts.

If you’ve been hang­ing out in any of the closed-off startup groups, I’m sure you’ve seen a num­ber of these posts as well.

So How Big Is This Market?

My rough es­ti­mate is that, across the sites, fo­rums, and re­sellers I looked at, there are prob­a­bly tens of mil­lions of these cred­its be­ing of­fered.

Unfortunately, when you try to of­fer nice things, abuse is­n’t far be­hind. Tokens have be­come a pseudo-cur­rency, and there is enough liq­uid­ity in the mar­ket to al­low for a lot of abuse. As we see the mar­ket turn and com­pa­nies be­come more aware of costs, crack­downs on this type of abuse prob­a­bly aren’t far be­hind.

Sources

Company and site names be­low are as they pre­sent them­selves pub­licly. Screenshots are from my own out­reach and from brows­ing the sites as a prospec­tive buyer and seller.

Previous piece: An Inside Look at the Relay Market Powering Token Resellers and Fraud

Credit mar­ket­places: AI Credits, AICreditMart

Bulk-discount routers: CheapCredits (cheapcredits.ai), Tokvana, Neokens

Direct out­reach: email ex­change with a bro­ker of­fer­ing $100k/day in spend

Share

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.