10 interesting stories served every morning and every evening.

Qwen/Qwen3.8-27B-FP8 · Hugging Face

huggingface.co

This repos­i­tory con­tains FP8-quantized model weights and con­fig­u­ra­tion files for the post-trained model in the Hugging Face Transformers for­mat. These ar­ti­facts are com­pat­i­ble with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. The quan­ti­za­tion method is fine-grained fp8 quan­ti­za­tion with block size of 128, and its per­for­mance met­rics are nearly iden­ti­cal to those of the orig­i­nal model.

This repos­i­tory con­tains FP8-quantized model weights and con­fig­u­ra­tion files for the post-trained model in the Hugging Face Transformers for­mat.

These ar­ti­facts are com­pat­i­ble with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

The quan­ti­za­tion method is fine-grained fp8 quan­ti­za­tion with block size of 128, and its per­for­mance met­rics are nearly iden­ti­cal to those of the orig­i­nal model.

For users seek­ing man­aged, scal­able in­fer­ence with­out in­fra­struc­ture main­te­nance, the of­fi­cial Qwen API ser­vice is pro­vided by Qwen Cloud. In par­tic­u­lar, Qwen3.8 – 27B will be avail­able as a hosted ver­sion with more pro­duc­tion fea­tures, e.g., 1M con­text length by de­fault, of­fi­cial built-in tools. For more in­for­ma­tion, please re­fer to the Qwen3.8 – 27B Overview. The ser­vice is com­ing soon. Stay tuned for up­dates.

For users seek­ing man­aged, scal­able in­fer­ence with­out in­fra­struc­ture main­te­nance, the of­fi­cial Qwen API ser­vice is pro­vided by Qwen Cloud.

In par­tic­u­lar, Qwen3.8 – 27B will be avail­able as a hosted ver­sion with more pro­duc­tion fea­tures, e.g., 1M con­text length by de­fault, of­fi­cial built-in tools. For more in­for­ma­tion, please re­fer to the Qwen3.8 – 27B Overview. The ser­vice is com­ing soon. Stay tuned for up­dates.

Following the wide­spread com­mu­nity adop­tion of the Qwen3.5 and Qwen3.6 se­ries, we are pleased to in­tro­duce Qwen3.8, the most ca­pa­ble gen­er­a­tion in the Qwen open-model fam­ily to date.

Built on the ar­chi­tec­tural foun­da­tion of Qwen3.5, Qwen3.8 de­liv­ers sub­stan­tial gains across cod­ing, pro­fes­sional work, re­search, and long-hori­zon agen­tic tasks. Qwen3.8 – 27B brings these ad­vances to a com­pact, de­ploy­ment-friendly dense model: a na­tive vi­sion-lan­guage model that un­der­stands im­ages and videos, with flex­i­ble think­ing con­trol, de­signed to carry com­plex, multi-step tasks through to com­ple­tion with greater re­li­a­bil­ity.

Qwen3.8 Highlights

Qwen3.8 – 27B fea­tures the fol­low­ing en­hance­ments:

Core Capabilities: Comprehensive im­prove­ments across cod­ing, pro­fes­sional work, re­search, and long-hori­zon agen­tic tasks.

Agent Execution: Stronger au­tonomous plan­ning and bet­ter han­dling of en­vi­ron­ment feed­back, lead­ing to more re­li­able end-to-end task com­ple­tion.

Downstream Compatibility: Broader sup­port for pop­u­lar har­nesses and de­vel­op­ment tools, mak­ing it eas­ier to in­te­grate into your ex­ist­ing stack.

Flexible Thinking Control: Thinking mode is on by de­fault and can be dis­abled per re­quest; rea­son­ing depth can be tuned with rea­son­ing_­ef­fort, and rea­son­ing con­text from his­tor­i­cal mes­sages is re­tained via pre­serve_­think­ing.

Vision-Language Understanding: Native sup­port for im­age and video un­der­stand­ing, from STEM di­a­grams and doc­u­ments to hour-scale videos.

Model Overview

Type: Causal Language Model with Vision Encoder

Training Stage: Pre-training & Post-training

Language Model Number of Parameters: 27B Hidden Dimension: 5120 Token Embedding: 248,320 (Padded) Number of Layers: 64 Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN)) Gated DeltaNet: Number of Linear Attention Heads: 48 for V and 16 for QK Head Dimension: 128

Gated Attention: Number of Attention Heads: 24 for Q and 4 for KV Head Dimension: 256 Rotary Position Embedding Dimension: 64

Feed Forward Network: Intermediate Dimension: 17,408

LM Output: 248,320 (Padded) MTP (Multi-Token Prediction): trained with mul­ti­ple steps

Number of Parameters: 27B

Hidden Dimension: 5120

Token Embedding: 248,320 (Padded)

Number of Layers: 64

Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))

Gated DeltaNet: Number of Linear Attention Heads: 48 for V and 16 for QK Head Dimension: 128

Number of Linear Attention Heads: 48 for V and 16 for QK

Head Dimension: 128

Gated Attention: Number of Attention Heads: 24 for Q and 4 for KV Head Dimension: 256 Rotary Position Embedding Dimension: 64

Number of Attention Heads: 24 for Q and 4 for KV

Head Dimension: 256

Rotary Position Embedding Dimension: 64

Feed Forward Network: Intermediate Dimension: 17,408

Intermediate Dimension: 17,408

LM Output: 248,320 (Padded)

MTP (Multi-Token Prediction): trained with mul­ti­ple steps

Context Length: 262,144 na­tively and ex­ten­si­ble up to 1,000,000 to­kens.

Benchmark Results

Text Performance

Agentic ter­mi­nal cod­ing

Terminal Bench 2.1 (Terminus)

Agentic cod­ing

SWE-bench Pro

Repo-level code gen­er­a­tion

NL2Repo-Bench

Agentic cod­ing

DeepSWE 1.1

Software en­gi­neer­ing

QwenSWEBench

Long-horizon of­fice work

CoWorkBench

Professional job tasks

JobBench

Frontier agen­tic tasks

Agents’ Last Exam

Pass@1

20.4

Score

42.9

Pass@1

10.6

Score

27.3

Pass@1

13.2

Score

33.6

Instruction fol­low­ing

IFBench

Scientific rea­son­ing

GPQA Diamond

Multidisciplinary rea­son­ing

HLE

Competitive cod­ing

LiveCodeBench v6

SWE-bench Pro: Except for Opus4.6 Max, which uses the of­fi­cially re­ported score, all mod­els are eval­u­ated with the Claude Code har­ness at temp=1.0, top_p=0.95, and a 256K con­text win­dow. Problematic tasks were cor­rected, and all base­line mod­els were re-eval­u­ated on the re­fined bench­mark.

NL2Repo-Bench: Evaluated with the Claude Code har­ness. To pre­vent re­ward hack­ing, we dis­able Bash com­mands that at­tempt to ac­cess the spe­cific repos­i­tory, such as pip down­load, pip in­stall, and git clone.

DeepSWE 1.1: Evaluated with the Claude Code har­ness at temp=1.0, top_p=0.95, and a 256K con­text win­dow.

QwenSWEBench: In-house cod­ing bench­mark for eval­u­at­ing mod­els’ soft­ware en­gi­neer­ing ca­pa­bil­i­ties. Evaluated with the Claude Code har­ness. Reporting avg@3 with an 8-hour time­out, max_­to­kens=32,768, tem­per­a­ture=1.0, and a 256K con­text win­dow.

CoWorkBench: In-house cowork bench­mark for eval­u­at­ing long-hori­zon tasks across com­puter sci­ence, fi­nance, law, med­ical, and other pro­duc­tiv­ity do­mains.

HLE: Judged by GPT-4o.

The best re­sult in each row is shown in bold.

Empty cells (–) in­di­cate that re­sults are not yet avail­able or not ap­plic­a­ble.

VL Performance

Computer use

OSWorld-Verified

Browser use

WebArena-Verified

Mobile use

AndroidWorld

Application recre­ation

RecreationBench

Multimodal tool use

ClawEval-MM

Pass@3

57.4

Average

How we tracked down a 16-year-old SQLite bug

tailscale.com

At the end of last year, our up­time was pretty shaky. You can see this trend on our sta­tus page, and that in­sta­bil­ity con­tin­ued into the new year. Many of these out­ages were caused by a sin­gle bug, deep in SQLite. It took months of in­tense foren­sics to track it down.

Now we’re in sum­mer, we’re con­fi­dent that we’ve found the bug, that we un­der­stand it—and more im­por­tantly, that we’ve fixed it.

We know our cus­tomers ex­pect Tailscale to be a re­li­able ser­vice, and for sev­eral months we did­n’t live up to that promise. That’s dis­rup­tive, and we’re sorry. We’re pub­lish­ing this blog post to ex­plain what went wrong, how we re­sponded, and how we ul­ti­mately helped to un­cover a long-stand­ing bug in the heart of the SQLite data­base.

Tailscale’s data­base ar­chi­tec­ture

While our clients in­ter­act with our con­trol plane as a sin­gle pub­lic end­point (controlplane.tailscale.com), in­ter­nally, our con­trol plane is split into a se­ries of co­or­di­na­tion servers (or shards”). Each tail­net lives on one in­ter­nal shard at a time, but can mi­grate seam­lessly from one to an­other. These shards are an in­ter­nal im­ple­men­ta­tion de­tail: you don’t know what shard your tail­net is on, and you never need to.

Each shard has an SQLite data­base that holds all the in­for­ma­tion about the tail­nets on that shard. A sin­gle Go process ex­clu­sively ac­cesses that data­base, and serves the con­trol plane for those tail­nets. This sin­gle-writer de­sign is ex­actly how SQLite is meant to be used.

We’ve used SQLite as our pri­mary data­base since 2022, and we chose it be­cause it’s well-known, re­li­able, and widely used. SQLite is boring tech­nol­ogy”—in a good way. Many com­pa­nies use SQLite in much larger de­ploy­ments with­out is­sue, and we ex­pected the same stress-free us­age.

In our cur­rent backup pipeline, we take a com­plete snap­shot of the data­base every few min­utes, then up­load the en­tire SQLite file to an S3 bucket. We’d been run­ning this setup with­out in­ci­dent since early 2023.

Fast for­ward to August last year, when a data pipeline that reads those S3 back­ups re­ported an er­ror in one of our data­bases. We ran SQLite’s PRAGMA in­tegri­ty_check com­mand against the backup, and found it was in­deed cor­rupted. SQLite cor­rup­tion is pos­si­ble, but it’s highly un­usual and not some­thing you should en­counter in nor­mal op­er­a­tion. We re­paired the af­fected data­base, and in­ves­ti­gated the cause, but to no avail.

When op­er­at­ing at scale, even rare events can oc­cur with some fre­quency, so we should have been un­sur­prised when it hap­pened again—and again, and again, and again. In to­tal, we faced 19 sep­a­rate in­stances of data­base cor­rup­tion over six months be­fore we fi­nally re­solved the un­der­ly­ing bug.

When you hear the phrase database cor­rup­tion”, it’s nat­ural to worry about data loss. Because our con­trol plane only han­dles con­fig­u­ra­tion data, these data­bases con­tain meta­data about your tail­net and de­vices, but never your pri­vate en­cryp­tion keys or net­work traf­fic. In the ear­li­est in­ci­dents, the re­cov­ery process meant a hand­ful of newly added de­vices or con­fig­u­ra­tion changes did­n’t per­sist, and a small amount of meta­data had to be re-en­tered.

Whenever cor­rup­tion oc­curred, we had to stop the con­trol plane process on the shard while we re­paired or re­stored the data­base. This was painful for tail­nets on that shard, be­cause their en­tire con­trol plane dis­ap­peared dur­ing that re­cov­ery win­dow. In the early in­ci­dents, that down­time was over an hour, but we grad­u­ally sped up the re­cov­ery process over sub­se­quent in­ci­dents.

Each tail­net is a mesh net­work, where de­vices make peer-to-peer WireGuard® con­nec­tions to each other. When a de­vice joins the tail­net, it has to get a list of other de­vices from the con­trol plane be­fore it can es­tab­lish new con­nec­tions—so if a de­vice came on­line dur­ing the SQLite down­time, it could­n’t con­nect. While the data­base was be­ing re­paired, de­vices al­ready on­line re­mained con­nected to each other, but they could­n’t learn about changes to the net­work. Those tail­nets also tem­porar­ily lost ac­cess to the web-based ad­min con­sole and the Tailscale API.

There’s also a broader im­pact on trust. We post a global in­ci­dent on our sta­tus page even when only a small num­ber of tail­nets are af­fected. Many peo­ple saw a sta­tus page event for an in­ci­dent that did­n’t af­fect them. Indeed, the ma­jor­ity of shards and tail­nets were never in­volved in a data­base cor­rup­tion in­ci­dent! Nonetheless, re­peated down­time erodes trust, whether or not you’re di­rectly af­fected.

From the very first in­stance of cor­rup­tion, we knew this was a se­ri­ous threat to our re­li­a­bil­ity, and we threw a lot of en­gi­neer­ing time at the prob­lem—but the fix was­n’t easy.

Trying to find the fault

This bug re­sisted all our ini­tial at­tempts to find it.

We looked at re­cent changes, but there weren’t any that seemed rel­e­vant. Nobody had been work­ing on our low-level code that in­ter­acts with SQLite, be­cause it had all been writ­ten years ago and pre­sented no is­sues up un­til that point. We re-re­viewed all of that code with a fine-toothed comb to look for pre­vi­ously missed bugs, but we did­n’t find any­thing that would cause the cor­rup­tion we were see­ing.

We looked for com­mon fac­tors be­tween cor­rup­tion in­ci­dents, but we could­n’t find any. It was­n’t tied to a sin­gle shard, or cus­tomer, or tail­net fea­ture, or time of day, or load level. We were at a loss for what might be trig­ger­ing the be­hav­iour.

This lack of re­li­able trig­ger con­di­tions meant we could­n’t re­pro­duce the bug syn­thet­i­cally. Instead, we had to rely on de­ploy­ing pas­sive, foren­sic teleme­try in our live en­vi­ron­ment to catch the cor­rup­tion red-handed. Gathering live di­ag­nos­tics for a data­base is­sue is the last thing we wanted to do, but we had no choice.

As an ad­di­tional com­pli­ca­tion, the cor­rup­tion did­n’t oc­cur on a reg­u­lar sched­ule. Sometimes in­ci­dents would be hours apart, other times weeks. This made it dif­fi­cult to pre­dict progress or plan fur­ther work, be­cause we were never sure when we’d get our next di­ag­nos­tic dump. We had a six-week pe­riod be­tween October and December when there were no cor­rup­tion in­ci­dents, be­fore they re­turned as an un­wel­come Christmas pre­sent.

Because this would­n’t be a quick or easy fix, we reached out to the SQLite de­vel­op­ers for a pro­fes­sional sup­port con­tract. This was a great de­ci­sion. It gave us di­rect ac­cess to their deep ex­per­tise and ex­pe­ri­ence, and we had many de­tailed tech­ni­cal con­ver­sa­tions about our ar­chi­tec­ture and our in­ci­dents.

Between Tailscale en­gi­neer­ing and the SQLite core de­vel­op­ers, we mapped out sev­eral the­o­ries for what might be caus­ing the cor­rup­tion—in­clud­ing bro­ken POSIX locks on close(), mis­man­ag­ing mem­ory owned by SQLite, or ac­ci­den­tally us­ing SQLite from mul­ti­ple threads while dis­abling thread safety. After every in­ci­dent, we gath­ered more data, added more di­ag­nos­tics, and sys­tem­at­i­cally ruled out these the­o­ries. We were grad­u­ally con­verg­ing on the true bug.

The trans­ac­tions that did­n’t bark

While we were in­ves­ti­gat­ing the root cause, we still had a live plat­form to run. We took ag­gres­sive steps to au­to­mate re­cov­ery and min­i­mize down­time:

Configuring our con­trol plane shards to hard-stop im­me­di­ately upon en­coun­ter­ing cor­rup­tion

Deploying an au­to­mated backup mon­i­tor that con­tin­u­ously ran PRAGMA in­tegri­ty_check over our back­ups

Improving our run­books and on-call train­ing

These ef­forts cut our re­sponse time to un­der an hour—and then we dis­cov­ered an un­ex­pected clue.

We wanted a way to re­store ser­vice that did­n’t in­volve rolling back to the last known-good backup (which would lose a lot of data) or re­pair­ing the known-cor­rupted data­base (which was po­ten­tially risky).

To do this, we built a trans­ac­tion log­ging pipeline. We streamed every SQL state­ment that mod­i­fied the data­base to a sep­a­rate log file. Because SQLite is a sin­gle-writer data­base with se­ri­al­is­able trans­ac­tions, our trans­ac­tion his­tory was com­pletely lin­ear and de­ter­min­is­tic. (This would­n’t be true in a multi-writer data­base like Postgres or MySQL.) Replaying those trans­ac­tions against the lat­est known-good backup should re­store the data­base to its most re­cent state, safely by­pass­ing the cor­rup­tion.

This pipeline worked, but then it did some­thing even bet­ter: it gave us a clue.

In two in­ci­dents, our trans­ac­tion logs failed to re­play cleanly. Upon closer in­spec­tion, we dis­cov­ered that data writ­ten and com­mit­ted by one trans­ac­tion was in­ex­plic­a­bly in­vis­i­ble to later trans­ac­tions. A write had van­ished into thin air with­out rais­ing an er­ror. That should be im­pos­si­ble!

The writ­ing on the WAL

As these in­ci­dents were on­go­ing, the SQLite de­vel­op­ers had been de­vel­op­ing a new de­bug­ging tool. For a while, we’d sus­pected that the bug was some­where in the check­point process. They were build­ing a new tool to give bet­ter vis­i­bil­ity into what was hap­pen­ing dur­ing check­points.

To un­der­stand what this tool found, we need to briefly ex­plain how SQLite check­points work.

A SQLite data­base is made of a se­ries of pages”, tiny blocks of in­for­ma­tion. When you up­date the data­base, some of those pages need to be re­placed with new pages with the up­dated in­for­ma­tion.

For bet­ter per­for­mance and greater con­cur­rency, we run SQLite with Write-Ahead Logging, which means new pages aren’t writ­ten di­rectly to the data­base file. Instead, they’re writ­ten to the write-ahead log” or WAL file”.

New pages can’t be writ­ten to the WAL file in­def­i­nitely; at some point they have to be copied back to the main data­base file. This process is called checkpointing”.

In most de­ploy­ments, SQLite it­self de­cides when to do a check­point, and the process is in­vis­i­ble to the end user and de­vel­oper. In our con­trol plane, we take man­ual con­trol of the check­point process so we can run fast and con­sis­tent back­ups. This non-stan­dard ap­proach seemed sus­pi­cious as we steadily elim­i­nated po­ten­tial causes.

One clue was that dur­ing cor­rup­tion in­ci­dents, our met­rics showed that SQLite would re­port copy­ing more pages from the WAL file than were ac­tu­ally avail­able. If there are 10 pages in the WAL file and 20 pages get copied to the data­base, some­thing is clearly wrong.

To un­der­stand what was hap­pen­ing dur­ing these faulty check­points, the SQLite de­vel­op­ers cre­ated a new de­bug­ging tool for the vir­tual filesys­tem layer.

SQLite is split into sev­eral lay­ers. The top layer is the parser and code gen­er­a­tor, which con­verts SQL state­ments into SQLite’s in­ter­nal data struc­tures. These data struc­tures get passed to the pager, which splits them into the in­di­vid­ual pages to be writ­ten to disk. Actually writ­ing them to disk is han­dled by the OS in­ter­face, or virtual filesys­tem”. Currently SQLite has two main­stream vir­tual filesys­tem im­ple­men­ta­tions—Unix and Windows.

If you’re in­ter­ested in a deeper dive on these in­ter­nals, I rec­om­mend this lec­ture by Richard Hipp, the pri­mary au­thor of SQLite.

This ap­proach al­lows you to re­place dif­fer­ent lay­ers with dif­fer­ent im­ple­men­ta­tions, or wrap an ex­ist­ing layer to get more in­for­ma­tion. To help di­ag­nose our prob­lem, the SQLite de­vel­op­ers cre­ated a wrap­per around the vir­tual filesys­tem that writes ad­di­tional trac­ing in­for­ma­tion and logs about changes to the data­base. This wrap­per is called the tm­stm­pvfs shim, and the source code is avail­able in the SQLite pub­lic repos­i­tory.

We de­ployed the shim into our live en­vi­ron­ment, and waited for the next cor­rup­tion to oc­cur. Fortunately, we did­n’t have to wait long.

The WAL-Reset bug

After our next cor­rup­tion in­ci­dent, the ad­di­tional logs from the new tm­stm­pvfs shim al­lowed the SQLite de­vel­op­ers to find and fix the bug: a rare data race in the SQLite source code be­tween a check­point and a write trans­ac­tion.

In par­tic­u­lar, if a write oc­curs at a spe­cific time dur­ing a check­point, the check­point­ing process gets con­fused—it thinks some of the pages have been copied from the WAL into the main data­base file, but they haven’t. Those pages never get writ­ten to the data­base file, and that data is per­ma­nently lost. The data­base file be­comes cor­rupt, be­cause other pages which ref­er­ence those pages—such as an in­dex—are writ­ten to the data­base.

The SQLite de­vel­op­ers named this the WAL-Reset bug”, and they es­ti­mate it was pre­sent in SQLite for at least 16 years. It could ex­ist that long be­cause it was rare—so rare, the SQLite de­vel­op­ers had to add code to de­lib­er­ately trig­ger it in their test­ing en­vi­ron­ments. Their fix adds an ad­di­tional check to the check­point­ing func­tion which de­tects when the WAL has been re­set by an­other thread.

They con­firmed that this bug caused all of the baf­fling be­hav­iour we’d seen. It ex­plained the cor­rup­tion, the trans­ac­tion logs that would­n’t ap­ply cleanly, and the in­con­sis­tent check­point sta­tis­tics. They also ex­plained why we were more likely to hit the bug than other SQLite users: we take man­ual con­trol of the check­point­ing process, and we check­point very ag­gres­sively. Even a bug trig­gered by a rare con­di­tion was bound to hit us even­tu­ally.

This was an ex­cit­ing mo­ment. After months of con­fu­sion and un­cer­tainty, we fi­nally had a plau­si­ble the­ory for why the cor­rup­tion was oc­cur­ring, and a fix we could de­ploy to pre­vent it.

The SQLite de­vel­op­ers re­leased the fix as SQLite 3.52.0, and we pre­pared to de­ploy it as soon as it was avail­able.

Fixed, with a false alarm

We rolled out SQLite 3.52.0 care­fully—first to a few ca­nary shards, then, when we saw it run­ning smoothly, we de­ployed it to the rest of the con­trol plane.

Our backup mon­i­tor promptly turned red, and re­ported cor­rup­tion in 13 dif­fer­ent data­bases. This was ex­tremely alarm­ing, but we fol­lowed our re­cov­ery pro­ce­dures to fix all the sup­posed cor­rup­tion, and every­thing was happy. It turned out these data­bases had not suf­fered real cor­rup­tion, but were sub­ject to a sec­ond prob­lem in the ver­sion of SQLite.

We shared our er­rors with the SQLite de­vel­op­ers, which un­cov­ered a bug in SQLite re­lated to stale ex­pres­sion in­dexes. If you cre­ate an in­dex on a com­puted value, and then the com­pu­ta­tion changes, the in­dex will con­tain mis­matched val­ues, which gets re­ported as cor­rup­tion by PRAGMA in­tegri­ty_check.

In our case, we were stor­ing some high-pre­ci­sion time­stamps as text, con­vert­ing them to a float­ing-point num­ber in a VIRTUAL gen­er­ated col­umn, and the SQLite 3.52.0 re­lease that fixed our data race also made an op­ti­mi­sa­tion that sub­tly changed the round­ing be­hav­iour for text-to-float­ing-point con­ver­sions. Our ca­nary shards did­n’t have any time­stamps that trig­gered the changed round­ing be­hav­iour, so we missed this in our phased roll­out.

Because this change caused false cor­rup­tion warn­ings, the SQLite de­vel­op­ers with­drew the 3.52.0 re­lease and in­stead pub­lished 3.51.3, which only con­tained a fix for the WAL-Reset bug.

We fixed the is­sue on our side by re­duc­ing the pre­ci­sion of our time­stamps to in­te­ger sec­onds; text-to-in­te­ger con­ver­sions are un­am­bigu­ous. Meanwhile, the SQLite de­vel­op­ers cre­ated an au­to­mated, self-heal­ing in­dex fea­ture in 3.53.0, which pre­vents the stale ex­pres­sion in­dex prob­lem.

Party time!

With the fix rolled out to our en­tire con­trol plane, we were ready to de­clare vic­tory, but we were still cau­tious. An ab­sence of cor­rup­tion in­ci­dents does­n’t mean things are fixed—we’d al­ready had one six-week pe­riod of de­cep­tive calm.

We wanted pos­i­tive proof that this data race was ac­tively oc­cur­ring in our pro­duc­tion en­vi­ron­ment. Now that we un­der­stood the cause of the bug—a col­li­sion be­tween a write trans­ac­tion and a WAL-reset—we patched our SQLite dri­ver to log a warn­ing when these two op­er­a­tions over­lap. If the warn­ing fired but the data­base re­mained un­cor­rupted, we’d know the fix had saved us from a po­ten­tial cor­rup­tion in­ci­dent.

We de­ployed the warn­ing, and we waited. And we waited. And waited. And waited. As weeks slipped by, we be­gan to won­der why we did­n’t see it. Was the warn­ing bro­ken? Was our the­ory wrong? Was the true bug still lurk­ing in the dark­ness?

Then, two months later, the alert we were wait­ing for fi­nally fired:

This alert proved that the pre­cise con­di­tions for the WAL-Reset bug do oc­cur in our pro­duc­tion en­vi­ron­ment, which means it was the likely cul­prit for our six months of shaky up­time.

Since that weirdly joy­ous alert fired, we’ve run for an­other four months with­out any data­base in­ci­dents, as of this writ­ing. Finally, we could breathe a sigh of re­lief.

Off the well-trod­den path

Nobody wanted us to spend six months look­ing for bugs in SQLite. This was an im­mensely frus­trat­ing ex­pe­ri­ence for both our cus­tomers and staff, and we’re all glad to put this in­sta­bil­ity be­hind us.

This in­ves­ti­ga­tion is a use­ful re­minder: run­ning bor­ing tech­nol­ogy in a non-stan­dard way is a risk. The com­mon paths and stan­dard con­fig­u­ra­tions are in­cred­i­bly well-tested and re­li­able. Most peo­ple use SQLite in a stan­dard con­fig­u­ra­tion and never face this sort of is­sue. Everything we were do­ing was a pub­lic, doc­u­mented, sup­ported con­fig­u­ra­tion—but by tak­ing man­ual con­trol of the check­point­ing process and run­ning at our own ag­gres­sive pace, we stepped off the well-trod­den op­er­a­tional path.

Resolving these in­ci­dents was a mas­sive, cross-func­tional ef­fort in­volv­ing dozens of peo­ple—in­clud­ing Tailscale’s en­gi­neer­ing and sup­port teams, and the core main­tain­ers of SQLite. It is to all of their credit that the im­pact of these in­ci­dents was not much worse.

We know that re­peated down­time erodes trust, no mat­ter how many peo­ple are af­fected, and we’re grate­ful to our cus­tomers for their pa­tience and sup­port while we chased this down.

Frustrating as this pe­riod was, we’re left in a stronger po­si­tion than we were be­fore. The long-stand­ing bug in SQLite has been patched, and we fixed dozens of other in­ci­den­tal is­sues that we spot­ted while look­ing for it. We funded the open-source SQLite VFS shim that helped iso­late the race con­di­tion al­most im­me­di­ately, and will help track down sim­i­lar bugs in the fu­ture. Finally, we’ve re­fined our data­base backup and re­cov­ery processes, and live-tested them over a dozen times.

Hopefully there won’t be an­other data­base in­ci­dent like this—but if there is, we’ll be ready.

Introducing Muse Glimmer: An Open Agentic Model That Runs on Your Device

research.meta.ai

Today, we’re in­tro­duc­ing Muse Glimmer, the next model from Meta Superintelligence Labs, and open sourc­ing the model weights un­der a per­mis­sive Apache 2.0 li­cense.

Muse Glimmer is a 30-billion-parameter model op­ti­mized for al­ways-on lo­cal agent work­flows. It’s small enough to run on a Mac or PC with a sin­gle con­sumer GPU, en­abling use cases that range from lo­cal agents and func­tion call­ing, to lo­cal cod­ing, and LLM-as-a-judge eval­u­a­tion. Muse Glimmer de­liv­ers strong per­for­mance on key agen­tic use cases and bench­marks com­pared with lead­ing mod­els in its size cat­e­gory.

Foundation mod­els have achieved re­mark­able ca­pa­bil­i­ties across rea­son­ing, code gen­er­a­tion, and tool use — yet most de­ploy­ments still de­pend on cloud in­fra­struc­ture and net­work ac­cess. Running mod­els lo­cally en­ables you to use AI any­where, any­time, with or with­out an in­ter­net con­nec­tion. This is in­creas­ingly vi­able: the open source com­mu­nity has shown that smaller mod­els, when trained ef­fec­tively, can ap­proach fron­tier-level per­for­mance on tar­geted tasks. Muse Glimmer is op­ti­mized for these lo­cal use cases.

Keeping with our long tra­di­tion of shar­ing fun­da­men­tal AI re­search, we’re re­leas­ing Muse Glimmer open weights to­day on Hugging Face, along with de­vel­oper doc­u­men­ta­tion to help you start build­ing and run­ning your own agents. Muse Glimmer is built to work with the tools de­vel­op­ers al­ready use. Optimized in­te­gra­tions on llama.cpp, MLX, and ExecuTorch will land in the com­ing days, so you can go from down­load to work­ing agent in min­utes.

How We Trained Muse Glimmer

An agent that man­ages your sched­ule, drafts your mes­sages, or­ga­nizes your files, and learns how you work needs deep ac­cess to per­sonal con­text. It also needs sev­eral ca­pa­bil­i­ties work­ing in con­cert: long-hori­zon ex­e­cu­tion, pre­cise tool call­ing, mul­ti­modal un­der­stand­ing, long-con­text mem­ory, and in­struc­tion fol­low­ing.

We de­signed Muse Glimmer to bal­ance ca­pa­bil­ity against the mem­ory and com­pute con­straints of lo­cal hard­ware. This re­quired a com­pact ar­chi­tec­ture, a novel dis­til­la­tion recipe that trans­fers agen­tic rea­son­ing from a much larger teacher model, and in­fer­ence op­ti­miza­tions — in­clud­ing quan­ti­za­tion — to meet la­tency ex­pec­ta­tions. We achieved this in the fol­low­ing phases:

Pre-Training. We trained Muse Glimmer on Muse Spark’s out­puts us­ing logit dis­til­la­tion, lever­ag­ing a sim­i­lar data mix as the teacher.

Mid-Training. We trained the model on longer-con­text, more agent-heavy data with richer rea­son­ing traces, along­side or­ganic data.

Post-Training. We com­bined su­per­vised fine-tun­ing with a mix of on-pol­icy dis­til­la­tion and re­in­force­ment learn­ing across gen­eral, rea­son­ing, cod­ing, and agen­tic do­mains.

Muse Glimmer was eval­u­ated un­der the stan­dards set out in Meta’s Advanced AI Scaling Framework and as­sessed for open-weight re­lease across all rel­e­vant cat­e­gories.

Built for Agents: What Muse Glimmer Can Do

Building ef­fec­tive agents re­quires key ca­pa­bil­i­ties work­ing to­gether to achieve the user’s goals. Muse Glimmer is trained and eval­u­ated across each of the fol­low­ing:

End-to-end Agentic Task Completion. Muse Glimmer achieves strong suc­cess rates on full-task bench­marks in­clud­ing DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which mea­sure its abil­ity to work within scaf­folds, write and de­bug code, and re­solve multi-turn re­quests from start to fin­ish.

Reliable Tool Use. The model han­dles a wide range of func­tion calls, in­vok­ing tools with pre­cise schemas through­out ex­tended work­flows.

Multi-Step Reasoning. Muse Glimmer chains rea­son­ing over long hori­zons, sus­tain­ing co­her­ent plans across com­plex, ex­tended work­flows.

Failure Recovery. When a tool call fails or re­turns an un­ex­pected re­sult, the model is trained to di­ag­nose the er­ror and retry rather than halt.

Multimodal Input and Reasoning. Through a ded­i­cated per­cep­tion en­coder, the model ac­cepts in­ter­leaved text and im­ages. This en­ables agents to in­ter­pret screen­shots, charts, and doc­u­ments along­side con­ver­sa­tion.

Scaffold Compatibility. Muse Glimmer works across OpenClaw and other agen­tic or­ches­tra­tion pat­terns.

Controllable Effort. Muse Glimmer sup­ports dif­fer­ent rea­son­ing strengths to se­lect the right bal­ance be­tween qual­ity and speed.

Multilingual. Muse Glimmer is trained on data from more than 100 lan­guages.

Performance

We eval­u­ated Muse Glimmer across a broad range of bench­marks to as­sess the di­verse ca­pa­bil­i­ties re­quired for ef­fec­tive au­tonomous agent be­hav­ior. Compared with Gemma4 – 31B and Qwen3.6 – 27B, Muse Glimmer per­forms strongly for its size class on sev­eral widely used LLM bench­marks.

For more de­tail about our eval­u­a­tions, see our re­port.

Optimized for Local Deployments

A lo­cal agent is truly use­ful if it’s fast enough to feel re­spon­sive. An agent that takes min­utes to re­ply or plan its next step breaks the flow of real work. We ap­plied two op­ti­miza­tions to make Muse Glimmer run at prac­ti­cal speeds on con­sumer hard­ware with­out sac­ri­fic­ing qual­ity.

Fitting the Model on Your Device.

At full pre­ci­sion, a 30-billion pa­ra­me­ter model would re­quire over 55 GB of mem­ory — far more than any con­sumer GPU of­fers. We use quan­ti­za­tion tech­niques to com­press the mod­el’s weights to ap­prox­i­mately 4-bit pre­ci­sion, shrink­ing the lan­guage model to un­der 20 GB. This leaves enough head­room for the mod­el’s work­ing mem­ory (its KV cache”), the per­cep­tion en­coder for im­age un­der­stand­ing, and the spec­u­la­tive de­cod­ing drafter to run si­mul­ta­ne­ously within a 24 GB or 32 GB en­ve­lope. We val­i­dated that this com­pres­sion in­tro­duces min­i­mal to no degra­da­tion on agen­tic tasks.

Faster Generation Through Speculative Decoding.

Language mod­els nor­mally gen­er­ate text one to­ken at a time, which can feel slow dur­ing long rea­son­ing chains or multi-step tool calls. Muse Glimmer ships with a light­weight drafter” model based on DFlash — a small com­pan­ion net­work that pro­poses en­tire blocks of to­kens at once. The main model then ver­i­fies these pro­pos­als in par­al­lel, ac­cept­ing cor­rect to­kens and cor­rect­ing wrong ones. This tech­nique lets Muse Glimmer gen­er­ate text sig­nif­i­cantly faster than stan­dard to­ken-by-to­ken gen­er­a­tion while pro­duc­ing iden­ti­cal out­put qual­ity. We pro­vide quan­tized drafter ver­sions to in­cur a smaller mem­ory over­head in the re­lease.

The Result:

We mea­sure the speed of our K-Quant-17GB model along­side the quan­tized DFlash drafter on MacBook M4-Max, M5-Max and on a RTX-5090. The model is fast enough for fluid con­ver­sa­tion and real-time agent in­ter­ac­tion, all run­ning en­tirely on your de­vice.

Get Started With Muse Glimmer Today

Muse Glimmer is avail­able now, and you can down­load the weights on Hugging Face. In the com­ing days, run it lo­cally through part­ners like Ollama, LM Studio, and Unsloth, de­ploy it with edge frame­works in­clud­ing llama.cpp, ExecuTorch, and MLX, serve it at scale with vLLM and SGLang, or get started quickly through part­ners like Together AI, Fireworks AI, and OpenRouter. You can even cus­tomize it for your use case by lever­ag­ing PyTorch’s TorchTitan train­ing fea­ture to tune the model fur­ther.

We’re also work­ing with our part­ners in­clud­ing AMD, Arm, Dell, Intel, and NVIDIA to op­ti­mize per­for­mance across de­vices. In ad­di­tion, we’re re­leas­ing doc­u­men­ta­tion so de­vel­op­ers have the re­sources they need to get started and build re­spon­si­bly with Muse Glimmer. This in­cludes guid­ance on set­ting up cus­tom scaf­folds, so it’s even eas­ier to start build­ing and de­ploy­ing per­sonal agents on day one. You can learn more and find re­sources to build on Meta’s AI Developer Center.

This work builds on Meta’s long track record of open AI re­search, ex­tend­ing it into agen­tic AI and giv­ing de­vel­op­ers ac­cess to lo­cal agen­tic ca­pa­bil­i­ties. As al­ways, we wel­come feed­back from the com­mu­nity and can’t wait to see what de­vel­op­ers build with this open weights model.

Download the Model on Hugging Face Developer Documentation

Firefox is now the last major browser that still supports uBlock Origin

www.pcworld.com

When you pur­chase through links in our ar­ti­cles, we may earn a small com­mis­sion. This does­n’t af­fect our ed­i­to­r­ial in­de­pen­dence.

News

Aug 13, 2026

Firefox re­cently an­nounced via Bluesky post: Our sup­port for uBlock Origin is­n’t go­ing any­where.” The mo­ment comes in re­sponse to news that Microsoft Edge is soon go­ing to lock out uBlock Origin and other ad-block­ing ex­ten­sions that run on Manifest V2 ar­chi­tec­ture.

Once Microsoft Edge moves to Manifest V3, ad-block­ing ex­ten­sions won’t have ac­cess to the func­tions needed to prop­erly iden­tify and block ads that oc­cur while brows­ing web­sites and watch­ing videos.

Microsoft’s move is­n’t sur­pris­ing, as Edge is based on Chromium, the open-source browser en­gine that pow­ers most web browsers to­day, in­clud­ing Opera, Brave, Vivaldi, and Samsung Browser. Google ini­ti­ated the mi­gra­tion from Manifest V2 to V3 in Chrome/Chromium, and Microsoft Edge is now fol­low­ing Google’s lead.

But Firefox is one of the few web browsers re­main­ing that is­n’t based on Chromium, and it’s now the only ma­jor browser to still sup­port uBlock Origin. Neither Safari nor DuckDuckGo—the two other ma­jor non-Chromium browsers out there—sup­port uBlock Origin.

For die-hard uBlock Origin fans, Firefox ap­pears to be the only browser left with­out com­pro­mises. With any other browser, you’ll need to set­tle for uBlock Origin Lite (with fewer fea­tures and less ad-block­ing suc­cess) or what­ever built-in ad-block­ing fea­ture comes with the browser.

This ar­ti­cle orig­i­nally ap­peared on our sis­ter pub­li­ca­tion PC för Alla and was trans­lated and lo­cal­ized from Swedish.

z.ai

Client Challenge

www.lemonde.fr

A re­quired part of this site could­n’t load. This may be due to a browser ex­ten­sion, net­work is­sues, or browser set­tings. Please check your con­nec­tion, dis­able any ad block­ers, or try us­ing a dif­fer­ent browser.

DeepSeek V4 Pro 0813 - API Pricing & Benchmarks

openrouter.ai

Not avail­able in this work­space

Gemini 3.7 Flash

ai.google.dev

Gemini 3.7 Flash is the next it­er­a­tion in the Gemini 3 se­ries of highly-ca­pa­ble, na­tively mul­ti­modal, rea­son­ing mod­els.

Documentation

Visit the Latest model page for full cov­er­age of fea­tures and ca­pa­bil­i­ties.

gem­ini-3.7-flash

Inputs

Text, Image, Video, Audio, and PDF

Output

Text

Input to­ken limit

1,048,576

Output to­ken limit

65,536

Audio gen­er­a­tion

Not sup­ported

Caching

Supported

Code ex­e­cu­tion

Supported

Computer use

Supported (Preview)

File search

Supported

Function call­ing

Supported

Grounding with Google Maps

Supported

Image gen­er­a­tion

Not sup­ported

Live API

Not sup­ported

Search ground­ing

Supported

Structured out­puts

Supported

Thinking

Supported (low, medium, high)

Note: min­i­mal is not sup­ported and re­turns an er­ror.

URL con­text

Supported

Batch API

Supported

Flex in­fer­ence

Supported

Priority in­fer­ence

Supported

Stable: gem­ini-3.7-flash

Except as oth­er­wise noted, the con­tent of this page is li­censed un­der the Creative Commons Attribution 4.0 License, and code sam­ples are li­censed un­der the Apache 2.0 License. For de­tails, see the Google Developers Site Policies. Java is a reg­is­tered trade­mark of Oracle and/​or its af­fil­i­ates.

Last up­dated 2026 – 08-13 UTC.

Introducing Gemini 3.7 Flash

blog.google

Aug 13, 2026

|

Our most in­tel­li­gent work­horse model yet for cod­ing and agents.

Your browser does not sup­port the au­dio el­e­ment.

Listen to ar­ti­cle

[[duration]] min­utes

This con­tent is gen­er­ated by Google AI. Generative AI is ex­per­i­men­tal

Today, we’re build­ing on the progress of our widely used Flash se­ries by in­tro­duc­ing Gemini 3.7 Flash, our most in­tel­li­gent work­horse model yet for cod­ing and agents.

This re­lease comes just three weeks af­ter Gemini 3.6 Flash, and is a di­rect re­sult of de­vel­oper feed­back and al­go­rith­mic in­no­va­tions that we look for­ward to bring­ing to fu­ture mod­els. 3.7 Flash de­liv­ers sub­stan­tial im­prove­ments across soft­ware en­gi­neer­ing, knowl­edge work, and web de­vel­op­ment work­flows — with an in­tro­duc­tory price of half the orig­i­nal 3.6 Flash cost per mil­lion to­kens.

Better in­tel­li­gence for com­plex work­flows

3.7 Flash shows strong gains over 3.6 Flash in cod­ing tasks like de­bug­ging and is­sue res­o­lu­tion. It also achieves higher first-pass code ac­cu­racy and has im­proved per­for­mance in gen­er­at­ing pro­duc­tion-ready code as seen in FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%).

In web de­vel­op­ment, 3.7 Flash gen­er­ates more func­tional lay­outs and fea­ture-com­plete apps in fewer prompts. For UI gen­er­a­tion, the model shows high de­sign ad­her­ence and par­ity based on a ref­er­ence in­put, whether it’s a screen­shot, an im­age, or a full de­sign sys­tem. It out­per­forms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowl­edge-dense fields like fi­nance, law, and bio­sciences, 3.7 Flash de­liv­ers im­proved rea­son­ing and ac­cu­racy. It sig­nif­i­cantly out­per­forms 3.6 Flash on the GDP.pdf bench­mark (34.0% vs 22.0%), an eval for test­ing a mod­el’s abil­ity to process com­plex doc­u­ments. It also sur­passes 3.6 Flash in AutomationBench, demon­strat­ing it can more ef­fec­tively com­plete real-world busi­ness work­flows (30.4% vs 17.0%).

Better de­vel­oper ex­pe­ri­ence and price

Gemini 3.7 Flash de­liv­ers a no­tice­ably im­proved de­vel­oper ex­pe­ri­ence over 3.6 Flash. It bet­ter adapts to road­blocks, clar­i­fies in­tent when needed, and fol­lows in­struc­tions with greater fi­delity. It thinks more dili­gently, putting in more ef­fort into multi-step plan­ning and tool calls. A more dis­ci­plined ex­e­cu­tion means less man­ual over­sight and fewer re­tries across en­gi­neer­ing work­flows.

3.7 Flash is avail­able through the end of the year at an in­tro­duc­tory price

1

of $0.75/1M in­put to­kens and $3.75/1M out­put to­kens. This price com­bined with the en­hanced model per­for­mance en­ables de­vel­op­ers and cus­tomers to scale pro­duc­tion-ready agents cost ef­fec­tively.

Early cus­tomer feed­back is high­light­ing 3.7 Flash’s per­for­mance and pre­ci­sion, achiev­ing re­sults that are sig­nif­i­cantly bet­ter than 3.6 Flash at a low cost.

Improving Gemini Spark with 3.7 Flash

Gemini Spark, avail­able to Google AI Pro and Ultra sub­scribers in over 160 coun­tries, will be us­ing Gemini 3.7 Flash start­ing to­day. We launched Spark at I/O as your per­sonal AI agent that runs 24/7, tak­ing ac­tion on your be­half while un­der your di­rec­tion. This model up­date makes Spark more ef­fi­cient for knowl­edge work with im­proved tool use for Google Workspace apps, de­liv­er­ing im­proved ac­cu­racy and out­put qual­ity for com­plex, multi-skill work­flows.

With 3.7 Flash, Gemini Spark can turn ideas into ac­tion more ef­fi­ciently by con­sol­i­dat­ing files, draft­ing emails, and up­dat­ing sta­tus doc­u­ments.

Built with safety in mind

We con­tin­u­ally work to im­prove the cov­er­age and ro­bust­ness of Frontier Safety safe­guards. Gemini 3.7 Flash is ship­ping with up­dated safe­guards against mis­use in the do­mains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cy­ber of­fense, while en­abling ben­e­fi­cial use cases, in ac­cor­dance with our ap­proach to biore­silience and our cy­ber pro­gram.

For more in­for­ma­tion, see the 3.7 Flash model card.

Try it to­day

Developers: Explore agent-first work­flows in Google Antigravity or start build­ing to­day in the Gemini API via Google AI Studio and Android Studio. Get started with our de­vel­oper guide.

Enterprises: Access 3.7 Flash in Gemini Enterprise Agent Platform and the Gemini Enterprise app.

Individuals: Available via Spark, your 24/7 per­sonal agent in the Gemini app for Google AI Pro and Ultra sub­scribers in sup­ported coun­tries.

Detailed bench­marks

Get the lat­est news from Google in your in­box

Sign up for our newslet­ters with prod­uct up­dates, event in­for­ma­tion, spe­cial of­fers, and more.

Your in­for­ma­tion will be used in ac­cor­dance with Google’s pri­vacy pol­icy. You may opt out at any time.

AI is removing the middle class of software engineering

blog.florianherrengt.com

It’s 2020. You’re the most se­nior per­son on your team, in charge of code qual­ity and ar­chi­tec­ture. You’ve set up good en­gi­neer­ing prac­tices, you thor­oughly re­view PRs from peo­ple who are less ex­pe­ri­enced than you and work hard to main­tain a healthy code­base.

Then at some point, you go on hol­i­day. When you come back, the code­base is a mess. Everyone merged each oth­er’s PRs with­out re­ally pay­ing much at­ten­tion, some­one added a bunch of new ta­bles to the data­base to de­nor­malise it be­cause it was eas­ier and they added server­less or Kafka to the stack with­out any solid ev­i­dence that they needed ei­ther.

It’s okay. You can fix this.

Fast for­ward to 2026. You haven’t been on hol­i­day. It’s just a nor­mal Monday morn­ing. You make your­self a nice cof­fee, open your com­puter and find your­self with 7 PRs to re­view. You open the first one: +24506 – 3938 lines, ac­com­pa­nied by some AI-generated de­scrip­tion of what they’re sup­posed to do. Somehow, your team has made more changes since Friday than they used to make while you were away for a few weeks.

AI re­moved the speed limit

AI makes pro­jects with weak en­gi­neer­ing cul­ture fail much faster.

There used to be a time when peo­ple sat down and talked about how they’d do some­thing. Now they can just prompt an agent for a few hours and open a PR.

The most tragic as­pect of this way of work­ing is that, to the un­trained eye, it works.

If you pull the branch and test it, you’ll prob­a­bly get some­thing some­what func­tional. So what do they do? They keep go­ing. Again and again. Until the pro­ject reaches a point where no one knows how any­thing works.

Just like some­one buy­ing a new lux­ury car on a credit card. You don’t see the debt. You just see the car that looks great.

But then users start to re­port a weird bug. It’s the 4th time your team has been try­ing to fix it. I mean… ask­ing AI to fix it. Unfortunately, it seems like not even Fable can fig­ure it out.

You go talk to the per­son who worked on this fea­ture.

So where does the data come from?”

Hmm… ac­tu­ally I don’t know. Let me ask Claude.”

You sit next to each other watch­ing an end­less wall of text ap­pear on the screen. Neither of you has any idea whether any of it is true but Claude seems very con­fi­dent.

Let’s just turn on ul­tra­code and ask it to dou­ble-check?”

This one will take a while. You start talk­ing about the lat­est drama on X.

You fi­nally get an an­swer back.

Does this make any sense to you?”

I’m not sure.”

Didn’t you build this like… last week?”

Silence.

This pro­ject has be­come so con­vo­luted, with so many lay­ers and ser­vices, that no one on your team could pos­si­bly start to un­der­stand what’s go­ing on.

So, what do you do?

Fixing it would re­quire such a colos­sal amount of work that it would be im­pos­si­ble to even start jus­ti­fy­ing it to any­one in man­age­ment.

And what are you even think­ing about? It would end up in the ex­act same state again in just a few months any­way.

Let’s just ask Claude to fix it.”

Okay. I’ll cre­ate a loop and goal so it does­n’t stop un­til it’s checked that every­thing works.”

Sounds good”

Actually, I ran out of Fable us­age for to­day so I’ll run it to­mor­row”

You grab an­other cof­fee and walk back to your com­puter. You now have 13 PRs left to re­view. You see some­thing you don’t quite un­der­stand, so you mes­sage the per­son who wrote it.

Why are we do­ing this here?”

They send you a link. It’s a Claude con­ver­sa­tion.

Somewhere in that con­ver­sa­tion, buried be­tween Claude con­fi­dently rec­om­mend­ing one ar­chi­tec­ture, apol­o­gis­ing, chang­ing its mind, your coworker ask­ing it to re­con­sider again and an­other 15 rounds of changes, is ap­par­ently the de­sign de­ci­sion be­hind this code.

Which part should I read?”

Probably all of it.”

Does this sound fa­mil­iar?

Whenever I talk about this, some­one even­tu­ally tells me that no­body ever fully un­der­stood large sys­tems any­way. It’s true.

You were never ex­pected to un­der­stand every ser­vice and every data­base. But at least some­one did and would ex­plain it to you.

Now they ask an LLM be­cause they don’t ac­tu­ally know them­selves.

You can’t af­ford bad en­gi­neers any­more

In every team, there are com­pe­tent peo­ple who make the pro­ject pos­si­ble. There are also peo­ple who es­sen­tially make it harder for every­one else. And now any­one can pro­duce more code in a day than they used to in a year.

In the story above, every­one is fail­ing:

The en­gi­neer open­ing a 25,000-line PR should have stopped the agent long be­fore it got there. They should have un­der­stood what it was do­ing, bro­ken the work into smaller pieces and ques­tioned every new ab­strac­tion it in­tro­duced.

The per­son re­view­ing it should have re­fused to re­view some­thing that large in­stead of giv­ing in.

The per­son adding Kafka should have been able to ex­plain ex­actly why it was needed.

The per­son who built the fea­ture should have been able to ex­plain where the data came from with­out send­ing a link to a Claude con­ver­sa­tion.

But what’s the prob­lem then? Just use AI to fix it. Well, it’s not that easy…

Before any­one jumps on this, none of this means tech­ni­cal debt is al­ways bad. The im­por­tant part is that you know it’s a short­cut.

Anyway, re­vert­ing a bad de­ci­sion is hard. Very hard.

For ex­am­ple, how long would it take an LLM to add a bunch of ta­bles and columns to the data­base? 10 min­utes?

But once you start stor­ing data there, you can’t just re­move them. You have to come up with a mi­gra­tion plan, make sure you don’t dis­rupt the sys­tem be­cause peo­ple are pay­ing to use this every day. You have to think about what you’ll do if the mi­gra­tion fails. Make sure you don’t end up with or­phaned for­eign keys. It’s just so much harder to fix. Even with the best model you can get.

And while you’re fix­ing it, more PRs keep com­ing in. More code, more ab­strac­tions, more de­ci­sions. A per­son can gen­er­ate 20,000 lines of code in an af­ter­noon, but you still have to sit there and un­der­stand what those lines ac­tu­ally do.

By the time you’ve un­tan­gled one bad de­ci­sion, five more have been merged.

The new AI econ­omy

Of course, bad en­gi­neers were al­ways a li­a­bil­ity.

It has been like this for decades, well be­fore OpenAI or Anthropic ex­isted. Bad de­ci­sions com­pounded, un­nec­es­sary com­plex­ity ac­cu­mu­lated and teams ended up main­tain­ing sys­tems no­body re­ally un­der­stood.

The dif­fer­ence is that there used to be a limit to how fast you could do it.

Today, im­ple­men­ta­tion is cheap. You are paid to make good de­ci­sions. To build soft­ware that will scale while man­ag­ing com­plex­ity.

Ask your­self why com­pa­nies are pay­ing six-fig­ure salaries for en­gi­neers in London or San Francisco in the first place.

If all they needed was some­one who could turn a spec­i­fi­ca­tion into work­ing code, why were they pay­ing that much when they could al­ready get it done cheaply else­where?

Why are the tech com­pa­nies claim­ing that software is solved” still pay­ing top salaries to at­tract the best peo­ple they can?

My bet is that AI pushes salaries fur­ther apart. To be em­ploy­able, there’s a bar you have to clear and that bar is what­ever the cur­rent best model du jour can do.

Good en­gi­neers have be­come more valu­able be­cause AI lets them move much faster. They don’t need as many peo­ple around them just to do the im­ple­men­ta­tion work any­more.

At the same time, bad en­gi­neers have be­come much more ex­pen­sive to hire.

I wrote about this be­fore when I said the vibe coder ca­reer path is doomed.

You need to con­tribute be­yond what every­one al­ready gets by giv­ing an agent a prompt.

If you lack the judg­ment re­quired to eval­u­ate the LLMs rec­om­men­da­tion, ask­ing for more judg­ment does­n’t solve the prob­lem.

At some point, some­one still has to know what is go­ing on. And that’s the most valu­able per­son on the team.

The peo­ple who don’t will be­come much cheaper to hire or get re­placed en­tirely while the money gets fun­nelled to­wards an in­creas­ingly smaller num­ber of peo­ple who can ac­tu­ally be trusted.

I don’t think this is go­ing to be lim­ited to soft­ware en­gi­neer­ing ei­ther. I be­lieve the same thing is go­ing to hap­pen across most knowl­edge work. AI will make the best peo­ple much more pro­duc­tive and the bad ones al­most im­pos­si­ble to hire. Before, there was a good chance some­one would catch their bad de­ci­sions be­fore they went too far. Now they can make changes faster than any­one around them can re­al­is­ti­cally re­view or un­der­stand them.

Answers to the most com­mon ob­jec­tions

The dif­fer­ence is the speed.

It is the dif­fer­ence be­tween crash­ing at 30 km/​h and crash­ing at 200 km/​h. Before AI, a bad en­gi­neer would strug­gle to pro­duce code that even com­piled. When they did pro­duce some­thing, it took them a long time and the blast ra­dius was lim­ited. The dam­age was bounded by how fast a hu­man could type.

Now a bad en­gi­neer can pro­duce 10,000 lines of work­ing code be­fore lunch. The dam­age we can do in an af­ter­noon used to take them months. The speed at which bad de­ci­sions com­pound has changed com­pletely while the speed at which you can fix them has not.

Several peo­ple ar­gued that the real prob­lem is lack of process. If you had proper tests, CI, code re­view and ar­chi­tec­tural re­views, AI-generated slop would­n’t get through.

We had all of those things. None of them dis­ap­peared.

The prob­lem is that they were de­signed for a world where pro­duc­ing a mas­sive amount of change was im­pos­si­ble. Code re­view don’t work any­more when some­one opens 10 PRs a day with an AI-generated de­scrip­tion. Tests work when they cover the be­hav­iours you thought to test. They do not catch the be­hav­iours no­body thought to test.

How many times have you had a com­pletely green CI with full cov­er­age and still shipped a bug?

The dif­fi­culty of pro­duc­ing code was by it­self one of the lim­it­ing fac­tors.

If the peo­ple try­ing to un­der­stand changes and guard qual­ity are now the bot­tle­neck, you have three op­tions. Generate less, find a gen­uinely bet­ter way to val­i­date or ac­cept lower qual­ity.

Everytime I criti­sise AI some­one says I’m a Luddism. I’m re­fus­ing to adapt to the new world, cling­ing to the old ways.

I use AI heav­ily. I have said this re­peat­edly. I use it every day and I have no in­ter­est in go­ing back to writ­ing every­thing by hand.

The point is that we have made pro­duc­ing large changes ex­tremely cheap and fast while un­der­stand­ing those changes is still slow, dif­fi­cult work. We have no short­cut for build­ing a cor­rect men­tal model of what a change does.

Maybe one day we will find one. As of to­day, I do not think we have one.

You can be a heavy AI user and still recog­nise that there are se­ri­ous prob­lems with how it is be­ing used.

Someone gen­er­ates 10 PRs in a day, the num­bers look in­cred­i­ble, surely this per­son is ten times more pro­duc­tive.

Not nec­es­sar­ily. The sup­posed 10x en­gi­neer may sim­ply be some­one steal­ing pro­duc­tiv­ity from every­one around them.

If I gen­er­ate 10 PRs in a day but three en­gi­neers now have to spend the next two days re­view­ing them, fig­ur­ing out what I changed, cor­rect­ing bad as­sump­tions, de­bug­ging re­gres­sions and ex­plain­ing why half of it needs to be re­done, I have not be­come 10x more pro­duc­tive. I have just moved the work onto other peo­ple.

Worse, I am con­sum­ing the time of the peo­ple who are usu­ally the hard­est to re­place and whose at­ten­tion is al­ready scarce.

PR count, lines changed and fea­tures completed” are ter­ri­ble mea­sures of pro­duc­tiv­ity. You can make your own num­bers look in­cred­i­ble while re­duc­ing the through­put of the en­tire team.

Some peo­ple pushed back on this by say­ing the real prob­lem is bad or­gan­i­sa­tions with bro­ken in­cen­tives. It’s true. If your com­pany re­wards ticket count over qual­ity, the care­ful en­gi­neer looks like the bad em­ployee and the sloppy one gets pro­moted.

Someone pointed out that try­ing to hold the line on qual­ity can get you la­belled as toxic. Everyone else is ship­ping fast and you are the per­son say­ing wait”.

I would much rather have some­one on my team who ships less but whose work I can trust than some­one much faster whose changes leave me won­der­ing what prob­lems we’re go­ing to dis­cover later.

You need to be flex­i­ble and com­pro­mise when the busi­ness trade-off makes sense. But you also need a back­bone. If you think some­thing is go­ing to cause real prob­lems, bring­ing it up is part of the job.

And when pro­duc­tion breaks, and it will, I need the per­son who made the change to ac­tu­ally un­der­stand it well enough to help fix it. Not show up with no idea what is go­ing on.

We al­ready trust com­pil­ers to pro­duce ma­chine code we do not read. Why is this dif­fer­ent?

A com­piler takes code and trans­lates it into an­other rep­re­sen­ta­tion while pre­serv­ing its se­man­tics. The com­piler is not de­cid­ing what your sys­tem should do. It is de­ter­min­is­tic.

An LLM is mak­ing de­ci­sions. It is choos­ing ar­chi­tec­tures, pick­ing ab­strac­tions, de­cid­ing where to put things. When you ask Claude to build a fea­ture, it is not trans­lat­ing your in­tent into code. It is mak­ing dozens of de­sign de­ci­sions on your be­half.

If in five years I can give an agent a com­plete spec­i­fi­ca­tion and re­li­ably ver­ify the re­sult­ing code against it, then sure, re­view­ing code may be­come ob­so­lete and I would hap­pily stop do­ing it. We are not there yet.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.