10 interesting stories served every morning and every evening.

AI;DR (AI; Didn’t Read)

www.rickmanelius.com

I’m SUPER jeal­ous that I did­n’t think of this first…

Alas! Hat tip to se­clilc for tweet­ing this gem out two days ago.

lil c@se­clilc

AI;DR

(AI; did­n’t read)

4:13 PM · Aug 15, 2026 · 346K Views

83 Replies · 2.09K Reposts · 16.6K Likes

I’ve been think­ing about it ever since. Why? Because there is grow­ing grum­bling among every­one about AI writ­ing. And it’s not just oth­ers; it’s me! I am get­ting to the point where I phys­i­cally flinch (sometimes drop­ping my shoul­ders and hunch­ing, or hav­ing a slight eye twitch) when some­one I re­spect sends me un­fil­tered and unedited AI out­put.

Look, I get it. It’s Q3 2026, and we should ex­pect that every­one is uti­liz­ing AI at SOME point in their process (sourcing ideas, cre­at­ing out­lines, re­fin­ing prose, etc.).

However, I have a new pol­icy.

If you’re not both­ered enough to re­view and edit it…

…then I’m not go­ing to bother read­ing it.

Yes, there are cer­tain sit­u­a­tions in which we should ex­pect 100% AI-generated copy. Customer sup­port would be a per­fect ex­am­ple. We’re not look­ing for ar­ti­sanal did you make sure to re­set your phone” style di­a­logue.

But if you’re my col­league and we’re in a Slack dis­cus­sion and you post a wall of Claude out­put, then I’m afraid I re­ceived a dif­fer­ent mes­sage than you in­tended.

The same is true for peo­ple’s newslet­ters and so­cial con­tent. It’s your name on it; are you proud of the prose and weird AI-isms sprin­kled through­out it? If so, great. But I can ask Claude di­rectly if I wanted to.

TL;DR (too long; did­n’t read) was the so­lu­tion for so­cial me­dia.

AI;DR (AI; did­n’t read) is the so­lu­tion for AI slop.

May you em­brace this pol­icy your­self and seek out those will­ing to care enough to pri­or­i­tize a hu­man touch when they talk to you.

No posts

A Preview of DuckDB v2.0

duckdb.org

Mark Raasveldt and Hannes Mühleisen

2026 – 08-17

| 15 min

TL;DR: DuckDB v2.0 is com­ing this fall. In this post, we pre­view its head­line fea­tures: DuckDB as a server, trig­gers, the VARIANT type, asyn­chro­nous I/O, a new SQL parser, a new stor­age for­mat, and much more.

DuckDB v2.0 will be named Cyanoptera” af­ter the cin­na­mon teal (Anas cyanoptera), a strik­ingly red­dish-brown duck found in the west­ern Americas.

A ma­jor ver­sion bump is not some­thing we do lightly, and it is not just cer­e­mony: v2.0 ships a new SQL parser, a new de­fault stor­age for­mat, a re­worked C API, and a small num­ber of care­fully cho­sen break­ing changes. But above all, it is a fea­ture re­lease, built from over 10,000 com­mits since we re­leased v1.5 in March. Where last year was the year of the lake­house, this re­lease kicks off the year of DuckDB as a server. We pre­viewed many of these fea­tures in the State of the Duck” talk at DuckCon #7, if you pre­fer to watch in­stead of read.

DuckDB is mov­ing rather quickly, and we can only cover a small frac­tion of the changes here. Condensing all new fea­tures down to a short­list is al­ways a fight over what gets in, and yes, we know that what fol­lows is tech­ni­cally a lis­ti­cle (Ten Things Coming to DuckDB v2.0, Number Eight Will Shock You). We are not proud of the for­mat, but it works, so here it is, start­ing with the SQL-level fea­tures and work­ing down into the en­gine.

1. DuckDB as a Server: Quack and CONNECT

DuckDB has been an in-process data­base since day one. But peo­ple have asked us — very per­sis­tently — for a client/​server mode, and we have fi­nally caved. The quack ex­ten­sion im­ple­ments DuckDB’s na­tive pro­to­col for talk­ing to other DuckDBs. It was re­leased as a pre­view shortly be­fore DuckCon #7, grad­u­ates to sta­ble in v2.0, and it is a big part of where DuckDB is headed: any DuckDB process can serve its data­bases over the net­work, and any other DuckDB can at­tach to it and route queries there us­ing the new CONNECT state­ment. For ex­am­ple:

DuckDB server

CALL quack­_serve( to­ken = my_token’ );

quack:

DuckDB client

ATTACH quack:server.example.com’ AS qk (TOKEN my_token’);

CONNECT qk; SELECT count(*) FROM events; — ex­e­cutes on the server, — re­sults stream back DISCONNECT;

CONNECT is the suc­ces­sor to the re­mote.query($$…$$) workaround we showed when Quack was first re­vealed — we looked at that syn­tax and said: no, this can­not be it. And CONNECT is not lim­ited to Quack: it points your ses­sion at any re­mote data­base that sup­ports it, and the new re­mote push­down op­ti­mizer (#22914) ships SQL di­rectly to PostgreSQL and MySQL in­stead of pulling ta­bles over the wire:

CONNECT postgres://localhost/mydb’; SELECT count(*) FROM or­ders; — runs on the PostgreSQL server DISCONNECT;

If you have worked with an­a­lyt­i­cal sys­tems in the past, you may as­sume that DuckDB can­not han­dle trans­ac­tional work­loads. But DuckDB has been built as a trans­ac­tional, multi-con­nec­tion data­base with full MVCC and trans­ac­tion iso­la­tion since day one. Most users just never needed that in a sin­gle-user sce­nario. It turns out DuckDB han­dles trans­ac­tions well: it’s fast enough to com­pete with gen­eral-pur­pose data­bases like PostgreSQL on quite a few work­loads, and the client/​server pat­tern fi­nally lets that ma­chin­ery shine in multi-ten­ant, long-run­ning de­ploy­ments.

Running DuckDB long-term also comes with new chal­lenges, which is why v2.0 pushes on bet­ter met­rics, logs, and ob­serv­abil­ity (see, e.g., the met­rics layer re­work in #22799) that let you look at a DuckDB in­stance and see what it is ac­tu­ally do­ing. People even built stand­alone clients for the Quack pro­to­col within weeks of the pre­view. We thought we were ex­tend­ing DuckDB to talk to other DuckDBs; the world said no, no, no, and built their own clients. Who would have thought.

2. VARIANT Becomes a First-Class Citizen

The VARIANT type shipped in DuckDB v1.5, and the way to think about it is JSON on steroids. Basically, imag­ine if JSON were fast. Like JSON, a VARIANT col­umn can store dif­fer­ently-shaped data in every row. Unlike JSON, it is not a text for­mat: DuckDB au­to­mat­i­cally de­tects the com­mon struc­ture hid­den in your semi-struc­tured data and shreds” it, so it com­presses well in stor­age and ex­e­cutes fast in queries, all with­out you ever de­clar­ing a schema. This makes VARIANT a nat­ural fit for real-time log in­ges­tion, where streams of JSON-ish records share struc­ture but evolve over time.

In v2.0, this pipeline works end to end: shred­ded ex­e­cu­tion straight from stor­age (#20912), ex­trac­tion push­down into scans (#22478), shred­ded VARIANT read­ing and writ­ing for Parquet, and a fam­ily of vari­ant_* func­tions:

CREATE TABLE events (payload VARIANT); INSERT INTO events VALUES (‘{“user”: {“id”: 42, tags”: [“a”, b”]}}’::JSON::VARIANT);

SELECT vari­ant_­type(pay­load), vari­ant_keys(pay­load) FROM events;

SELECT * FROM events WHERE vari­ant_­con­tains(pay­load, {‘user’: {‘id’: 42}}::VARIANT);

Longer term, likely soon af­ter v2.0 (but don’t hold us to it), we plan to back the reg­u­lar JSON type with VARIANT, so ex­ist­ing JSON work­loads get all of these ben­e­fits with­out chang­ing a sin­gle query.

3. Triggers

Triggers have been a long-stand­ing fea­ture re­quest, and DuckDB v2.0 de­liv­ers them in full: BEFORE and AFTER trig­gers, FOR EACH ROW and FOR EACH STATEMENT, tran­si­tion ta­bles via REFERENCING OLD/NEW TABLE, mul­ti­ple trig­gers per event, RETURNING on trig­gered ta­bles, and DROP TRIGGER.

The clas­sic use case is au­dit ta­bles: some­thing hap­pens in the sys­tem, and a trig­ger records what changed. For ex­am­ple:

CREATE TABLE tar­get (id INTEGER, val INTEGER); CREATE TABLE au­dit (id INTEGER, old_­val INTEGER, new_­val INTEGER);

CREATE TRIGGER trg_au­dit AFTER UPDATE ON tar­get REFERENCING OLD TABLE AS o NEW TABLE AS n FOR EACH STATEMENT INSERT INTO au­dit SELECT n.id, o.val, n.val FROM o JOIN n ON o.id = n.id;

INSERT INTO tar­get VALUES (1, 10), (2, 20); UPDATE tar­get SET val = val * 10 WHERE id <= 2; SELECT * FROM au­dit;

Triggers fit nat­u­rally with long-run­ning DuckDB ser­vices, and we are also plan­ning to use them in­ter­nally to build sev­eral up­com­ing fea­tures. They are fully ex­posed at the SQL level too, so you can build your own cool stuff with them.

4. SQL Dialect Additions

As al­ways, DuckDB’s SQL di­alect keeps grow­ing. A few fa­vorites from this re­lease cy­cle:

With NEAREST joins (#24137), top-k sim­i­lar­ity search be­comes a join clause, handy for vec­tor and em­bed­ding work­loads:

SELECT q.user_id, t.prod­uc­t_id FROM users q INNER JOIN prod­ucts t APPROX NEAREST 2 BY SIMILARITY ar­ray_­co­sine_sim­i­lar­ity(q.em­bed­ding, t.em­bed­ding);

DML in­side CTEs (#21634, #21997, #24217) lets you use INSERT, UPDATE, DELETE, and COPY as pipeline steps:

WITH moved AS MATERIALIZED ( DELETE FROM stag­ing RETURNING * ) INSERT INTO archive SELECT * FROM moved;

Nested schemas (#23492, #24222) al­low schemas within schemas:

CREATE SCHEMA fi­nance; CREATE SCHEMA fi­nance.re­ports; CREATE TABLE fi­nance.re­ports.q3 (revenue DECIMAL);

The new vari­able syn­tax (#21194) lets you write $x any­where an ex­pres­sion is al­lowed, no more get­vari­able(…) ver­biage:

SET VARIABLE thresh­old = 100; SELECT * FROM or­ders WHERE amount > $threshold;

The JSON mu­ta­tion func­tions json_set, json_in­sert, json_re­place, and json_re­move (#23786) fi­nally let you mod­ify JSON doc­u­ments in place:

SELECT json_set(‘{“a”:1}‘, $.b’, 2’);

And re­cur­sive CTEs with USING KEY ag­gre­ga­tion (#19481) en­able it­er­a­tive al­go­rithms in pure SQL, backed by the rewrit­ten re­cur­sive CTE en­gine de­scribed be­low:

WITH RECURSIVE tbl(a, b) USING KEY (a, avg(b)) AS ( SELECT 1, 5 UNION SELECT a, b - 1 FROM tbl WHERE b > 0 ) TABLE tbl;

There is more: SQL-standard FETCH FIRST 2 ROWS ONLY (#23533), OVERLAY() (#22456), UNNEST in GROUP BY (#23644), and well-de­fined MERGE / UPDATEFROM se­man­tics for multi-matched rows (#24058).

5. Asynchronous I/O

Interacting with ob­ject stores like S3 is cen­tral to the DuckDB ex­pe­ri­ence: your data has to come from some­where, and it of­ten sits in ob­ject stor­age. DuckDB has long been able to read from ob­ject stores in par­al­lel, but syn­chro­nous ac­cess placed a limit on how fast this could go. DuckDB v2.0 in­tro­duces asyn­chro­nous I/O through­out the en­gine. We de­scribed the de­sign in de­tail in a ded­i­cated blog post.

Thanks to asyn­chro­nous ac­cess, the I/O layer now scales in­de­pen­dently from the query pro­cess­ing layer, which means far more par­al­lelism for re­mote reads and dra­mat­i­cally faster queries on net­work stor­age. Parquet sup­port came first (#23662), with CSV (#23961) and DuckDB’s own file for­mat (#24654) fol­low­ing, along with asyn­chro­nous Parquet writes (#23283) and new MMAP and DIRECT_IO modes (#22988). Local stor­age ben­e­fits a lit­tle too, but net­work stor­age is where you will see the big gains.

6. Faster Queries Across the Board

As with every re­lease, a lot of work went into mak­ing your ex­ist­ing queries faster with­out you do­ing any­thing. To pick some high­lights: par­tial ag­gre­gates are now pushed be­low joins (#22572) and re­dun­dant ag­gre­ga­tions are reused (#24543), the re­cur­sive CTE en­gine has been rewrit­ten (#22211), ag­gre­ga­tions now spill to disk when they out­grow mem­ory (#24499), and the Windows CLI got ap­prox­i­mately 2.2× faster at multi-threaded re­sult ma­te­ri­al­iza­tion (#24036).

How much faster can this get? Here is a mi­crobench­mark you can run on a lap­top: sin­gle-source reach­a­bil­ity over a graph with one mil­lion edges, writ­ten as a plain re­cur­sive CTE.

CREATE TABLE edges AS SELECT (range % 100_000)::INTEGER AS src, ((range * 13 + 7) % 100_000)::INTEGER AS dst FROM range(1_000_000);

WITH RECURSIVE reach­able(node) AS ( SELECT 0 UNION SELECT dst FROM edges, reach­able WHERE src = node ) SELECT count(*) FROM reach­able;

As you can see, DuckDB v2.0 is about 40× faster (!) for the same re­cur­sive query.

Row-group prun­ing has been mas­sively ex­panded: min-max in­dexes (zone maps) and Parquet Bloom fil­ters now skip data for structs, lists, dec­i­mals, UUIDs, IN fil­ters, and even func­tion pred­i­cates:

– these now prune row groups in­stead of scan­ning them: SELECT * FROM logs WHERE con­tains(mes­sage, ERROR); SELECT * FROM t WHERE sub­str(code, 1, 3) = NL-’; SELECT * FROM data/*.parquet’ WHERE id IN (1, 5, 9);

Query plan­ning also be­comes par­ti­tion-aware (#22336). Lakehouse for­mats (DuckLake, Iceberg and plain Hive-partitioned Parquet on S3) are all par­ti­tioned, and ex­ploit­ing that par­ti­tion­ing is of­ten the dif­fer­ence be­tween scan­ning a dataset and skip­ping most of it. In v2.0, the plan­ner and op­ti­mizer take full ad­van­tage of ex­ist­ing par­ti­tion­ing, and par­ti­tioned writes have been re­worked as well (#22225, #22620).

7. Storage Format v2.0

DuckDB v2.0 bumps the de­fault stor­age for­mat ver­sion to v2.0.0 (#22875). The head­line change is buffer-man­aged ART in­dexes (#21458, #23605): in­dexes are no longer pinned in mem­ory, which means large in­dexed ta­bles open in­stantly and their in­dexes are paged in on de­mand.

Column meta­data is now loaded lazily (#22333), so wide ta­bles open faster too. The DICT_FSST string com­pres­sion method is en­abled by de­fault (#23733), deletes are stored com­pactly (#24336), and the stor­age layer per­forms much stronger cor­rup­tion val­i­da­tion on read. In short: data­bases with big in­dexes and wide ta­bles open faster and use far less mem­ory.

8. A Brand New SQL Parser

DuckDB has fa­mously al­ways used a parser de­rived from PostgreSQL’s. We have de­cided that enough is enough: v2.0 ships our own mod­ern, ex­ten­si­ble PEG-based parser (#22194), an idea we first ex­plored in our 2024 post on run­time-ex­ten­si­ble parsers. This change ties into the ex­ten­sion ecosys­tem: ex­ten­sions can now hook into the gram­mar it­self, so ex­pect ex­ten­sions that ex­pose en­tirely new SQL syn­tax. It also brings bet­ter er­ror mes­sages with pre­cise source lo­ca­tions, and the first di­alect com­pat­i­bil­ity mode:

SET di­alec­t_­com­pat­i­bil­i­ty_­mode = spark’;

You should not ac­tu­ally no­tice any­thing from the parser swap as we de­signed it to be com­pat­i­ble with the old one. If you do no­tice, please file an is­sue.

9. Timezones, Calendars, and Collations Without ICU

Timezone-aware time­stamps, cal­en­dars, and col­la­tions in DuckDB have al­ways been pow­ered by the ICU li­brary. ICU is a fine li­brary, but we only ever used a small slice of it, while still car­ry­ing it around in every DuckDB dis­tri­b­u­tion. In v2.0, the ICU li­brary is gone en­tirely: the icu ex­ten­sion now im­ple­ments time­zones, cal­en­dars, and col­la­tions it­self (#24463, #24403), with the time­zone data built di­rectly from the IANA data­base and com­pressed down to around 45 kB. Everything keeps work­ing ex­actly as be­fore:

SELECT 2026 – 08-14 12:00:00’::TIMESTAMPTZ AT TIME ZONE Europe/Paris’; SELECT * FROM names ORDER BY name COLLATE de;

Besides be­ing much smaller and eas­ier to keep up to date, the new im­ple­men­ta­tion is also sim­ply faster. Here’s a quick mi­crobench­mark on a MacBook that con­verts 25 mil­lion time­stamps to a time­zone and fil­ters 5 mil­lion strings with a German col­la­tion:

10. Write Extensions Once, Host Them Yourself

Extensions are one of the best things about DuckDB, but to­day, most of them, in­clud­ing our own, build against the un­sta­ble C++ API. That means ex­ten­sion au­thors have to re-tar­get and re­build for every DuckDB re­lease, and com­mu­nity ex­ten­sions can silently dis­ap­pear when their au­thors stop keep­ing up. DuckDB v2.0 broad­ens the sta­ble C API far enough that ex­ten­sions can be writ­ten once, built once, pub­lished once, and keep work­ing, es­sen­tially un­til the end of time.

To make this sus­tain­able over the long run, the C API is now gen­er­ated from a de­clar­a­tive, ver­sioned spec­i­fi­ca­tion (#24135): every func­tion in duckdb.h, duck­d­b_ex­ten­sion.h, and the ex­ten­sion ABI is de­scribed in YAML in the api_spec/ di­rec­tory, with its full life­cy­cle on record, and CI ver­i­fies the com­mit­ted head­ers against the spec so API and ABI can no longer drift apart. The re­lease also brings uni­fied sym­bol ver­sion­ing (#24435), cus­tom al­lo­ca­tion han­dlers (#23945), and sta­tic link­ing of C API ex­ten­sions into your ap­pli­ca­tion (#22251).

So what does build­ing an ex­ten­sion against the sta­ble C API look like? Here is a com­plete ex­ten­sion: a sin­gle file that reg­is­ters a vec­tor­ized scalar func­tion, com­piled once against duck­d­b_ex­ten­sion.h.

#include duckdb_extension.h”

DUCKDB_EXTENSION_EXTERN

// a scalar func­tion that adds two BIGINTs, one vec­tor at a time sta­tic void AddNumbers(duckdb_function_info info, duck­d­b_­da­ta_chunk in­put, duck­d­b_vec­tor out­put) { idx_t count = duck­d­b_­da­ta_chunk_get_­size(in­put); in­t64_t *a = (int64_t *) duck­d­b_vec­tor_get_­data(duck­d­b_­da­ta_chunk_get_vec­tor(in­put, 0)); in­t64_t *b = (int64_t *) duck­d­b_vec­tor_get_­data(duck­d­b_­da­ta_chunk_get_vec­tor(in­put, 1)); in­t64_t *result = (int64_t *) duck­d­b_vec­tor_get_­data(out­put); for (idx_t row = 0; row < count; row++) { re­sult[row] = a[row] + b[row]; } }

DUCKDB_EXTENSION_ENTRYPOINT(duckdb_connection con, duck­d­b_ex­ten­sion_info info, duck­d­b_ex­ten­sion_ac­cess *access) { duck­d­b_s­calar_­func­tion f = duck­d­b_cre­ate_s­calar_­func­tion(); duck­d­b_s­calar_­func­tion_set_­name(f, add_numbers”); duck­d­b_­log­i­cal_­type big­int = duck­d­b_cre­ate_­log­i­cal_­type(DUCK­D­B_­TYPE­_BIG­INT); duck­d­b_s­calar_­func­tion_ad­d_­pa­ra­me­ter(f, big­int); duck­d­b_s­calar_­func­tion_ad­d_­pa­ra­me­ter(f, big­int); duck­d­b_s­calar_­func­tion_set_re­turn_­type(f, big­int); duck­d­b_de­stroy_­log­i­cal_­type(&big­int); duck­d­b_s­calar_­func­tion_set_­func­tion(f, AddNumbers); duck­d­b_reg­is­ter_s­calar_­func­tion(con, f); duck­d­b_de­stroy_s­calar_­func­tion(&f); re­turn true; }

LOAD ad­d_num­bers; SELECT ad­d_num­bers(40, 2);

For brevity, we skipped NULL han­dling here. See the de­mo_­capi ex­ten­sion for the full ver­sion.

For brevity, we skipped NULL han­dling here. See the de­mo_­capi ex­ten­sion for the full ver­sion.

The bi­nary this com­piles to keeps work­ing across DuckDB ver­sions. You do not need re-tar­get or re­build it every time a new DuckDB ver­sion comes out. And nowa­days, with all the AI tool­ing around, build­ing an ex­ten­sion has never been eas­ier.

So you have writ­ten your ex­ten­sion. But how should you dis­trib­ute it? Until now, DuckDB could only in­stall ex­ten­sions from the built-in repos­i­to­ries (core, core_nightly, com­mu­nity, …). In v2.0, you will be able to reg­is­ter your own trusted repos­i­to­ries (#24777, cur­rently work-in-progress), so an or­ga­ni­za­tion can host and sign its own ex­ten­sions and have them in­stall and load just like the built-in ones:

SET al­low_ex­ten­sion_repos­i­to­ries = allowed’; CREATE EXTENSION REPOSITORY my_repo FROM https://​ex­ten­sions.ex­am­ple.org; INSTALL my_ext FROM my_repo; LOAD my_repo/​my_ext;

A repos­i­tory is a name, a URL pre­fix, and one or more RSA pub­lic keys that are trusted to sign the ex­ten­sions served from it. The pre­fix can point at any­thing DuckDB can read: a lo­cal path, https, s3, you name it. At CREATE time, DuckDB fetches the repos­i­to­ry’s pub­lic keys and pins them into the repos­i­tory de­f­i­n­i­tion, print­ing each key’s SHA-256 fin­ger­print so you can com­pare it against one pub­lished out of band. If you would rather not trust the net­work at all, you can pass the key di­rectly:

CREATE EXTENSION REPOSITORY my_repo FROM s3://my-bucket/extensions’ USING PUBLIC KEY ––-BEGIN PUBLIC KEY––- …’;

Pinned repos­i­to­ries sur­vive restarts, sup­port key ro­ta­tion by trust­ing mul­ti­ple keys, and can be au­dited at any time through the duck­d­b_ex­ten­sion_repos­i­to­ries() table func­tion, or re­moved again with DROP EXTENSION REPOSITORY. Together with the sta­ble C API, the ex­ten­sion story rounds out nicely: write your ex­ten­sion once, sign it, host it wher­ever you like, and INSTALL it any­where.

Bonus: DuckDB Foundation — Advisory Board

Starting this fall, we will add a stake­holder ad­vi­sory board to the DuckDB Foundation. The ad­vi­sory board will pro­vide in­put on the de­vel­op­ment roadmap of DuckDB, DuckLake, and Quack. This al­lows key stake­hold­ers to have a say in the pro­jects’ di­rec­tion.

Final Thoughts

These are only a few high­lights, and this post is only a pre­view. Some de­tails may still shift be­fore the re­lease this fall, and there are many more fea­tures and im­prove­ments that we could not cover here. DuckDB v2.0 will also come with a small set of break­ing changes, in­clud­ing the new de­fault stor­age for­mat and the com­pleted lambda syn­tax tran­si­tion, which we will cover in de­tail in the re­lease an­nounce­ment.

There have been more than 10,000 com­mits by many con­trib­u­tors since we re­leased v1.5. We would like to thank our com­mu­nity for the de­tailed is­sue re­ports, feed­back, and con­tri­bu­tions that shaped this re­lease. If you want a taste be­fore the fall, the pre­view builds have most of these fea­tures to­day, and if some­thing breaks, you know where the is­sue tracker is.

Recent Posts

Thank You for 40 000 Stars on GitHub

The DuckDB team

Asynchronous I/O in DuckDB: Work, Thread, Work

Pedro Holanda

Announcing DuckDB 1.5.5

The DuckDB team

Incident with GitHub.com

www.githubstatus.com

Resolved

This in­ci­dent has been re­solved. Thank you for your pa­tience and un­der­stand­ing as we ad­dressed this is­sue. A de­tailed root cause analy­sis will be shared as soon as it is avail­able.

Posted Aug 17, 2026 – 21:15 UTC

Update

We are con­tin­u­ing to ap­ply mit­i­ga­tions to ad­dress spo­radic Copilot au­then­ti­ca­tion fail­ures in some ap­pli­ca­tions. We ex­pect full re­cov­ery within the next 30 min­utes. Copilot us­age via the GitHub CLI and GitHub App are un­af­fected.

Posted Aug 17, 2026 – 20:45 UTC

Update

Issues is op­er­at­ing nor­mally.

Posted Aug 17, 2026 – 20:22 UTC

Update

We are con­tin­u­ing to in­ves­ti­gate spo­radic fail­ures af­fect­ing Copilot au­then­ti­ca­tion in some ap­pli­ca­tions. Copilot us­age via the GitHub CLI and GitHub App are un­af­fected.

Posted Aug 17, 2026 – 20:08 UTC

Update

We are con­tin­u­ing to in­ves­ti­gate spo­radic au­then­ti­ca­tion fail­ures. We have par­tially dis­abled au­then­ti­ca­tion to­ken re­tries and have seen im­prove­ment, and we are mon­i­tor­ing im­pact be­fore fully ap­ply­ing this mit­i­ga­tion.

Posted Aug 17, 2026 – 19:13 UTC

Update

API Requests is op­er­at­ing nor­mally.

Posted Aug 17, 2026 – 19:01 UTC

Update

API Requests is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 18:48 UTC

Update

The degra­da­tion af­fect­ing Git Operations has been mit­i­gated. We are mon­i­tor­ing to en­sure sta­bil­ity.

Posted Aug 17, 2026 – 18:23 UTC

Update

We iden­ti­fied the prob­lem­atic com­po­nent and have taken cor­rec­tive ac­tions, but we are see­ing resid­ual im­pact in the form of spo­radic au­then­ti­ca­tion fail­ures. We are con­tin­u­ing to ap­ply ad­di­tional mit­i­ga­tions and in­ves­ti­gate the re­main­ing im­pact.

Posted Aug 17, 2026 – 18:11 UTC

Update

Issues is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 17:36 UTC

Update

We iden­ti­fied the prob­lem­atic com­po­nent and have taken cor­rec­tive ac­tions, but we are see­ing resid­ual im­pact across nu­mer­ous ser­vices. We are con­tin­u­ing to ap­ply ad­di­tional mit­i­ga­tions and in­ves­ti­gate the re­main­ing im­pact.

Posted Aug 17, 2026 – 17:34 UTC

Update

Git Operations is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 17:30 UTC

Update

The degra­da­tion af­fect­ing API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mit­i­gated. We are mon­i­tor­ing to en­sure sta­bil­ity.

Posted Aug 17, 2026 – 16:59 UTC

Update

We iden­ti­fied the prob­lem­atic com­po­nent and have taken cor­rec­tive ac­tions. There are strong signs of re­cov­ery but we are still work­ing to com­pletely re­store ser­vice, with er­ror rates still re­main­ing slightly el­e­vated. We will post fur­ther up­dates as re­cov­ery con­tin­ues.

Posted Aug 17, 2026 – 16:36 UTC

Update

We are ex­pe­ri­enc­ing high er­ror rates around 20% for web ex­pe­ri­ences and api traf­fic. Archive down­loads and raw repos­i­tory con­tent down­loads are ex­pe­ri­enc­ing an ap­prox­i­mate 50% er­ror rate. SAML and OIDC au­then­ti­ca­tion, SCIM, and Team Sync are also im­pacted. We are still work­ing to iden­tify the root cause and will con­tinue to post up­dates as we learn more and per­form mit­i­ga­tion.

Posted Aug 17, 2026 – 16:16 UTC

Update

We are ex­pe­ri­enc­ing high er­ror rates around 20% for web ex­pe­ri­ences and api traf­fic. Archive down­loads and raw repos­i­tory con­tent down­loads are ex­pe­ri­enc­ing an ap­prox­i­mate 50% er­ror rate. SAML and OIDC au­then­ti­ca­tion, SCIM, and Team Sync are also im­pacted. We are cur­rently per­form­ing mit­i­ga­tions and will post up­dates as we progress.

Posted Aug 17, 2026 – 15:42 UTC

Update

Webhooks is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 15:40 UTC

Update

Git Operations is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 15:21 UTC

Update

Pages is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 15:10 UTC

Update

API Requests is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 15:01 UTC

Update

Webhooks is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 14:58 UTC

Update

We are ex­pe­ri­enc­ing high er­ror rates around 20% for web ex­pe­ri­ences and api traf­fic. Archive down­loads and raw repos­i­tory con­tent down­loads are ex­pe­ri­enc­ing an ap­prox­i­mate 50% er­ror rate. SAML and OIDC au­then­ti­ca­tion, SCIM, and Team Sync are also im­pacted. We are cur­rently per­form­ing mit­i­ga­tions based on our in­ves­ti­ga­tion thus far and are mon­i­tor­ing for im­prove­ment.

Posted Aug 17, 2026 – 14:58 UTC

Update

Actions is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 14:58 UTC

Update

Pull Requests is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 14:54 UTC

Update

Issues is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 14:49 UTC

Update

Pull Requests is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 14:45 UTC

Update

Copilot is ex­pe­ri­enc­ing de­graded avail­abil­ity. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 14:31 UTC

Update

We are ex­pe­ri­enc­ing high er­ror rates around 20% for web ex­pe­ri­ences and api traf­fic. Archive down­loads and raw repos­i­tory con­tent down­loads are ex­pe­ri­enc­ing an ap­prox­i­mate 50% er­ror rate. SAML and OIDC au­then­ti­ca­tion, SCIM, and Team Sync are also im­pacted. Investigations are on-go­ing and we will con­tinue to pro­vide up­dates as we dis­cover more in­for­ma­tion.

Posted Aug 17, 2026 – 14:24 UTC

Update

We are ex­pe­ri­enc­ing high er­ror rates around 20% for web ex­pe­ri­ences and api traf­fic. Archive down­loads and raw repos­i­tory con­tent down­loads are ex­pe­ri­enc­ing an ap­prox­i­mate 50% er­ror rate. Investigations are on-go­ing into the root cause, and up­dates will con­tinue to be pro­vided as we in­ves­ti­gate.

Posted Aug 17, 2026 – 14:04 UTC

Update

Pull Requests is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 13:58 UTC

Update

Issues is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 13:46 UTC

Update

We are see­ing an ap­prox­i­mate 20% er­ror rate across nu­mer­ous ex­pe­ri­ences in­clud­ing Pull Requests, Issues, and oth­ers. Investigations are cur­rently un­der way and we will be post­ing up­dates as they be­come avail­able

Posted Aug 17, 2026 – 13:45 UTC

Update

Webhooks is ex­pe­ri­enc­ing de­graded per­for­mance. We are con­tin­u­ing to in­ves­ti­gate.

Posted Aug 17, 2026 – 13:44 UTC

Update

Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+

techcrunch.com

In Brief

Posted:

1:57 PM PDT · August 16, 2026

Stripe has fi­nal­ized a deal to ac­quire OpenRouter, ac­cord­ing to a new re­port in Bloomberg.

OpenRouter helps cus­tomers se­lect dif­fer­ent AI mod­els to per­form dif­fer­ent tasks, de­pend­ing on their spe­cific needs and bud­get. The com­pany an­nounced in May that it had raised a $113 mil­lion Series B, at a re­ported $1.3 bil­lion val­u­a­tion. (Investors in­clude Sequoia, Andreessen Horowitz, Menlo Ventures, and Alphabet’s CapitalG.)

At the time, OpenRouter CEO Alex Atallah de­scribed the com­pany as the equiv­a­lent of Stripe for AI, be­cause it pro­vides cus­tomers with a sin­gle ac­cess point for dif­fer­ent sys­tems and pre­vents lock-in. The startup also claimed to have 8 mil­lion global users and to pro­vide ac­cess to more than 400 mod­els.

The Wall Street Journal re­ported last month that Stripe and OpenRouter were in ac­qui­si­tion talks. Now, Bloomberg said those dis­cus­sions have led to a deal price of more than $7 bil­lion.

A Stripe spokesper­son told TechCrunch that the com­pany does not com­ment on ru­mors or spec­u­la­tion.

Topics

Subscribe for the in­dus­try’s biggest tech news

Latest in AI

Universal Health Coverage Could Save $1 Trillion and 114,000 Lives Every Year, Yale Study Projects

ysph.yale.edu

By Matt Kristoffersen

August 13, 2026

A sin­gle-payer uni­ver­sal health care sys­tem could cover every American, save more than 100,000 lives a year, and still cost $1 tril­lion less than the sys­tem it would re­place, ac­cord­ing to a new preprint study led by re­searchers at the Yale School of Public Health.

For the study, which has not yet been peer re­viewed, the re­searchers mod­eled what would hap­pen if the United States adopted a na­tional pub­lic in­sur­ance pro­gram like the one pro­posed in the Medicare for All Act. Using 2024 spend­ing, in­sur­ance cov­er­age, and mor­tal­ity data, they es­ti­mate that the uni­ver­sal cov­er­age would re­duce an­nual health ex­pen­di­tures by $1.04 tril­lion, or nearly 20% — even af­ter ac­count­ing for the ad­di­tional care that unin­sured and un­der­in­sured peo­ple would re­ceive.

Healthcare costs have been ris­ing faster than in­fla­tion, and a stag­ger­ing share of that spend­ing is con­sumed by ad­min­is­tra­tive mid­dle­men, soar­ing drug prices, and emer­gency care for con­di­tions that should have been treated ear­lier and for less.Al­i­son Galvani, PhDBurnett and Stender Families Professor of Epidemiology (Microbial Diseases) and Director, Center for Infectious Disease Modeling and Analysis

Healthcare costs have been ris­ing faster than in­fla­tion, and a stag­ger­ing share of that spend­ing is con­sumed by ad­min­is­tra­tive mid­dle­men, soar­ing drug prices, and emer­gency care for con­di­tions that should have been treated ear­lier and for less.

Alison Galvani, PhD

Burnett and Stender Families Professor of Epidemiology (Microbial Diseases) and Director, Center for Infectious Disease Modeling and Analysis

Healthcare costs have been ris­ing faster than in­fla­tion, and a stag­ger­ing share of that spend­ing is con­sumed by ad­min­is­tra­tive mid­dle­men, soar­ing drug prices, and emer­gency care for con­di­tions that should have been treated ear­lier and for less,” said se­nior au­thor Alison Galvani, the Burnett and Stender Families Professor of Epidemiology at YSPH and the di­rec­tor of the Yale Center for Infectious Disease Modeling and Analysis. Medicare for All strips out those sources of waste while pro­vid­ing every­one with health­care, sav­ing over a tril­lion dol­lars and 114,000 lives every year.”

The model iden­ti­fied five ma­jor sources of sav­ings: lower phar­ma­ceu­ti­cal prices, Medicare-level pay­ments to providers, re­duced ad­min­is­tra­tive over­head, less fraud­u­lent billing, and fewer avoid­able emer­gency de­part­ment vis­its and hos­pi­tal­iza­tions. Even un­der more con­ser­v­a­tive as­sump­tions about drug prices and fraud re­duc­tion, the re­searchers pro­jected that the sav­ings would reach at least $663 bil­lion a year. Both fig­ures ac­count for an es­ti­mated $304 bil­lion in ad­di­tional spend­ing to meet un­met med­ical needs, re­im­burse care that now goes un­paid, and pro­vide uni­ver­sal den­tal cov­er­age.

A uni­ver­sal health­care pol­icy would also save tens of thou­sands of lives, the re­searchers pro­ject. They es­ti­mate that ad­e­quate cov­er­age for all could avert about 62,863 deaths an­nu­ally — and that nearly half, or 29,631, would be among peo­ple who al­ready hold in­sur­ance. These are the un­der­in­sured: the more than 45 mil­lion work­ing-age adults whose de­ductibles and cost-shar­ing put care be­yond fi­nan­cial reach any­way. Reversing cov­er­age roll­backs and other health poli­cies en­acted since 2025 would avert a fur­ther 51,311 deaths each year, the re­searchers pro­ject, bring­ing the an­nual to­tal to 114,174.

The study builds on find­ings Galvani, co-au­thor Meagan Fitzpatrick, and other col­leagues pub­lished in The Lancet in 2020, which pro­jected that uni­ver­sal health care would save $450 bil­lion and more than 68,000 lives an­nu­ally. The larger es­ti­mates in the new study re­flect bal­loon­ing health ex­pen­di­tures, a widen­ing gap be­tween com­mer­cial and Medicare pay­ment rates, and new es­ti­mates on re­cent pol­icy changes and the un­der­in­sured.

The au­thors cau­tion that di­rect es­ti­mates of ex­cess mor­tal­ity among un­der­in­sured adults are un­avail­able, re­quir­ing them to model that risk. Their spend­ing analy­sis also does not ac­count for tran­si­tion costs, ad­min­is­tra­tive job losses, or how providers might re­spond to Medicare pay­ment rates.

A sys­tem that cov­ers every­one, costs $1 tril­lion less, and averts more than 100,000 deaths an­nu­ally re­quires no new dis­cov­ery to im­ple­ment, only en­act­ment.

A sys­tem that cov­ers every­one, costs $1 tril­lion less, and averts more than 100,000 deaths an­nu­ally re­quires no new dis­cov­ery to im­ple­ment, only en­act­ment.

Even with those lim­i­ta­tions, the re­searchers ar­gue that the United States al­ready spends enough to pro­vide uni­ver­sal cov­er­age. The prob­lem, they con­clude, is how that money is al­lo­cated.

A sys­tem that cov­ers every­one, costs $1 tril­lion less, and averts more than 100,000 deaths an­nu­ally re­quires no new dis­cov­ery to im­ple­ment, only en­act­ment,” they wrote.

The study’s other au­thors are Abhishek Pandey, se­nior re­search sci­en­tist in epi­demi­ol­ogy (microbial dis­eases); Chad Wells, post­doc­toral re­search as­so­ci­ate; and Yang Ye, as­so­ci­ate re­search sci­en­tist in epi­demi­ol­ogy (microbial dis­eases), at YSPH.

Article outro

Author

Matt Kristoffersen

Featured in this ar­ti­cle

Olo (color)

en.wikipedia.org

From Wikipedia, the free en­cy­clo­pe­dia

Olo is an imag­i­nary color that can only be seen us­ing spe­cial­ized tools that ex­clu­sively ac­ti­vate the M cone cells on the retina.

It is im­pos­si­ble to view olo un­der nor­mal view­ing con­di­tions, due to the over­lap­ping sen­si­tiv­i­ties of M cone cells and S and L cone cells in all wave­lengths of vis­i­ble light that evoke them. In other words, there is no mono­chro­matic stim­u­lus (the purest type of stim­u­lus that hu­mans can per­ceive) that ac­ti­vates only the M cones. This means that olo is out­side the vis­i­ble gamut. To get around this, re­searchers mapped a por­tion of the retina and in­di­vid­u­ally iden­ti­fied each cone cell as ei­ther an S, M, or L cone. They then used lasers to de­liver tiny doses of light at­tempt­ing to tar­get specif­i­cally the M cone cells, mostly avoid­ing the S and L cone cells.[1][2]

The re­searchers who ex­pe­ri­enced olo said that the clos­est color to olo in the sRGB gamut is hexa­dec­i­mal code #00FFCC.[2]

Olo was dis­cov­ered on April 18, 2025 by sci­en­tists at UC Berkeley.[1][3] The color is named af­ter its the­o­ret­i­cal LMS color space co­or­di­nates (0, 1, 0), which spells olo” in leet speak.[4][3]

Only the five sub­jects of the Berkeley ex­per­i­ment have of­fi­cially seen olo.[1][5] Professor Ren Ng, a co-au­thor of the study, de­scribed olo as more sat­u­rated than any color that you can see in the real world”;[5] the five sub­jects of the ex­per­i­ment sim­i­larly de­scribed the color as a blue-green of un­prece­dented sat­u­ra­tion”.[1] Ng and his team are ex­plor­ing whether the tech­nol­ogy used to gen­er­ate olo could be adapted to en­hance color per­cep­tion in in­di­vid­u­als with color blind­ness to man­age the symp­toms of color blind­ness. He fur­ther sug­gested that this ap­proach could even lead to a form of en­hanced vi­sion known as tetra­chro­macy, where in­di­vid­u­als may per­ceive a broader range of col­ors.[6]

Experts in the field have de­scribed the tech­nique used to cre­ate olo as a sig­nif­i­cant tech­ni­cal achieve­ment. The Berkeley team gen­er­ated the color by pre­cisely stim­u­lat­ing in­di­vid­ual cone cells in the retina us­ing lasers, cre­at­ing a color be­yond the hu­man vis­i­ble gamut.[7]

However, some sci­en­tists, in­clud­ing Professor John Barbur from City St George’s, University of London, have ques­tioned whether olo truly rep­re­sents a new” color, say­ing that its ex­is­tence is open to ar­gu­ment”.[5] Skepticism within the sci­en­tific com­mu­nity re­gard­ing the clas­si­fi­ca­tion of olo as a gen­uinely novel color has been noted.[3][8]

The idea of olo has drawn at­ten­tion be­yond the sci­en­tific com­mu­nity, with artists ex­press­ing in­ter­est in cre­at­ing paints in­spired by the color. The Berkeley re­search team has also re­ceived global in­ter­est, with re­quests from re­porters seek­ing to ex­pe­ri­ence the phe­nom­e­non first­hand.[9]

1 2 3 4 Fong, James; Doyle, Hannah K.; Wang, Congli; Boehm, Alexandra E.; Herbeck, Sofie R.; Pandiyan, Vimal Prabhu; Schmidt, Brian P.; Tiruveedhula, Pavan; Vanston, John E.; Tuten, William S.; Sabesan, Ramkumar; Roorda, Austin; Ng, Ren (2025 – 04-18). Novel color via stim­u­la­tion of in­di­vid­ual pho­tore­cep­tors at pop­u­la­tion scale”. Science Advances. 11 (16) ead­u1052. doi:10.1126/​sci­adv.adu1052. PMC 12007580. PMID 40249825.

1 2 Krywko, Jacek. Parshall, Allison (ed.). Only Five People Have Seen This New Impossible Color”. Scientific American. Retrieved 2025 – 05-04.

1 2 3 Sample, Ian (2025 – 04-18). Hue new? Scientists claim to have found colour no one has seen be­fore”. The Guardian. ISSN 0261 – 3077. Retrieved 2025 – 05-01.

↑ Lanese, Nicoletta (2025 – 04-18). ‘Olo’ is a brand-new color only ever seen by 5 peo­ple”. Live Science. Retrieved 2025 – 05-04.

1 2 3 Scientists claim to have dis­cov­ered new colour’ no one has seen be­fore”. BBC. 19 April 2025. Retrieved 8 May 2025.

Have sci­en­tists dis­cov­ered a new colour called olo’?”. Al Jazeera. 26 April 2025. Retrieved 18 May 2025.

Brand-new colour cre­ated by trick­ing hu­man eyes with laser”. Nature. 18 April 2025. Retrieved 8 May 2025.

↑ Hashemi, Sara. Scientists Say They’ve Discovered a New Color—an Unprecedented’ Hue Only Ever Seen by Five People”. Smithsonian Magazine. Retrieved 2025 – 05-04.

The Profound’ Experience of Seeing a New Color”. The Atlantic. 23 April 2025. Retrieved 8 May 2025.

How Bluesky draws its logo on screenshots

timmarinin.net

Sometimes I take a screen­shot of a post I like, ei­ther to send it to friends/​meme chan­nel or to save a durable” copy. Like this one (I’ve cropped out the rest of the in­ter­face):

I no­ticed the Bluesky logo in the right cor­ner and thought that it was weird that the logo does­n’t bother me when I use the app. Then I looked at the post in the app again—logo was­n’t there, re­placed by the Follow” but­ton.

I re­mem­bered that a few apps hide their logo where the iPhone notch is, so that it does­n’t stick out, un­less you take a screen­shot. But here the logo is placed in the open, so how do they do it?

I tried to take an­other screen­shot, this time mid-switch­ing to the other app:

Did they some­how set up a lis­tener for two but­tons I’m press­ing to take a screen­shot and do a switcheroo at the last mo­ment? I’m not an iOS de­vel­oper, so I’m not sure what’s pos­si­ble and what is not over there.

At this point I was mildly in­trigued. Thankfully, I re­mem­bered that Bluesky app is open source (or at least the code is avail­able to look at).

The an­swer was in the file lit­er­ally called GrowthHack.tsx , in­tro­duced in January 2026 by mozzius. But it merely used a de­pen­dency, so to un­der­stand I looked into pack­age expo-pri­vacy-sen­si­tive, also by them.

The pack­age cre­ates UITextField with is­Se­cure­Tex­tEn­try prop­erty set to true and ren­ders the ac­tual con­tent (the but­ton) into that field’s  .layer. When I take the screen­shot, iOS hides this UITextField by blank­ing the layer, al­low­ing the Bluesky logo to flut­ter its wings through (it was here the whooole time). For other plat­forms it sim­ply ren­ders con­tent as-is, with­out mask­ing.

Why does­n’t it work when I switch be­tween the apps? I sup­pose that iOS takes a snap­shot it­self at the start of the ges­ture (without trig­ger­ing blank­ing), and when I do a screen­shot, there is no live UITextField in­stance to re­act to that, only the in­ert snap­shot. But once again, I’m not an iOS de­vel­oper.

Nifty trick or an abuse of API meant for pri­vacy? The peo­ple in the thread adding the be­hav­ior mostly did­n’t like it, be­fore the thread got locked. I think it’s cute.

I googled a bit, and the trick is well-known. Telegram im­ple­mented sim­i­lar thing for its secret” chats, as did Signal, so I don’t ex­pect it to be patched by Apple any time soon.

Wiz Red Agent Finds Its Way Into Snowflake’s Internal Jira Through a Flaw in a GitHub Copilot–Assisted PR

www.wiz.io

As part of on­go­ing se­cu­rity re­search con­ducted through Snowflake’s HackerOne vul­ner­a­bil­ity dis­clo­sure pro­gram, Wiz Research’s Red Agent”—an au­tonomous, AI-powered se­cu­rity re­search tool—iden­ti­fied a crit­i­cal GitHub Actions work­flow vul­ner­a­bil­ity in one of Snowflake’s pub­lic repos­i­to­ries.

This in­ci­dent high­lights a new re­al­ity in soft­ware de­vel­op­ment: Critical vul­ner­a­bil­i­ties can still be in­tro­duced and ap­proved within work­flows in­volv­ing AI cod­ing agents, while au­tonomous AI se­cu­rity agents can rapidly dis­cover and ex­ploit them in the wild.

Upon re­spon­si­ble dis­clo­sure on June 23, 2026 by Wiz, Snowflake re­me­di­ated the vul­ner­a­bil­ity on the same day, ro­tated the af­fected cre­den­tial, and ver­i­fied via de­tailed au­dit logs that Wiz was the sole ac­tor dur­ing the ex­po­sure win­dow. Wiz con­firmed that all data ac­cessed dur­ing proof-of-con­cept test­ing was se­curely deleted.

August 17, 2026, 1957 UTC up­date: This blog has been up­dated to clar­ify that Copilot was a co-au­thor that checked the merged PR and code change, and iden­ti­fied it as all-clear with­out notic­ing the crit­i­cal vul­ner­a­bil­i­ties. It’s un­clear whether the code-change was AI-assisted.

Executive Summary

Wiz Red Agent iden­ti­fied a script in­jec­tion vul­ner­a­bil­ity in snowflakedb/​snowflake-con­nec­tor-net. The is­sue al­lowed an unau­then­ti­cated user to ex­e­cute ar­bi­trary com­mands within a GitHub Actions run­ner by open­ing a GitHub is­sue with a spe­cially crafted ti­tle.

Crucially, the vul­ner­a­bil­ity be­came live on June 18, 2026 - just five days be­fore its dis­cov­ery -when PR #1218 was merged. The fi­nal squash com­mit cred­its Copilot Autofix pow­ered by AI as a co-au­thor. The merged PR re­placed the repos­i­to­ry’s san­i­tized in­put pat­tern with di­rect string ex­pan­sion, yet GitHub’s AI-assisted se­cu­rity re­view did not flag the re­sult­ing crit­i­cal vul­ner­a­bil­ity.

Exposure Walk-Through

Discovery

Wiz Red Agent’s CI/CD ca­pa­bil­ity scanned Snowflake’s GitHub or­ga­ni­za­tion and flagged the ji­ra_is­sue.yml Workflow in snowflakedb/​snowflake-con­nec­tor-net as vul­ner­a­ble to script in­jec­tion via un­trusted in­put in run: blocks.

The Code Change

The work­flow trig­gered on is­sues: opened - mean­ing any GitHub user could fire it by open­ing an is­sue - and in­ter­po­lated the at­tacker-con­trolled is­sue ti­tle di­rectly into a shell script:

The sed es­cap­ing runs af­ter GitHub’s tem­plate ex­pan­sion, a sin­gle quote in the ti­tle breaks out of echo …’ and al­lows ar­bi­trary com­mand ex­e­cu­tion.

The in­jectable pat­tern was in­tro­duced just days ear­lier, on June 18, 2026, com­mit 4a1b8ce (PR #1218: SNOW-2069227: Update jira work­flows”) - co-au­thored by Copilot Autofix pow­ered by AI.

It re­moved the repos­i­to­ry’s ex­ist­ing safe pat­tern, which passed the is­sue ti­tle through an env: vari­able and built the JSON pay­load with jq. Instead it used the di­rect ${{ github.event.is­sue.ti­tle }} in­ter­po­la­tion shown above. In other words, an AI autofix” com­mit cre­ated the very in­jec­tion vec­tor.

The Open Security Gate”

The work­flow had an if: con­di­tion that ap­peared pro­tec­tive:

However, on is­sues events, github.event.pul­l_re­quest is al­ways null.

So the con­di­tion re­duces to (null != whitesource-for-github-com[bot]’). This is al­ways true, and every GitHub user passes the gate.

Exploitation

We crafted an is­sue ti­tle that, af­ter tem­plate ex­pan­sion, breaks out of the echo string and ex­fil­trates the Jira cre­den­tials via an out-of-band call­back:

Crucially, when Red Agent’s cicd ca­pa­bil­ity ini­tially at­tempted ex­fil­tra­tion us­ing a stan­dard com­ment char­ac­ter (#), the run­ner re­turned a bash syn­tax er­ror be­cause the com­ment con­sumed the clos­ing par­en­thet­i­cal of TITLE=$(…). Rather than stop­ping or fail­ing, Red Agent:

au­tonomously an­a­lyzed the syn­tax ex­e­cu­tion er­ror

au­tonomously an­a­lyzed the syn­tax ex­e­cu­tion er­ror

ad­justed its pay­load to use ; echo ′ to prop­erly close the shell block, and

ad­justed its pay­load to use ; echo ′ to prop­erly close the shell block, and

suc­cess­fully re­ceived the out-of-band call­back

suc­cess­fully re­ceived the out-of-band call­back

Within sec­onds, our lis­tener re­ceived the call­back from a GitHub Actions run­ner (Azure IP 20.106.182.197) con­tain­ing base64-en­coded cre­den­tials.

Note: Our first at­tempt used # to com­ment out the rest of the line, which caused an un­ex­pected EOF bash er­ror be­cause it also ate the clos­ing ) of TITLE=$(…). The fix was us­ing ; echo ′ to prop­erly close the shell syn­tax.

The ex­fil­trated to­ken au­then­ti­cated as qa@snowflake.net to snowflake­com­put­ing.at­lass­ian.net, grant­ing read ac­cess across Snowflake’s en­gi­neer­ing, se­cu­rity com­pli­ance, and bug bounty track­ing pro­jects.

Remediation & Forensics

Same-Day Patching: Snowflake patched the work­flow on June 23, 2026 (1dc7766, PR #1402), fully restor­ing the safe env: vari­able and jq –arg pars­ing pat­tern.

Same-Day Patching: Snowflake patched the work­flow on June 23, 2026 (1dc7766, PR #1402), fully restor­ing the safe env: vari­able and jq –arg pars­ing pat­tern.

Credential Revocation: The JIRA to­ken in ques­tion was re­voked and ro­tated.

Credential Revocation: The JIRA to­ken in ques­tion was re­voked and ro­tated.

Forensic Verification: Comprehensive au­dit log analy­sis con­firmed that no ex­ter­nal third par­ties ac­cessed the end­point dur­ing the 5-day ex­po­sure win­dow. All anom­alous queries were strictly matched to Wiz’s test­ing IPs.

Forensic Verification: Comprehensive au­dit log analy­sis con­firmed that no ex­ter­nal third par­ties ac­cessed the end­point dur­ing the 5-day ex­po­sure win­dow. All anom­alous queries were strictly matched to Wiz’s test­ing IPs.

Key Takeaways

AI Code Generation Demands Rigorous Oversight: AI cod­ing tools pre­dict code based on prob­a­bilis­tic pat­terns, which can in­ad­ver­tently rein­tro­duce dep­re­cated or in­se­cure shell pat­terns. AI-generated PRs must un­dergo the same sta­tic analy­sis and se­cu­rity scrutiny as hu­man code.

AI Code Generation Demands Rigorous Oversight: AI cod­ing tools pre­dict code based on prob­a­bilis­tic pat­terns, which can in­ad­ver­tently rein­tro­duce dep­re­cated or in­se­cure shell pat­terns. AI-generated PRs must un­dergo the same sta­tic analy­sis and se­cu­rity scrutiny as hu­man code.

Collapsing Discovery Windows: The vul­ner­a­bil­ity was live for only five days be­fore an au­to­mated agent dis­cov­ered and val­i­dated it. Security op­er­a­tions must adapt to a land­scape where au­to­mated dis­cov­ery oc­curs in hours, re­quir­ing rapid patch cy­cles and short-lived cre­den­tials.

Collapsing Discovery Windows: The vul­ner­a­bil­ity was live for only five days be­fore an au­to­mated agent dis­cov­ered and val­i­dated it. Security op­er­a­tions must adapt to a land­scape where au­to­mated dis­cov­ery oc­curs in hours, re­quir­ing rapid patch cy­cles and short-lived cre­den­tials.

Preventing AI Security Regressions: Automated AI as­sis­tants of­ten lack his­tor­i­cal con­text re­gard­ing why spe­cific code pat­terns were cho­sen. In this in­ci­dent, an au­to­mated PR re­moved a safe env: + jq pars­ing pat­tern that had been ex­plic­itly im­ple­mented to pre­vent shell in­jec­tion. Security teams must im­ple­ment Guardrails that block AI agents from re­plac­ing struc­tured data parsers with di­rect string in­ter­po­la­tion.

Preventing AI Security Regressions: Automated AI as­sis­tants of­ten lack his­tor­i­cal con­text re­gard­ing why spe­cific code pat­terns were cho­sen. In this in­ci­dent, an au­to­mated PR re­moved a safe env: + jq pars­ing pat­tern that had been ex­plic­itly im­ple­mented to pre­vent shell in­jec­tion. Security teams must im­ple­ment Guardrails that block AI agents from re­plac­ing struc­tured data parsers with di­rect string in­ter­po­la­tion.

Disclosure Timeline

June 18, 2026 - The vul­ner­a­bil­ity be­came live when PR #1218 was merged, co-au­thored by Copilot Autofix

June 18, 2026 - The vul­ner­a­bil­ity be­came live when PR #1218 was merged, co-au­thored by Copilot Autofix

June 23, 2026 - Wiz iden­ti­fied, ex­ploited, and re­ported vul­ner­a­bil­ity to Snowflake via HackerOne (report #3819931)

June 23, 2026 - Wiz iden­ti­fied, ex­ploited, and re­ported vul­ner­a­bil­ity to Snowflake via HackerOne (report #3819931)

June 23, 2026 - Slack no­ti­fi­ca­tion sent to Snowflake se­cu­rity team

June 23, 2026 - Slack no­ti­fi­ca­tion sent to Snowflake se­cu­rity team

June 23, 2026 (same day) - Snowflake patches the vul­ner­a­ble script-in­jec­tion work­flow (commit 1dc7766, PR #1402), restor­ing the safe env: + jq –arg pat­tern.

June 23, 2026 (same day) - Snowflake patches the vul­ner­a­ble script-in­jec­tion work­flow (commit 1dc7766, PR #1402), restor­ing the safe env: + jq –arg pat­tern.

June 24, 2026 - Jira to­ken ro­tated

June 24, 2026 - Jira to­ken ro­tated

July 25, 2026 - Public dis­clo­sure dead­line (30 days af­ter the June 25 res­o­lu­tion, per Snowflake’s dis­clo­sure pol­icy)

July 25, 2026 - Public dis­clo­sure dead­line (30 days af­ter the June 25 res­o­lu­tion, per Snowflake’s dis­clo­sure pol­icy)

Snowflake’s Response

Snowflake ap­pre­ci­ates Wiz’s re­spon­si­ble re­port­ing of and col­lab­o­ra­tion around these find­ings through our vul­ner­a­bil­ity dis­clo­sure and bug bounty pro­gram, HackerOne. Wiz Research re­ported a se­cu­rity vul­ner­a­bil­ity in one of Snowflake’s pub­lic GitHub repos­i­to­ries. The dis­clo­sure was re­ceived on June 23, 2026, and it was im­me­di­ately in­ves­ti­gated and re­me­di­ated, and our in­ves­ti­ga­tion found no ev­i­dence of unau­tho­rized ac­cess. Protecting our sys­tems re­mains a top pri­or­ity, and we re­main com­mit­ted to con­tin­u­ally strength­en­ing our soft­ware de­vel­op­ment and se­cu­rity prac­tices. We are work­ing to­gether with Wiz to share these learn­ings with the broader in­dus­try to en­cour­age wide­spread adop­tion of these se­cu­rity best prac­tices.

Snowflake ap­pre­ci­ates Wiz’s re­spon­si­ble re­port­ing of and col­lab­o­ra­tion around these find­ings through our vul­ner­a­bil­ity dis­clo­sure and bug bounty pro­gram, HackerOne. Wiz Research re­ported a se­cu­rity vul­ner­a­bil­ity in one of Snowflake’s pub­lic GitHub repos­i­to­ries. The dis­clo­sure was re­ceived on June 23, 2026, and it was im­me­di­ately in­ves­ti­gated and re­me­di­ated, and our in­ves­ti­ga­tion found no ev­i­dence of unau­tho­rized ac­cess. Protecting our sys­tems re­mains a top pri­or­ity, and we re­main com­mit­ted to con­tin­u­ally strength­en­ing our soft­ware de­vel­op­ment and se­cu­rity prac­tices. We are work­ing to­gether with Wiz to share these learn­ings with the broader in­dus­try to en­cour­age wide­spread adop­tion of these se­cu­rity best prac­tices.

Qwen3.8 27B - Intelligence, Performance & Price Analysis

artificialanalysis.ai

Intelligence

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.1.1 in­cor­po­rates 9 eval­u­a­tions: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

Reasoning mod­els are in­di­cated by a light­bulb icon

Artificial Analysis Intelligence Index v4.1.1 in­cludes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index method­ol­ogy for fur­ther de­tails, in­clud­ing a break­down of each eval­u­a­tion and how we run them.

Artificial Analysis Intelligence Index by Open Weights / Proprietary

Artificial Analysis Intelligence Index v4.1.1 in­cor­po­rates 9 eval­u­a­tions: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR

Reasoning mod­els are in­di­cated by a light­bulb icon

Artificial Analysis Intelligence Index v4.1.1 in­cludes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index method­ol­ogy for fur­ther de­tails, in­clud­ing a break­down of each eval­u­a­tion and how we run them.

Indicates whether the model weights are avail­able. Models are la­belled as Commercial Use Restricted’ if the weights are avail­able but com­mer­cial use is lim­ited (typically re­quires ob­tain­ing a paid li­cense).

Intelligence Evaluations

Intelligence eval­u­a­tions mea­sured in­de­pen­dently by Artificial Analysis · Higher is bet­ter

Agentic tool use

Reasoning & knowl­edge

Knowledge

1 - hal­lu­ci­na­tion rate

Long con­text rea­son­ing

Quantitative analy­sis on spread­sheets & doc­u­ments

Reasoning mod­els are in­di­cated by a light­bulb icon

While model in­tel­li­gence gen­er­ally trans­lates across use cases, spe­cific eval­u­a­tions may be more rel­e­vant for cer­tain use cases.

Artificial Analysis Intelligence Index v4.1.1 in­cludes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index method­ol­ogy for fur­ther de­tails, in­clud­ing a break­down of each eval­u­a­tion and how we run them.

AA-Omniscience

AA-Omniscience Index

AA-Omniscience Index (higher is bet­ter) mea­sures knowl­edge re­li­a­bil­ity and hal­lu­ci­na­tion. It re­wards cor­rect an­swers, pe­nal­izes hal­lu­ci­na­tions, and has no penalty for re­fus­ing to an­swer. Scores range from -100 to 100, where 0 means as many cor­rect as in­cor­rect an­swers, and neg­a­tive scores mean more in­cor­rect than cor­rect.

Reasoning mod­els are in­di­cated by a light­bulb icon

AA-Omniscience Index (higher is bet­ter) mea­sures knowl­edge re­li­a­bil­ity and hal­lu­ci­na­tion. It re­wards cor­rect an­swers, pe­nal­izes hal­lu­ci­na­tions, and has no penalty for re­fus­ing to an­swer. Scores range from -100 to 100, where 0 means as many cor­rect as in­cor­rect an­swers, and neg­a­tive scores mean more in­cor­rect than cor­rect.

Openness Index

Artificial Analysis Openness Index: Score

Openness Index as­sesses model open­ness on a 0 to 100 nor­mal­ized scale (higher is more open)

Reasoning mod­els are in­di­cated by a light­bulb icon

Intelligence Index Comparisons

Intelligence Index vs. Cost per Intelligence Index Task

Artificial Analysis Intelligence Index · Weighted av­er­age cost (USD) per Artificial Analysis Intelligence Index task

Most at­trac­tive quad­rant

Pareto line

Reasoning mod­els are in­di­cated by a light­bulb icon

Weighted av­er­age cost per Intelligence Index task. Each eval­u­a­tion’s cost is cal­cu­lated from in­put, cache hit, cache write, rea­son­ing, and an­swer to­ken prices, di­vided by task count, and weighted by its Intelligence Index weight.

Artificial Analysis Intelligence Index v4.1.1 in­cludes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR. See Intelligence Index method­ol­ogy for fur­ther de­tails, in­clud­ing a break­down of each eval­u­a­tion and how we run them.

Token Use

Output Tokens per Intelligence Index Task

Weighted av­er­age num­ber of out­put to­kens used to run one task in the Artificial Analysis Intelligence Index

Reasoning mod­els are in­di­cated by a light­bulb icon

The num­ber of to­kens re­quired per Intelligence Index task. This is cal­cu­lated by mul­ti­ply­ing the out­put to­kens per eval by the rel­a­tive weights of each bench­mark in the Intelligence Index, then di­vid­ing by task count (excluding re­peats).

Cost

Cost per Intelligence Index Task

Weighted av­er­age cost (USD) per Artificial Analysis Intelligence Index task, seg­mented by to­ken type. Lower is bet­ter

Reasoning mod­els are in­di­cated by a light­bulb icon

Weighted av­er­age cost per Intelligence Index task. Each eval­u­a­tion’s cost is cal­cu­lated from in­put, cache hit, cache write, rea­son­ing, and an­swer to­ken prices, di­vided by task count, and weighted by its Intelligence Index weight.

Cost to Run Artificial Analysis Intelligence Index

Cost (USD) to run all eval­u­a­tions in the Artificial Analysis Intelligence Index

Reasoning mod­els are in­di­cated by a light­bulb icon

The cost to run the eval­u­a­tions in the Artificial Analysis Intelligence Index, cal­cu­lated us­ing the mod­el’s in­put, cache hit, cache write, rea­son­ing, and an­swer to­ken prices and the num­ber of to­kens used across eval­u­a­tions (excluding re­peats).

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

Reasoning mod­els are in­di­cated by a light­bulb icon

Price per to­ken for cached prompts (previously processed), typ­i­cally of­fer­ing a sig­nif­i­cant dis­count com­pared to reg­u­lar in­put price, rep­re­sented as USD per mil­lion to­kens. The val­ues shown here are the cache hit price; cache write and cache stor­age are billed sep­a­rately and vary by provider — see Cache pric­ing by provider” for de­tail.

Context Window

Context Window

Context win­dow: to­kens limit · Higher is bet­ter

Reasoning mod­els are in­di­cated by a light­bulb icon

Larger con­text win­dows are rel­e­vant to RAG (Retrieval Augmented Generation) LLM work­flows which typ­i­cally in­volve rea­son­ing and in­for­ma­tion re­trieval of large amounts of data.

Maximum num­ber of com­bined in­put & out­put to­kens. Output to­kens com­monly have a sig­nif­i­cantly lower limit (varied by model).

Model Size (Open Weights Models Only)

Model Size: Total and Active Parameters

Comparison be­tween to­tal model pa­ra­me­ters and pa­ra­me­ters ac­tive dur­ing in­fer­ence

Reasoning mod­els are in­di­cated by a light­bulb icon

The to­tal num­ber of train­able weights and bi­ases in the model, ex­pressed in bil­lions. These pa­ra­me­ters are learned dur­ing train­ing and de­ter­mine the mod­el’s abil­ity to process and gen­er­ate re­sponses.

The num­ber of pa­ra­me­ters ac­tu­ally ex­e­cuted dur­ing each in­fer­ence for­ward pass, ex­pressed in bil­lions. For Mixture of Experts (MoE) mod­els, a rout­ing mech­a­nism se­lects a sub­set of ex­perts per to­ken, re­sult­ing in fewer ac­tive than to­tal pa­ra­me­ters. Dense mod­els use all pa­ra­me­ters, so ac­tive equals to­tal.

GPT 5.6 Sol is the best "vision" model OpenAI ever released

blog.roboflow.com

Last week, OpenAI an­nounced the GPT-5.6 lineup, in­tro­duc­ing the Sol, Terra, and Luna mod­els. During the re­lease stream, the team fo­cused heav­ily on com­puter use, show­ing mod­els ca­pa­ble of nav­i­gat­ing and op­er­at­ing desk­top ap­pli­ca­tions. OpenAI high­lighted UI agents and de­tailed 3D vi­su­al­iza­tions, but both de­pend on stronger vi­sual un­der­stand­ing.

To mea­sure their vi­sion ca­pa­bil­i­ties, we ran the mod­els through our up­com­ing VLM bench­mark, which we plan to re­lease in the next few weeks. The bench­mark cov­ers com­mon vi­sion tasks, in­clud­ing de­tec­tion, count­ing, OCR, and data ex­trac­tion. In this post, we take a closer look at how GPT-5.6 per­forms across each of them.

Sol is clearly the best vi­sion model OpenAI has re­leased so far. The jump is es­pe­cially vis­i­ble in ob­ject de­tec­tion and count­ing, where GPT-5.5 was far be­hind the strongest VLMs. Terra and Luna are not as strong as Sol, but both show mean­ing­ful progress over GPT-5.5.

Test Sol, Terra, and Luna in Roboflow Playground and com­pare their re­sults with mod­els such as Claude Fable 5 and Gemini 3.5 Flash across the same vi­sion tasks.

Roboflow Playground

Object Detection

Detection is where GPT-5.6 shows the clear­est jump. GPT-5.5 scored 13.8 mAP@50 in our bench­mark, while Sol reached 46.2. Terra and Luna fol­lowed closely at 44.7 and 43.3, mov­ing ob­ject de­tec­tion from a ma­jor weak­ness to a prac­ti­cal ca­pa­bil­ity.

Document lay­out de­tec­tion is one of the clear­est strengths of GPT-5.6. Sol han­dled ti­tles, para­graphs, ta­bles, im­ages, and sig­na­tures well. Many doc­u­ment work­flows start with lo­cat­ing the rel­e­vant parts of a page be­fore OCR or data ex­trac­tion be­gins.

GPT-5.6 also per­formed well on dense scenes. The pills and eggs ex­am­ples con­tain many sim­i­lar ob­jects packed closely to­gether, a com­mon weak­ness for VLM-based de­tec­tion. Unlike tra­di­tional de­tec­tors, VLMs gen­er­ate each class la­bel and set of co­or­di­nates as text. As ob­ject count grows, the re­sponse be­comes longer and the risk of missed ob­jects, du­pli­cates, or co­or­di­nate er­rors in­creases. Despite this, Sol de­tected most ob­jects across both scenes.

For the best de­tec­tion re­sults, prompt GPT-5.6 mod­els to re­turn ab­solute XYXY co­or­di­nates in im­age pix­els. This dif­fers from Gemini 3.5 Flash, which per­formed best with YXYX co­or­di­nates nor­mal­ized to a 0 – 1000 range. Using the wrong co­or­di­nate for­mat re­duced GPT-5.6 de­tec­tion per­for­mance by around 15 mAP points in our bench­mark.

In a few cases, GPT-5.6 Sol re­turned boxes in seem­ingly ran­dom parts of the im­age. Many had no over­lap, or al­most no over­lap, with the ground truth. Instead of match­ing the vis­i­ble ob­jects, the boxes of­ten formed un­nat­ural lay­outs, such as straight rows or evenly spaced groups.

We shared those ex­am­ples with OpenAI. Their team con­firmed that Sol be­comes less sta­ble on im­ages around 2,000 by 2,000 pix­els or larger, es­pe­cially at lower rea­son­ing ef­fort. Higher rea­son­ing ef­fort im­proves sta­bil­ity, but also in­creases to­ken use, la­tency, and cost. Resizing or crop­ping large im­ages be­fore send­ing them to the OpenAI API is the most prac­ti­cal workaround.

Object Counting

Counting im­proved across the full GPT-5.6 lineup. Sol scored 73.0% in our bench­mark, up from 64.9% for GPT-5.5, while Terra and Luna reached 67.6% and 66.2%. Luna, the cheap­est model in the lineup, still out­per­formed the pre­vi­ous OpenAI base­line.

As part of the bench­mark, we tested cases re­quir­ing more than spot­ting ob­jects and re­turn­ing a to­tal. Sol counted heav­ily over­lap­ping metal brack­ets, a dif­fi­cult case for both tra­di­tional ob­ject de­tec­tors and VLMs. Sol also counted bul­let holes only in­side se­lected scor­ing zones, show­ing an un­der­stand­ing of both which ob­jects to count and where the rule ap­plied.

Blister packs proved much harder. In sep­a­rate prompts, we asked Sol to count the empty slots and the pills still sealed in­side the pack­age. The re­peated lay­out, re­flec­tions, and small vi­sual dif­fer­ences be­tween filled and empty slots made both tasks dif­fi­cult.

The ab­nor­mal candy ex­am­ple ex­posed a dif­fer­ent type of fail­ure. Sol gave the wrong count, though it is un­clear whether the model mis­counted the can­dies or mis­un­der­stood the tar­get cat­e­gory.

OCR and Data Extraction

OCR per­for­mance stayed close to GPT-5.5. Sol achieved a 90.7% mean sim­i­lar­ity score, only 0.5 points be­hind GPT-5.5 at 91.2%, while Terra and Luna reached 88.8% and 88.4%. The gap was larger in text ex­trac­tion, where Sol scored 82.5% com­pared with 87.6% for GPT-5.5. Luna and Terra fol­lowed at 81.4% and 79.4%.

As part of the bench­mark, we sep­a­rated full tran­scrip­tion from tar­geted ex­trac­tion. OCR asks the model to tran­scribe all vis­i­ble text, while text ex­trac­tion asks for a spe­cific piece of in­for­ma­tion. Sol per­formed well on hand­writ­ten notes in both set­tings, pro­duc­ing a full tran­scrip­tion in one case and ex­tract­ing a re­quested date in an­other.

Sol per­formed well on text em­bed­ded in com­plex vi­sual scenes. It read a tire size se­quence printed along the curved sur­face of a dirty, worn tire. In an­other ex­am­ple, it ex­tracted the live score from a hockey broad­cast and re­turned the an­swer in the re­quested for­mat, test­ing both vi­sual read­ing and in­struc­tion fol­low­ing.

Some sim­ple-look­ing ex­trac­tion tasks still failed. Sol could not read the ex­pi­ra­tion date printed on a blis­ter pack. The text was small, ver­ti­cal, low con­trast, and af­fected by re­flec­tions, which may ex­plain the er­ror.

Trade-offs

The vi­sion gains come with higher to­ken us­age across the GPT-5.6 lineup. The dif­fer­ence mat­ters less in small tests, but be­comes more im­por­tant at scale, where to­ken vol­ume di­rectly in­creases pro­cess­ing costs.

Sol av­er­aged close to 10 sec­onds per im­age in our bench­mark. Terra re­duced that to around 6 sec­onds, while Luna fin­ished in slightly over 5 sec­onds. Luna of­fers the strongest la­tency-qual­ity bal­ance in the lineup, with speed close to Gemini 3.5 Flash while still out­per­form­ing GPT-5.5 on de­tec­tion and count­ing.

In our bench­mark, Sol cost roughly 2.5 cents per im­age, mak­ing it the sec­ond most ex­pen­sive model af­ter Claude Fable 5. Terra re­duced the av­er­age cost to about 1 cent per im­age, while Luna cost less than 0.5 cents.

At 0.8 cents per im­age, Gemini 3.5 Flash is much cheaper than Sol while still lead­ing our de­tec­tion and count­ing bench­marks. This makes it a strong op­tion for data-in­ten­sive work­loads where cost scales across large im­age batches. Roboflow Playground lets you test Sol, Terra, and Luna along­side Claude Fable 5, Gemini 3.5 Flash, and other VLMs on the same tasks.

Takeaways

With GPT-5.6, OpenAI is much closer to the lead­ing VLMs than be­fore. Detection moved from a weak point to a us­able ca­pa­bil­ity, and count­ing im­proved across the full model fam­ily.

There are still clear lim­its. Gemini 3.5 Flash re­mains a bet­ter prac­ti­cal choice for high-vol­ume de­tec­tion and count­ing in our bench­mark, es­pe­cially at its price.

GPT-5.6 shows OpenAI is now tak­ing vi­sion much more se­ri­ously. Sol still has flaws, es­pe­cially around cost, la­tency, and some un­sta­ble de­tec­tion cases, but the progress is hard to ig­nore. For agents, screen un­der­stand­ing, doc­u­ment work­flows, and vi­sual rea­son­ing, this re­lease makes OpenAI a much stronger op­tion than be­fore.

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.