10 interesting stories served every morning and every evening.

OpenRouter is Joining Stripe

openrouter.ai

Today, we are ex­cited to an­nounce that we are join­ing forces with Stripe, to power the next wave of GDP growth glob­ally.

OpenRouter is the first and largest model mar­ket­place and gate­way. We are the best way to dis­cover and use any AI model, with one in­ter­face, broad provider choice, model-ag­nos­tic ob­serv­abil­ity, cost man­age­ment, and rout­ing that im­proves price, per­for­mance, and up­time. We now process 10+ tril­lion to­kens per day from 400+ AI mod­els for a com­mu­nity of over 10 mil­lion de­vel­op­ers and com­pa­nies. Since our found­ing, we have seen at least 10x growth in in­fer­ence vol­ume every year.

We want to ex­plain why we made the de­ci­sion to join forces with Stripe and what it means for the mil­lions of de­vel­op­ers and com­pa­nies that build on us.

What this means for our users

OpenRouter will con­tinue to op­er­ate as it is: same mis­sion, same name, same prod­uct, same roadmap. If you build on OpenRouter to­day, noth­ing about your in­te­gra­tion changes.

OpenRouter ex­ists to give users every model on equal foot­ing, pro­vide open sig­nals about how they’re used in the mar­ket, help de­vel­op­ers or­ches­trate them to­gether, and make them ob­serv­able and man­age­able at scale. That com­mit­ment is core to how we op­er­ate, and it does­n’t bend to any model, any provider, or any par­ent com­pany. It also ex­tends to a grow­ing ecosys­tem of in­fer­ence-ad­ja­cent ser­vices, in­clud­ing AI-native web search, con­text man­age­ment, and more to come.

Routing de­ci­sions will re­main dri­ven by one thing: what’s best for you, the user.

Our mis­sion

We started OpenRouter in early 2023 on a sim­ple be­lief: in­tel­li­gence will be multi-model. No sin­gle model will win every task, and the fron­tier will move rapidly. That free­dom is crit­i­cal in­fra­struc­ture for the in­dus­try. AI is too im­por­tant for its fu­ture to be de­cided by whichever sin­gle model gets em­bed­ded first. AI has be­come the sin­gle largest dri­ver of eco­nomic growth in the US, and in­fer­ence is quickly be­com­ing the largest line item for every com­pany.

We en­vi­sion a healthy AI ecosys­tem where many mod­els thrive, where AI neu­ro­di­ver­sity is a strength, where a lab or an in­fer­ence provider with a break­through can reach mil­lions of de­vel­op­ers, and where no sin­gle model be­comes the de­fault by in­er­tia.

Our mis­sion is to re­al­ize this fu­ture, and it’s more im­por­tant than any­thing else. The op­por­tu­nity with Stripe al­lows us to ac­cel­er­ate it to­gether.

Why Stripe?

There are few com­pa­nies on earth we would have con­sid­ered sell­ing to; our mis­sion, our neu­tral­ity, and our lead in the mar­ket make the story for in­de­pen­dence strong. We would only join a com­pany if we thought we could do more to­gether, faster, with­out com­pro­mis­ing any of them.

Stripe is that com­pany. They are the best fi­nan­cial in­fra­struc­ture plat­form in the world. Their API set the stan­dard that de­vel­oper prod­ucts, in­clud­ing ours, have been mea­sured against ever since. This is a com­bi­na­tion of two plat­forms that de­vel­op­ers choose on merit, with cul­tures fo­cused on qual­ity, scale, and com­mit­ment to builders, and that will re­main es­sen­tial in a post-AGI econ­omy.

For years, OpenRouter has been called Stripe for LLMs.” Both com­pa­nies share com­mon DNA: we ab­stract com­plex in­fra­struc­ture and mar­ket dy­nam­ics into de­light­ful APIs, and we ob­sess over the de­vel­oper and user on the other side of it. Businesses trust Stripe to op­ti­mize every part of their rev­enue stack, across pay­ment meth­ods, au­tho­riza­tion, fraud, and more. Builders, cus­tomers, model labs, and providers trust OpenRouter to run a neu­tral, re­li­able layer across a fast-mov­ing ecosys­tem.

Stripe brings a large cus­tomer net­work, data on how in­ter­net busi­nesses grow, and years of ex­pe­ri­ence run­ning trusted global in­fra­struc­ture. There is also no one bet­ter at man­ag­ing fraud and abuse, some­thing we be­lieve will only be­come more chal­leng­ing for AI com­pa­nies to ad­dress. We can now serve de­vel­op­ers at a pace we could­n’t reach alone.

What’s next

To our cus­tomers: Thank you. We’re hon­ored to be part of your jour­ney, and we’re just get­ting started. OpenRouter’s prod­uct, mis­sion, and cur­rent com­mit­ments re­main un­changed. Joining Stripe helps us pur­sue them faster, and our abil­ity to sup­port you will only im­prove.

To our em­ploy­ees: We firmly be­lieve that the next few years will be the most im­por­tant time of our lives, and the most im­por­tant for our mis­sion. We are so grate­ful to have the priv­i­lege of be­ing alive dur­ing this trans­for­ma­tion for the world, and that I get to do it with you. You are some of the most bril­liant, cre­ative, cu­ri­ous, and de­ter­mined peo­ple there are, and we’re ex­cited to grow our team for an even greater global im­pact.

To every­one else who be­lieves in our mis­sion: come join us. Fitting into a role is­n’t as im­por­tant as hav­ing our val­ues: cu­rios­ity, rigor, agency, and trans­parency. AI will trans­form the way com­pa­nies are or­ga­nized, and OpenRouter will in­no­vate sig­nif­i­cantly here. And as we grow, we will re­lent­lessly aim to pre­serve the ve­loc­ity, agility, ef­fi­ciency, and tal­ent den­sity of the 90-person startup that we are to­day.

— Alex, Chris, Louis, and the OpenRouter team

The trans­ac­tion is sub­ject to cus­tom­ary clos­ing con­di­tions. We ex­pect to close in the com­ing weeks.

Don't paste the AI.

dontpastetheai.com

Prefer an­other lan­guage?

Don’t pastethe AI, please.

When some­one asks you some­thing, they want your an­swer. Not a wall of unedited ChatGPT out­put. A short re­ply from you beats a long one from a model, every time.

What’s hap­pen­ing

Someone asked you a real ques­tion. You popped it into a chat­bot, copied the an­swer, and sent it back. It felt fast. It felt help­ful. And it usu­ally is­n’t.

The per­son on the other side has the same tools you do. If they wanted the generic an­swer, they would have got­ten it in four sec­onds. They asked you be­cause they wanted your take on it… Your con­text, your taste, your judge­ment.

The world is filled with peo­ple that don’t want to read or think about things, don’t be one of them.

Try this in­stead

Use the AI. Really, go for it. It’s a great draft­ing part­ner. Just read what it gave you, then write your own take on it, don’t just be the proxy be­tween it and the an­swer.

Pull out the bit that ac­tu­ally an­swers the ques­tion. Drop the rest. Three sen­tences from you is plenty.

If a piece of the mod­el’s an­swer is gen­uinely use­ful, quote it and say why. I checked with Claude and this part lines up:” works great.

If you don’t have any­thing to add, it’s okay to say so. No strong opin­ion here” is a real, help­ful re­ply.

Want to send this to some­one?

If some­one just dropped a wall of model out­put in your DMs, Slack, or PR re­view, you can send them this link. No lec­ture re­quired.

Click to copy. They’ll get the hint.

Want a stronger ver­sion?

If you’d rather send some­one the ver­sion with feel­ings at­tached, we’ve got you cov­ered.

Take me to the an­gry ver­sion →

Not safe for send­ing to your man­ager. Probably.

GrapheneOS (@GrapheneOS@grapheneos.social)

grapheneos.social

To use the Mastodon web ap­pli­ca­tion, please en­able JavaScript. Alternatively, try one of the na­tive apps for Mastodon for your plat­form.

Go 1.27 is released - The Go Programming Language

go.dev

The Go Blog

Today the Go team is pleased to re­lease Go 1.27. You can find its bi­nary archives and in­stallers on the down­load page.

Go 1.27 brings ma­jor en­hance­ments across the lan­guage, tool­chain, run­time, and stan­dard li­brary. Below are some of the key high­lights.

Language changes

Go 1.27 in­tro­duces three no­table up­dates to the lan­guage spec­i­fi­ca­tion.

First, generic meth­ods are now sup­ported. For ex­am­ple, see math/​rand/​v2.Rand:

// Prior to Go 1.27, a sep­a­rate method on Rand had to be added for each type // (unsigned in­te­ger meth­ods omit­ted for brevity). func (r *Rand) Int32N(n in­t32) in­t32 func (r *Rand) Int64N(n in­t64) in­t64 func (r *Rand) IntN(n int) int

// Go 1.27 adds a new generic method that works for all in­te­ger types. func (r *Rand) N[Int int­Type](n Int) Int

Second, a key in a struct lit­eral may now be any valid field se­lec­tor for the struct type, al­low­ing fields in nested or em­bed­ded structs to be ini­tial­ized di­rectly:

type Habitat struct { Burrow string }

type Gopher struct { Name string Habitat // Embedded struct. }

// Go 1.27 al­lows us­ing Burrow as a key di­rectly. g := Gopher{ Name: Gopher”, Burrow: Burrow #42″, }

Finally, func­tion type in­fer­ence has been gen­er­al­ized to ap­ply in all as­sign­ment con­texts. Generic func­tions can now be used with­out ex­plicit type ar­gu­ments in com­pos­ite lit­er­als, type con­ver­sions, and chan­nel sends:

func GenericFormatter[T any](v T) string { re­turn fmt.Sprintf(“value: %v”, v) }

type IntFormatter func(int) string

// Go 1.27 in­fers T = int in com­pos­ite lit­er­als, con­ver­sions, and chan­nel sends. for­mat­ters := []IntFormatter{GenericFormatter} fn := IntFormatter(GenericFormatter) ch := make(chan IntFormatter, 1) ch <- GenericFormatter

Tool im­prove­ments

go fix in­cludes sev­eral new mod­ern­iz­ers: atom­ic­types, em­bedlit, slices­back­ward, and un­safe­funcs.

go doc now sup­ports pack­age@ver­sion queries such as go doc ex­am­ple.com/​pkg@v1.2.3.

go mod tidy now au­to­mat­i­cally con­sol­i­dates mul­ti­ple re­quire blocks in go.mod into a stan­dard di­rect and in­di­rect two-block struc­ture.

Performance and run­time

Size-specialized mem­ory al­lo­ca­tion re­duces small ob­ject (<80B) al­lo­ca­tion costs by up to 30%, im­prov­ing over­all per­for­mance by ~1% for al­lo­ca­tion-heavy pro­grams.

The gor­ou­tine­leak pro­file in run­time/​pprof is now gen­er­ally avail­able, al­low­ing au­to­matic de­tec­tion of per­ma­nently blocked gor­ou­tines.

Standard li­brary ad­di­tions

en­cod­ing/​json/​v2 pro­vides high-level JSON pro­cess­ing with con­fig­urable op­tions and stricter de­faults, along­side en­cod­ing/​json/​json­text for low-level stream­ing. The ex­ist­ing en­cod­ing/​json pack­age is now backed by the v2 im­ple­men­ta­tion for faster un­mar­shal­ing while main­tain­ing back­wards com­pat­i­bil­ity.

crypto/​mldsa im­ple­ments the post-quan­tum ML-DSA sig­na­ture scheme (FIPS 204), in­te­grated into crypto/​x509 and crypto/​tls.

uuid pro­vides na­tive sup­port for gen­er­at­ing and pars­ing UUIDs.

simd and ar­chi­tec­ture-spe­cific simd/​arch­simd pro­vide ex­per­i­men­tal SIMD sup­port.

net/​http/​httptest adds NewTestServer, pro­vid­ing an in-mem­ory fake net­work suit­able for use with the test­ing/​synctest pack­age.

Please read the Go 1.27 re­lease notes for the com­plete list of changes and de­tails.

Over the next few weeks, fol­low-up blog posts will cover some of the top­ics rel­e­vant to Go 1.27 in more de­tail. Check back later to read those posts.

Thanks to every­one who con­tributed to this re­lease by writ­ing code, fil­ing bugs, try­ing out ex­per­i­men­tal ad­di­tions, and test­ing re­lease can­di­dates. As al­ways, if you no­tice any prob­lems, please file an is­sue.

We hope you en­joy us­ing Go 1.27!

Remote workers report the highest well-being in study of 7,700 employees

www.colorado.edu

For years, many em­ploy­ers have wor­ried that work-from-home arrange­ments leave em­ploy­ees iso­lated, dis­con­nected from cowork­ers and more likely to leave their jobs.

But ac­cord­ing to a new study, re­mote work­ers are do­ing bet­ter than many em­ploy­ers re­al­ize.

Researchers an­a­lyzed sur­vey data from 7,704 em­ploy­ees at a large health­care or­ga­ni­za­tion. One pat­tern stood out: Employees who worked fully re­motely re­ported the high­est lev­els of well-be­ing, while those who worked en­tirely on­site re­ported the low­est. The study also found lit­tle ev­i­dence that re­mote work­ers felt less con­nected to col­leagues or work­place cul­ture.

This sug­gests you let peo­ple work re­motely if they want to work re­motely,” said Ste­fanie Johnson, pro­fes­sor of or­ga­ni­za­tional lead­er­ship and in­for­ma­tion an­a­lyt­ics at the Leeds School of Business and co-au­thor of the study, pub­lished in July 2026 in the jour­nal Fron­tiers in Psychology. Taking away peo­ple’s choice of how they work is prob­a­bly not go­ing to help them in terms of their well-be­ing.”

Stefanie Johnson

As com­pa­nies con­tinue to de­bate re­mote work, many lead­ers worry that em­ploy­ees need to be in the of­fice to stay con­nected, work well to­gether and re­main com­mit­ted to their or­ga­ni­za­tion. Johnson said the re­search does­n’t al­ways sup­port those con­cerns.

The data from our study and oth­ers sug­gest re­mote and hy­brid work re­sult in bet­ter out­comes than re­turn-to-of­fice man­dates,” she said. Leaders are not mak­ing de­ci­sions based on data. I think they are just re­turn­ing to what they are used to.”

Rethinking re­mote work

Johnson, who co-au­thored the study with Alyssa Lezcano, Stephanie Zajac and Courtney Holladay of the MD Anderson Leadership Institute in Houston, said the re­sults sur­prised her. She thought em­ploy­ees who split their time be­tween home and the of­fice might have the best of both worlds.

I ac­tu­ally thought you would be hap­pi­est if you were part time out of the of­fice,” she said. Then every once in a while you get to see peo­ple, get that hu­man con­nec­tion.”

Instead, the data pointed in a dif­fer­ent di­rec­tion.

Among em­ploy­ees in the study, well-be­ing was high­est for fully re­mote work­ers, fol­lowed by hy­brid em­ploy­ees and then on­site work­ers.

The find­ings also cast doubt on one of the main ar­gu­ments for bring­ing em­ploy­ees back to the of­fice: that peo­ple need to be to­gether in per­son to feel con­nected.

Employees in the study were asked to de­scribe their or­ga­ni­za­tion’s cul­ture in a hand­ful of words. Remote work­ers were slightly more likely than their hy­brid and on­site peers to use words as­so­ci­ated with team­work, in­clu­sion and sup­port.

People who are re­mote ac­tu­ally said more things that in­di­cated they had more pos­i­tive con­nec­tions, even though they were re­mote,” Johnson said.

Still, Johnson said face-to-face in­ter­ac­tion can play an im­por­tant role in help­ing cowork­ers build re­la­tion­ships, es­pe­cially if they are just start­ing out in their ca­reers.

Remote works bet­ter af­ter you know peo­ple,” she said. So there is still a ben­e­fit of hav­ing some face time.”

Staying power

The re­searchers also ex­am­ined em­ployee turnover one year af­ter the sur­vey was com­pleted.

They found that em­ploy­ees with higher well-be­ing were less likely to leave the or­ga­ni­za­tion. Work lo­ca­tion it­self was not a strong di­rect pre­dic­tor of turnover. Instead, re­mote work was as­so­ci­ated with higher well-be­ing, which in turn was as­so­ci­ated with lower turnover.

It makes sense. If you have higher well-be­ing, you’re less likely to leave your job,” Johnson said.

Participants com­pleted the work­place sur­vey in 2023, and re­searchers com­pared those re­sponses with ac­tual turnover records one year later. Of the em­ploy­ees sur­veyed, roughly half worked on­site, with the re­main­der split be­tween hy­brid and fully re­mote arrange­ments.

Flexibility mat­ters

The study did not ex­plore why re­mote work­ers re­ported higher well-be­ing, but Johnson points to a grow­ing body of re­search on au­ton­omy and flex­i­bil­ity.

One ex­pla­na­tion is that re­mote work­ers have greater con­trol over their work setup and daily sched­ule, she said.

If you have con­trol over your en­vi­ron­ment, you tend to have more pos­i­tive out­comes,” she said.

Working from home can also elim­i­nate many every­day stres­sors.

Spending a lot of time in traf­fic is neg­a­tively re­lated to well-be­ing,” Johnson said. There are so many lit­tle stres­sors as­so­ci­ated with be­ing in the of­fice.”

Those stres­sors can in­clude ar­rang­ing child care, hir­ing help for pets or man­ag­ing the lo­gis­tics of get­ting to and from work, she said.

Johnson said the study points to a broader les­son for em­ploy­ers nav­i­gat­ing re­turn-to-of­fice de­bates. Rather than fo­cus­ing only on where em­ploy­ees work, or­ga­ni­za­tions may get bet­ter re­sults by in­vest­ing in em­ployee well-be­ing.

I think flex­i­bil­ity is here to stay,” she said.

Civic Hygiene

shkspr.mobi

Imagine, just for a mo­ment, that the Government wanted to keep a record of every­one’s sex­u­al­ity. They need to know this de­tailed de­mo­graphic data be­cause it will be highly use­ful in civic plan­ning. It will help them work out what pro­vi­sion needs to be made for sex­ual health ser­vices, how many chil­dren are likely to be born, how many schools to build, etc.

You trust the Government, you voted for them, you and your friends have noth­ing to hide with re­gards to your sex­u­al­ity.

But! Shock hor­ror! After cre­at­ing the data­base, the Government loses the elec­tion and the ho­mo­phobes at UKIP get in to power!

Now they have a data­base of every gay in the vil­lage, and can ha­rass then, try to cure” them, or make their lives a liv­ing hell.

Far fetched? Not re­ally. With Cameron’s inane web fil­ter­ing plan, the black boxes” in ISPs which can record every click you make, and the sell­ing of the your NHS de­tails to pri­vate par­ties, we’re in a sit­u­a­tion where a ma­li­cious gov­ern­ment could cause se­ri­ous dam­age to us.

The se­cu­rity ex­pert Bruce Schneier wrote a won­der­ful ar­ti­cle for CNN on how the ex­ist­ing sur­veil­lance state is lead­ing to dis­as­trous breaches of our pri­vate in­for­ma­tion. He con­cludes by say­ing:

It’s bad civic hy­giene to build tech­nolo­gies that could some­day be used to fa­cil­i­tate a po­lice state.

– Bruce Schneier on CNN

It’s bad civic hy­giene to build tech­nolo­gies that could some­day be used to fa­cil­i­tate a po­lice state.

– Bruce Schneier on CNN

We have to be care­ful that the ap­pa­ra­tus we build can­not eas­ily be mis­used for evil pur­poses. Sure, even an in­nocu­ous toaster can be weaponised if some­one is will­ing enough, but we should not fall into the trap of mak­ing sys­tems which can eas­ily be turned against the peo­ple.

It’s prob­a­bly sen­si­ble to build a data­base of which car be­longs to which owner - it has an im­por­tant civil use and would be hard to abuse (although not im­pos­si­ble).

Should we have a na­tional data­base of, say, re­li­gious be­liefs? Almost in­stinc­tively the an­swer is no. The mem­o­ries of fas­cist dic­ta­tors haunt our col­lec­tive con­scious­ness. We have seen count­less times how race and re­li­gious iden­tity be­come death penal­ties. We would­n’t coun­te­nance it.

Civic hy­giene is­n’t about say­ing we dis­trust our cur­rent gov­ern­ment - it’s about not trust­ing the next gov­ern­ment.

Access Denied

www.casio.com

Reference #18.8cc82c17.1787237488.138790c6

https://​er­rors.edge­suite.net/​18.8c­c82c17.1787237488.138790c6

AliExpress webpage keeping multipoint Bluetooth headphones active with WebAudio fingerprinting

blog.laserphile.com

Recently I ran into a strange prob­lem with my Bluetooth head­phones. They sup­port mul­ti­point Bluetooth au­dio, so they can be con­nected to my PC and phone at the same time. Normally the PC takes pri­or­ity play­ing au­dio, with my phone be­ing able to play au­dio when noth­ing is play­ing on the PC.

Usually I lis­ten to mu­sic on my phone but with no­ti­fi­ca­tions or youtrube play­ing through the PC, this works re­li­ably un­til I open an AliExpress page in Firefox or Chrome (other browsers untested).

Shortly af­ter load­ing the AliExpress home­page, au­dio from my phone would stop play­ing. Closing the AliExpress tab fixes it im­me­di­ately. Muting the tab/​fire­fox/​Win­dows does not help, and there is no vis­i­ble video, mu­sic, or other me­dia play­ing on the page.

This seemed sus­pi­cious enough to in­ves­ti­gate.

Looking for hid­den me­dia

My first thought was an au­to­play­ing prod­uct video or ad­ver­tise­ment, so I checked for the usual sus­pects:

<audio> and <video> el­e­ments

calls to HTMLMediaElement.play()

ac­tive Media Session meta­data

me­dia re­quests

em­bed­ded frames con­tain­ing me­dia

None of these showed any­thing use­ful. There were no au­dio or video el­e­ments, no me­dia play­back calls, and nav­i­ga­tor.me­di­aSes­sion.play­back­State re­mained none.

A clue was that the prob­lem did not be­gin im­me­di­ately. It ap­peared af­ter the page had been sit­ting idle for sev­eral sec­onds. I in­stru­mented the page be­fore load­ing it and watched the Web Audio API in­stead of only look­ing for con­ven­tional me­dia el­e­ments.

The ba­sic idea was to wrap the AudioContext con­struc­tor and record when­ever a page cre­ated an au­dio-pro­cess­ing con­text:

const OriginalAudioContext = win­dow.Au­dio­Con­text;

win­dow.Au­dio­Con­text = class ex­tends OriginalAudioContext {

con­struc­tor(…args) {

su­per(…args);

con­sole.log(“Au­dio­Con­text cre­ated”, {

state: this.state,

stack: new Error().stack

});

}

};

I also wrapped AudioNode.prototype.connect() so I could see whether any­thing was con­nected to the con­tex­t’s au­dio des­ti­na­tion.

That fi­nally found it, two hid­den au­dio con­texts!

During an idle cap­ture of the AliExpress home­page, the page cre­ated two AudioContext ob­jects. Both en­tered the run­ning state and both con­nected nodes to AudioContext.destination.

At the same time there were still:

zero <audio> or <video> el­e­ments

zero me­dia play() calls

no ac­tive Media Session

no au­di­ble sound

The con­struc­tor stack traces pointed to two scripts:

https://​as­sets.aliex­press-me­dia.com/​g/​AWSC/​uab/​1.140.0/​col­lina.jshttps://​as­sets.aliex­press-me­dia.com/​g/​AWSC/​fireyejs/​1.231.67/​fireyejs.js

https://​as­sets.aliex­press-me­dia.com/​g/​AWSC/​uab/​1.140.0/​col­lina.js

https://​as­sets.aliex­press-me­dia.com/​g/​AWSC/​fireyejs/​1.231.67/​fireyejs.js

The first con­text was cre­ated by col­lina.js, while the sec­ond came from fireyejs.js. Both sit un­der an AWSC di­rec­tory and ap­pear to be part of Alibaba’s browser se­cu­rity and anti-abuse tool­ing.

The scripts are ex­tremely ob­fus­cated, but enough names and op­er­a­tions sur­vive for AI to work out what the au­dio code is do­ing.

What the au­dio code does

Both scripts build a WebAudio graph re­sem­bling this:

Sawtooth os­cil­la­tor    -> AnalyserNode    -> ScriptProcessorNode    -> GainNode set to zero    -> AudioContext.destination

Sawtooth os­cil­la­tor

-> AnalyserNode

-> ScriptProcessorNode

-> GainNode set to zero

-> AudioContext.destination

The os­cil­la­tor gen­er­ates a known wave­form. The analyser mea­sures the re­sult af­ter it has passed through the browser’s au­dio im­ple­men­ta­tion, and the script reads fre­quency data from it.

The gain is set to zero, so the user should not hear any­thing. However, the graph is still con­nected to the sys­tem au­dio des­ti­na­tion. Connecting it to the des­ti­na­tion causes the browser to ac­tively process the graph, even though the fi­nal vol­ume is zero.

This is very dif­fer­ent from an au­to­play­ing video. There is no me­dia el­e­ment for the browser’s nor­mal tab mute con­trol to stop. As far as the page is con­cerned, it is per­form­ing live au­dio pro­cess­ing.

In my case, that ap­pears to have been enough for Firefox or Windows to keep the Bluetooth au­dio path ac­tive, pre­vent­ing my mul­ti­point head­phones from switch­ing cleanly back to the phone.

This looks like fin­ger­print­ing

The WebAudio test is not the only mea­sure­ment in these scripts. Inspection of the bun­dles found code that queries or mea­sures:

can­vas ren­der­ing and to­DataURL()

WebGL ren­derer in­for­ma­tion, ex­ten­sions, and shader pre­ci­sion

au­dio os­cil­la­tor and analyser out­put

screen and view­port di­men­sions

de­vice pixel ra­tio

hard­ware con­cur­rency and de­vice mem­ory

in­stalled browser plu­g­ins

sup­ported au­dio and video for­mats

WebRTC be­hav­iour

browser per­for­mance tim­ing

mouse, touch, fo­cus, and scroll events

de­vice mo­tion and ori­en­ta­tion

prop­er­ties com­monly as­so­ci­ated with browser au­toma­tion

There is also code for se­ri­al­is­ing and en­crypt­ing re­sults, mak­ing re­quests to Alibaba teleme­try ser­vices, and send­ing data with fetch() or send­Bea­con().

This is a fairly com­pre­hen­sive browser and de­vice fin­ger­print.

Audio fin­ger­print­ing works be­cause small dif­fer­ences in browser ver­sions, op­er­at­ing sys­tems, au­dio li­braries, and hard­ware can pro­duce slightly dif­fer­ent re­sults from the same gen­er­ated sig­nal. It is not nec­es­sar­ily enough to uniquely iden­tify a de­vice by it­self, but it be­comes much more use­ful when com­bined with can­vas, WebGL, hard­ware, tim­ing, and in­ter­ac­tion data.

I can­not see what AliExpress does with the re­sult­ing data af­ter it reaches their servers. It may be used as a per­sis­tent de­vice iden­ti­fier, but it could also be one in­put into a fraud or bot-de­tec­tion score.

Why AliExpress would want this

AliExpress has plenty of rea­sons to dis­tin­guish nor­mal shop­pers from au­to­mated or sus­pi­cious clients as well as track­ing users brows­ing habits. The site has to deal with ac­count takeovers, fake ac­counts, scrap­ing, au­to­mated pur­chas­ing, pay­ment fraud, re­view ma­nip­u­la­tion, and abuse of coupons or new-cus­tomer pro­mo­tions. They also like most large busi­nesses make use of large datasets of user be­hav­iour to bet­ter mar­ket prod­ucts and ser­vices.

Cookies are not es­pe­cially re­li­able for this pur­pose be­cause they can be cleared, copied, or re­placed. A fin­ger­print made from many in­de­pen­dent browser mea­sure­ments is harder to ma­nip­u­late con­sis­tently.

Interaction data can also help de­ter­mine whether a browser is con­trolled by a per­son or au­toma­tion. From AliExpress’s per­spec­tive, this could re­duce fraud and al­low trusted cus­tomers through with­out show­ing a CAPTCHA every few pages. (Not that Aliexpress shies away from their AI gen­er­ated CAPTCHAs)

Personally I do not want a shop­ping home­page silently ex­er­cis­ing my graph­ics, au­dio, WebRTC, hard­ware, and mo­tion APIs, etc, to track my be­hav­iours, es­pe­cially if it has such an an­noy­ing ef­fect as block­ing my mu­sic. Per­haps if AliExpress was­n’t block­ing my mu­sic I never would’ve looked into what the site was do­ing.

Blocking it with uBlock Origin

I tested block­ing the two iden­ti­fied script fam­i­lies. With both re­quests blocked, the AliExpress home­page con­tin­ued to ren­der and no AudioContext ob­jects or des­ti­na­tion con­nec­tions ap­peared dur­ing the con­trol cap­ture.

In Firefox, I use the of­fi­cial uBlock Origin ex­ten­sion by Raymond Hill. To block the scripts open the uBlock dash­board, se­lect My fil­ters, and add:

! AliExpress AWSC fin­ger­print­ing scripts||as­sets.aliex­press-me­dia.com/​g/​AWSC/​uab/*/​col­lina.js$script,do­main=aliex­press.com||as­sets.aliex­press-me­dia.com/​g/​AWSC/​fireyejs/*/​fireyejs.js$script,do­main=aliex­press.com

! AliExpress AWSC fin­ger­print­ing scripts

||as­sets.aliex­press-me­dia.com/​g/​AWSC/​uab/*/​col­lina.js$script,do­main=aliex­press.com

||as­sets.aliex­press-me­dia.com/​g/​AWSC/​fireyejs/*/​fireyejs.js$script,do­main=aliex­press.com

Click Apply changes, close any ex­ist­ing AliExpress tabs, and open the site again. Existing tabs need to be closed be­cause block­ing a script does not shut down an au­dio con­text that it has al­ready cre­ated.

These rules are de­lib­er­ately nar­row. They block only the two ob­served script fam­i­lies and only when re­quested by AliExpress. I would not be sur­prised if this stops work­ing in the fu­ture, I’ll cross that bridge when it comes to it.

Because these scripts ap­pear to be con­nected with anti-fraud sys­tems, block­ing them may cause ex­tra CAPTCHAs or prob­lems dur­ing lo­gin or check­out. So far the home­page and or­di­nary prod­uct brows­ing still work, but I would tem­porar­ily dis­able the rules if AliExpress re­fuses a le­git­i­mate lo­gin or pay­ment.

Why I am block­ing it

The anti-fraud use case is un­der­stand­able, but this im­ple­men­ta­tion has sev­eral prob­lems.

It runs on the gen­eral shop­ping home­page be­fore I per­form a sen­si­tive ac­tion. It col­lects a broad set of de­vice and be­hav­ioural mea­sure­ments, the im­ple­men­ta­tion is de­lib­er­ately dif­fi­cult to in­spect, and there is no vis­i­ble in­di­ca­tion that the page has started a live au­dio-pro­cess­ing graph.

It also pro­duced a very real hard­ware side ef­fect. A silent fin­ger­print­ing test was able to in­ter­fere with Bluetooth mul­ti­point switch­ing, while the browser’s mute con­trol did noth­ing.

If a hid­den an­a­lyt­ics or se­cu­rity fea­ture can take own­er­ship of an au­dio path strongly enough to change how ex­ter­nal hard­ware be­haves, block­ing it seems like a rea­son­able trade-off.

I also can­not prove how long AliExpress stores the fin­ger­print or whether it is used across other Alibaba prop­er­ties. The client code proves that ex­ten­sive fin­ger­print-like mea­sure­ments are col­lected and trans­mit­ted, but server-side re­ten­tion and iden­tity link­age are not vis­i­ble from the browser. Can you re­ally trust any­one to have your best in­ter­ests at heart?

TL;DR

The AliExpress home­page silently cre­ates two run­ning WebAudio graphs from heav­ily ob­fus­cated Alibaba se­cu­rity scripts. The graphs gen­er­ate and analyse a wave­form as part of a much larger browser fin­ger­print, then con­nect through a zero-gain node to the sys­tem au­dio des­ti­na­tion pre­vent­ing the user from hear­ing any­thing.

On my setup, this ap­pears to keep the PCs Bluetooth au­dio path ac­tive and pre­vents mul­ti­point head­phones from switch­ing back to a phone. Muting the tab does not fix it be­cause there is no con­ven­tional me­dia el­e­ment to mute.

Blocking col­lina.js and fireyejs.js with the two uBlock Origin rules above pre­vented the hid­den au­dio con­texts from be­ing cre­ated and means I can hap­pily lis­ten to my mu­sic with­out be­ing in­ter­rupted while brows­ing Aliexpress.

Feature Request: Support AGENTS.md.

github.com

Codex, Amp, Cursor, and oth­ers are start­ing to stan­dard­ize around AGENTS.md (https://​agents.md/) — a uni­fied Markdown file that cod­ing agents can use to un­der­stand a code­base.

By con­trast, CLAUDE.md feels too spe­cific to Claude Code. It does­n’t work as well when col­lab­o­rat­ing with other de­vel­op­ers who aren’t us­ing Claude Code.

Unsloth Dynamic 3.0 GGUFs | Unsloth Documentation

unsloth.ai

Basics

🦥Unsloth Dynamic 3.0 GGUFs

Unsloth Dynamic v3.0 is the next it­er­a­tion of our Dynamic quan­ti­za­tion and a ma­jor im­prove­ment over Dynamic v2.0.

Today, we’re re­leas­ing Qwen3.8 – 27B Dynamic v3.0 quants that de­liver >10% top-1% bet­ter ac­cu­racy at the same size com­pared to every other provider. This is an up­date of our first shared early pre­view ver­sion of Dynamic v3.0. The new 3.0 GGUFs work with most in­fer­ence en­gines in­clud­ing llama.cpp and Unsloth Desktop.

Dynamic v3.0 over­all pre­serves more model qual­ity while keep­ing the same size, with stronger re­sults across met­rics like Divergence-300 @32 and KL Divergence.

Also a huge thanks to all your sup­port! We saw over 5.1 mil­lion Unsloth Qwen3.8 down­loads in just 5 days!

Our new method­ol­ogy com­poses of many new fea­tures and im­prove­ments. We now use a much higher-qual­ity ima­trix cal­i­bra­tion dataset from di­verse sources. The dataset is re­fined for agen­tic cod­ing, chat, and mul­ti­lin­gual per­for­mance. We also im­proved layer se­lec­tion and in­tro­duced many more quan­ti­za­tion tech­niques to pre­serve as much model qual­ity as pos­si­ble.

We do not train on the ima­trix cal­i­bra­tion dataset, and we do NOT use QAT or QAD. Everything is done through post-train­ing quan­ti­za­tion. Our ima­trix file used is avail­able for the com­mu­nity to test, eval­u­ate, and use. We en­cour­age re­searchers and de­vel­op­ers to cre­ate vari­a­tions and fine-tunes of Qwen3.8 us­ing our Unsloth quants/​ima­trix. You can read our over­fit­ting analy­sis as well.

We also re­moved the MTP mod­ule from smaller quants un­der UD-Q2_K_XL (8.37GB and lower) to con­verse around 500MB of disk space - you can use the Q4_0 MTP sep­a­rate mod­ule if needed

We also re­moved the MTP mod­ule from smaller quants un­der UD-Q2_K_XL (8.37GB and lower) to con­verse around 500MB of disk space - you can use the Q4_0 MTP sep­a­rate mod­ule if needed

We also made some smaller UD-1bit quants with UD-IQ1_S be­ing 6.2GB (without MTP) which re­tain around 72% top-1% ac­cu­racy yet be­ing 89% smaller.

We also made some smaller UD-1bit quants with UD-IQ1_S be­ing 6.2GB (without MTP) which re­tain around 72% top-1% ac­cu­racy yet be­ing 89% smaller.

UD-Q2_K_XL is around +8% more ac­cu­rate on top-1% than the next best and it’s 9.83GB and man­aged to cre­ate a work­ing HTML pro­gram with 1 small JS bug - pre­vi­ously it would break.

UD-Q2_K_XL is around +8% more ac­cu­rate on top-1% than the next best and it’s 9.83GB and man­aged to cre­ate a work­ing HTML pro­gram with 1 small JS bug - pre­vi­ously it would break.

🔀 Divergence-300 @32

We gen­er­ally re­port top-1% ac­cu­racy like how for Kimi-K3 Dynamic 1-bit reaches ~78.9% top-1 ac­cu­racy while be­ing 62% smaller.” However top-1% is an argmax on 1 pre­dic­tion, so it’s not re­ally ef­fec­tive on gaug­ing ac­tual in­fer­ence.

We cre­ated a dataset of 300 held out ex­am­ples (NOT in cal­i­bra­tion dataset) from Terminal-Bench 2.1 + DeepSWE + Harbor + MathArena 2025 – 26 + non-Latin/​long-doc prompts and we did greedy argmax de­cod­ing for 32 to­kens for BF16 vs all quants and providers. See over­fit­ting analy­sis for more de­tails on over­fit­ting.

This al­lows us to gauge if there is over­fit­ting and if quant out­puts are sim­i­lar to BF16′s tra­jec­to­ries over mul­ti­ple to­kens. This is a bet­ter met­ric than top-1% ac­cu­racy since we ex­tend KLD top-1% to more like KLD top-1% at 32 to­kens.

🔀 KL Divergence Benchmarks

We ran KLD bench­marks for all providers as well and re­port Top-1% and KLD mean. At all lev­els es­pe­cially on the smaller quant sizes, Unsloth UD-3 quants get up to +10% ex­tra top-1% ac­cu­racy at the same disk space!

All plots re­move the MTP head from the x axis when cal­cu­lat­ing disk space to pro­vide a fair com­par­i­son to every­one.

🕊️Not Overfitting

When com­par­ing to our older UD-2 on un­seen Wikitext and Code, we show great im­prove­ment on KLD - the big­ger ones not so much, so we still use our old UD-2 for the larger quants - we plan to ex­per­i­ment and im­prove them as well!

We also con­trol for over­fit­ting by us­ing to­tally dif­fer­ent datasets for cal­i­bra­tion and re­move all leak­ages as much as pos­si­ble. We test KLD on these un­seen datasets, and also we do NOT do QAD / QAT, just pure PTQ so over­fit­ting is less of a con­cern vs other QAD / QAT ap­proaches.

Similarly 🔀 Divergence-300 @32 uses an un­seen dataset of 300 prompts from DeepSWE, Terminal Bench and oth­ers, and acts as an­other dataset to gauge over­fit­ting - and shows our new UD-3 meth­ods do not over­fit.

Dynamic v2.0 (Old)

We’re in­tro­duc­ing Unsloth Dynamic v2.0 quan­ti­za­tion - a ma­jor up­grade to our pre­vi­ous quants. This new method out­per­forms lead­ing quan­ti­za­tion meth­ods and sets new bench­marks for Aider Polyglot, 5-shot MMLU and KL Divergence.

This means you can now run + fine-tune quan­tized LLMs while pre­serv­ing as much ac­cu­racy as pos­si­ble! You can run the 2.0 GGUFs on most in­fer­ence en­gines like llama.cpp, Unsloth Studio etc.

Sept 10, 2025 up­date: You asked for tougher bench­marks, so here’s Aider Polyglot re­sults! Our Dynamic 3-bit DeepSeek V3.1 GGUF scores 75.6%, sur­pass­ing many full-pre­ci­sion SOTA LLMs. Read more.

You can also view real-world use-case bench­marks con­ducted by Benjamin Marie for LiveCodeBench v6, MMLU Pro etc.:

You can see how Unsloth’s GGUFs per­forms bet­ter than the non-Un­sloth quants de­spite be­ing ~8GB smaller.

Detailed analy­sis of our bench­marks and eval­u­a­tion fur­ther be­low.

💡 What’s New in Dynamic v2.0?

Revamped Layer Selection for GGUFs + safeten­sors: Unsloth Dynamic 2.0 now se­lec­tively quan­tizes lay­ers much more in­tel­li­gently and ex­ten­sively. Rather than mod­i­fy­ing only se­lect lay­ers, we now dy­nam­i­cally ad­just the quan­ti­za­tion type of every pos­si­ble layer, and the com­bi­na­tions will dif­fer for each layer and model.

Revamped Layer Selection for GGUFs + safeten­sors: Unsloth Dynamic 2.0 now se­lec­tively quan­tizes lay­ers much more in­tel­li­gently and ex­ten­sively. Rather than mod­i­fy­ing only se­lect lay­ers, we now dy­nam­i­cally ad­just the quan­ti­za­tion type of every pos­si­ble layer, and the com­bi­na­tions will dif­fer for each layer and model.

Current se­lected and all fu­ture GGUF up­loads will uti­lize Dynamic 2.0 and our new cal­i­bra­tion dataset. The dataset con­tains more than >1.5M to­kens (depending on model) and com­prise of high-qual­ity, hand-cu­rated and cleaned data - to greatly en­hance con­ver­sa­tional chat per­for­mance.

Current se­lected and all fu­ture GGUF up­loads will uti­lize Dynamic 2.0 and our new cal­i­bra­tion dataset. The dataset con­tains more than >1.5M to­kens (depending on model) and com­prise of high-qual­ity, hand-cu­rated and cleaned data - to greatly en­hance con­ver­sa­tional chat per­for­mance.

Previously, our Dynamic quan­ti­za­tion (DeepSeek-R1 1.58-bit GGUF) was ef­fec­tive only for MoE ar­chi­tec­tures. Dynamic 2.0 quan­ti­za­tion now works on all mod­els (including MOEs & non-MoEs).

Previously, our Dynamic quan­ti­za­tion (DeepSeek-R1 1.58-bit GGUF) was ef­fec­tive only for MoE ar­chi­tec­tures. Dynamic 2.0 quan­ti­za­tion now works on all mod­els (including MOEs & non-MoEs).

Model-Specific Quants: Each model now uses a cus­tom-tai­lored quan­ti­za­tion scheme. E.g. the lay­ers quan­tized in Gemma 3 dif­fer sig­nif­i­cantly from those in Llama 4.

Model-Specific Quants: Each model now uses a cus­tom-tai­lored quan­ti­za­tion scheme. E.g. the lay­ers quan­tized in Gemma 3 dif­fer sig­nif­i­cantly from those in Llama 4.

To max­i­mize ef­fi­ciency, es­pe­cially on Apple Silicon and ARM de­vices, we now also add Q4_NL, Q5.1, Q5.0, Q4.1, and Q4.0 for­mats.

To max­i­mize ef­fi­ciency, es­pe­cially on Apple Silicon and ARM de­vices, we now also add Q4_NL, Q5.1, Q5.0, Q4.1, and Q4.0 for­mats.

To en­sure ac­cu­rate bench­mark­ing, we built an in­ter­nal eval­u­a­tion frame­work to match of­fi­cial re­ported 5-shot MMLU scores of Llama 4 and Gemma 3. This al­lowed ap­ples-to-ap­ples com­par­isons be­tween full-pre­ci­sion vs. Dynamic v2.0, QAT and stan­dard ima­trix GGUF quants.

All fu­ture GGUF up­loads will uti­lize Unsloth Dynamic 2.0, and our Dynamic 4-bit safe ten­sor quants will also ben­e­fit from this in the fu­ture.

📊 Why KL Divergence?

Accuracy is Not All You Need show­cases how prun­ing lay­ers, even by se­lect­ing un­nec­es­sary ones still yields vast dif­fer­ences in terms of flips”. A flip” is de­fined as an­swers chang­ing from in­cor­rect to cor­rect or vice versa. The pa­per shows how MMLU might not de­crease as we prune lay­ers or do quan­ti­za­tion,but that’s be­cause some in­cor­rect an­swers might have flipped” to be­come cor­rect. Our goal is to match the orig­i­nal model, so mea­sur­ing flips” is a good met­ric.

KL Divergence should be one of the gold stan­dards for re­port­ing quan­ti­za­tion er­rors as per the re­search pa­per Accuracy is Not All You Need”. Using per­plex­ity is in­cor­rect since out­put to­ken val­ues can can­cel out, so we must use KLD or harder bench­marks like Aider.

The pa­per also shows that in­ter­est­ingly KL Divergence is highly cor­re­lated with flips, and so our goal is to re­duce the mean KL Divergence whilst in­creas­ing the disk space of the quan­ti­za­tion as less as pos­si­ble.

⚖️ Calibration Dataset Overfitting

Most frame­works re­port per­plex­ity and KL Divergence us­ing a test set of Wikipedia ar­ti­cles. However, we no­ticed us­ing the cal­i­bra­tion dataset which is also Wikipedia re­lated causes quants to over­fit, and at­tain lower per­plex­ity scores. We uti­lize Calibration_v3 and Calibration_v5 datasets for fair test­ing which in­cludes some wiki­text data amongst other data. Also in­struct mod­els have unique chat tem­plates, and us­ing text only cal­i­bra­tion datasets is not ef­fec­tive for in­struct mod­els (base mod­els yes). In fact most ima­trix GGUFs are typ­i­cally cal­i­brated with these is­sues. As a re­sult, they nat­u­rally per­form bet­ter on KL Divergence bench­marks that also use Wikipedia data, since the model is es­sen­tially op­ti­mized for that do­main.

To en­sure a fair and con­trolled eval­u­a­tion, we do not to use our own cal­i­bra­tion dataset (which is op­ti­mized for chat per­for­mance) when bench­mark­ing KL Divergence. Instead, we con­ducted tests us­ing the same stan­dard Wikipedia datasets, al­low­ing us to di­rectly com­pare the per­for­mance of our Dynamic 2.0 method against the base­line ima­trix ap­proach.

🔢 MMLU Replication Adventure

Replicating MMLU 5 shot was night­mar­ish. We could not repli­cate MMLU re­sults for many mod­els in­clud­ing Llama 3.1 (8B) Instruct, Gemma 3 (12B) and oth­ers due to sub­tle im­ple­men­ta­tion is­sues. Llama 3.1 (8B) for ex­am­ple should be get­ting ~68.2%, whilst us­ing in­cor­rect im­ple­men­ta­tions can at­tain 35% ac­cu­racy.

Replicating MMLU 5 shot was night­mar­ish. We could not repli­cate MMLU re­sults for many mod­els in­clud­ing Llama 3.1 (8B) Instruct, Gemma 3 (12B) and oth­ers due to sub­tle im­ple­men­ta­tion is­sues. Llama 3.1 (8B) for ex­am­ple should be get­ting ~68.2%, whilst us­ing in­cor­rect im­ple­men­ta­tions can at­tain 35% ac­cu­racy.

Llama 3.1 (8B) Instruct has a MMLU 5 shot ac­cu­racy of 67.8% us­ing a naive MMLU im­ple­men­ta­tion. We find how­ever Llama to­k­enizes A” and _A” (A with a space in front) as dif­fer­ent to­ken ids. If we con­sider both spaced and non spaced to­kens, we get 68.2% (+0.4%)

Llama 3.1 (8B) Instruct has a MMLU 5 shot ac­cu­racy of 67.8% us­ing a naive MMLU im­ple­men­ta­tion. We find how­ever Llama to­k­enizes A” and _A” (A with a space in front) as dif­fer­ent to­ken ids. If we con­sider both spaced and non spaced to­kens, we get 68.2% (+0.4%)

Interestingly Llama 3 as per Eleuther AIs LLM Harness also ap­pends The best an­swer is” to the ques­tion, fol­low­ing Llama 3′s orig­i­nal MMLU bench­marks.

Interestingly Llama 3 as per Eleuther AIs LLM Harness also ap­pends The best an­swer is” to the ques­tion, fol­low­ing Llama 3′s orig­i­nal MMLU bench­marks.

There are many other sub­tle is­sues, and so to bench­mark every­thing in a con­trolled en­vi­ron­ment, we de­signed our own MMLU im­ple­men­ta­tion from scratch by in­ves­ti­gat­ing github.com/​hendrycks/​test di­rectly, and ver­i­fied our re­sults across mul­ti­ple mod­els and com­par­ing to re­ported num­bers.

There are many other sub­tle is­sues, and so to bench­mark every­thing in a con­trolled en­vi­ron­ment, we de­signed our own MMLU im­ple­men­ta­tion from scratch by in­ves­ti­gat­ing github.com/​hendrycks/​test di­rectly, and ver­i­fied our re­sults across mul­ti­ple mod­els and com­par­ing to re­ported num­bers.

✨ Gemma 3 QAT Replication, Benchmarks

The Gemma team re­leased two QAT (quantization aware train­ing) ver­sions of Gemma 3:

Q4_0 GGUF - Quantizes all lay­ers to Q4_0 via the for­mula w = q * block­_s­cale with each block hav­ing 32 weights. See llama.cpp wiki for more de­tails.

Q4_0 GGUF - Quantizes all lay­ers to Q4_0 via the for­mula w = q * block­_s­cale with each block hav­ing 32 weights. See llama.cpp wiki for more de­tails.

We bench­marked all Q4_0 GGUF ver­sions, and did ex­ten­sive ex­per­i­ments on the 12B model. We see the 12B Q4_0 QAT model gets 67.07% whilst the full bfloat16 12B ver­sion gets 67.15% on 5 shot MMLU. That’s very im­pres­sive! The 27B model is mostly nearly there!

MMLU 5 shot

26.12%

55.13%

67.07% (67.15% BF16)

70.64% (71.5% BF16)

Disk Space

0.93GB

2.94GB

7.52GB

16.05GB

Efficiency*

1.20

10.26

5.59

2.84

We de­signed a new Efficiency met­ric which cal­cu­lates the use­ful­ness of the model whilst also tak­ing into ac­count its disk size and MMLU 5 shot score:

We have to mi­nus 25 since MMLU has 4 mul­ti­ple choices - A, B, C or D. Assume we make a model that sim­ply ran­domly chooses an­swers - it’ll get 25% ac­cu­racy, and have a disk space of a few bytes. But clearly this is not a use­ful model.

On KL Divergence vs the base model, be­low is a table show­cas­ing the im­prove­ments. Reminder the closer the KL Divergence is to 0, the bet­ter (ie 0 means iden­ti­cal to the full pre­ci­sion model)

IQ1_S

1.035688

5.83

0.972932

6.06

IQ1_M

0.832252

6.33

0.800049

6.51

IQ2_XXS

0.535764

7.16

0.521039

7.31

IQ2_M

To add this web app to your iOS home screen tap the share button and select "Add to the Home Screen".

10HN is also available as an iOS App

If you visit 10HN only rarely, check out the the best articles from the past week.

Visit pancik.com for more.