Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Before Fable 5 was shut down, it pushed Gemma 4 to 255 tok/s on WebGPU. Some didn't believe it was real. Today we're releasing the demo and kernels it wrote for you to see yourself. Run it locally in your browser. Agentic kernel optimization is the future of on-device inference

485,900 Aufrufe • vor 3 Monaten •via X (Twitter)

35 Kommentare

Profilbild von Xenova
Xenovavor 3 Monaten

In case you hadn't noticed, we're working on something big. Stay tuned. 🔗 Link to the demo:

Profilbild von octalmage
octalmagevor 3 Monaten

it doesn't know...

Profilbild von Steve💙🇨🇦
Steve💙🇨🇦vor 3 Monaten

I used GPT 5.5 to take 4x Intel B70s running Minimax m2.7 from 14 to over 100 tok/s (decode rate). It took two weeks of 24/7 auto research however. That's a lot of tokens!! Fable and GPT Pro can do it much faster. Even GLM 5.2 can do it I'm finding though; just slowly.

Profilbild von The Singularity Project
The Singularity Projectvor 3 Monaten

Failed to load: No supported WebGPU variant for com.xenova.gemma4.DecodeOprojNorm; rejected fused_rows: when guard resolved to false; fused: when guard resolved to false

Profilbild von Fab
Fabvor 3 Monaten

How many GB will my browser load if i access the page?

Profilbild von Adria B.A.
Adria B.A.vor 3 Monaten

That one went to 500 tok/sec

Profilbild von xcaliburr
xcaliburrvor 3 Monaten

I'm much more interested if the output is still correct, quality normally deteriorates when speeds increases

Profilbild von Ian Danforth
Ian Danforthvor 3 Monaten

As far as I can tell it's fast and not very high quality. So interesting technical work, but I wouldn't use the model for anything.

Profilbild von Nathan Spencer
Nathan Spencervor 3 Monaten

Not working at all for me...

Profilbild von John Lussier
John Lussiervor 3 Monaten

I'd love to work on a challenge like this ala

Profilbild von Eyal Toledano
Eyal Toledanovor 3 Monaten

@huggingface Really well done. Have to exchange notes with you. E2B?

Profilbild von ansuman
ansumanvor 3 Monaten

bro retired too early!

Profilbild von The Singularity Project
The Singularity Projectvor 3 Monaten

WebGPU: Hardware accelerated Adapter selected with `powerPreference: "high-performance"`: ```js { vendor: "nvidia", architecture: "ampere", subgroupMinSize: 32, subgroupMaxSize: 128, features: [ "shader-f16", "subgroups", "timestamp-query", ... ] } ``` GPU: NVIDIA GeForce RTX 2050 Likely cause The embedded `DecodeOprojNorm` variant guard appears to require an exact fixed subgroup range: ```js device.features.has("subgroups") && device.adapterInfo.subgroupMinSize == 32 && device.adapterInfo.subgroupMaxSize == 32 ``` On this NVIDIA Ampere/D3D12 adapter, Chrome reports: ```js subgroupMinSize: 32 subgroupMaxSize: 128 So both `fused_rows` and `fused` variants are rejected before compilation, even though the adapter supports `subgroups` and `shader-f16`. Suggested fix Please add a compatible fallback or relax/add a variant for adapters where subgroup size includes 32 but `subgroupMaxSize > 32`, e.g. NVIDIA/D3D12. If the WGSL is safe for 32-lane subgroup assumptions, the guard might be closer to: ```js subgroupMinSize <= 32 && subgroupMaxSize >= 32 Otherwise, a separate NVIDIA/D3D12 variant or non-fixed-subgroup fallback would allow the demo to run on hardware-backed WebGPU adapters that expose a subgroup range rather than fixed 32. ---

Profilbild von Samian
Samianvor 3 Monaten

255 tok/s on gemma 4 in browser is wild if it holds up. what's the model size they're running and is this with prefill + decode or just decode

Profilbild von Jan Boon
Jan Boonvor 3 Monaten

Hmmm

Profilbild von Gerard Sans | Axiom 🇬🇧
Gerard Sans | Axiom 🇬🇧vor 3 Monaten

Claude Mythos leak update:

Profilbild von The Singularity Project
The Singularity Projectvor 3 Monaten

index.html:1210 Error: No supported WebGPU variant for com.xenova.gemma4.DecodeOprojNorm; rejected fused_rows: when guard resolved to false; fused: when guard resolved to false at Mn.selectVariant (gemma-4-e2b.js:18:3407) at Mn.selectVariantAndScope (gemma-4-e2b.js:18:3171) at Mn.prepare (gemma-4-e2b.js:17:22996) at Xe.add (gemma-4-e2b.js:5161:83352) at Xe.oprojNorm (gemma-4-e2b.js:5161:89523) at (gemma-4-e2b.js:5161:103613) at async #s (gemma-4-e2b.js:5161:116283) at async #u (gemma-4-e2b.js:5161:117692) at async e.streamTokenIdsFromCache (gemma-4-e2b.js:5161:117018) at async e.warmup (gemma-4-e2b.js:5161:121223) loadModel @ index.html:1210

Profilbild von Alice The Ai Expert
Alice The Ai Expertvor 3 Monaten

Fable 5 hit 255 tok/s Gemma 4 on WebGPU before shutdown. Demo + kernels are live run it in your browser

Profilbild von Kuldeep Singh
Kuldeep Singhvor 2 Monaten

Root-caused the Windows/NVIDIA gibberish (#1/#8): QatMatMul uses a bare subgroupAdd where the rest use the sgExact32 butterfly guard. D3D12 miscompiles it after a divergent store. 217 tok/s clean on a 5070. Fix + 2s reproducer:

Profilbild von bdvd 🇦🇷
bdvd 🇦🇷vor 3 Monaten

does not work... firefox or chrome, webgpu variant error.

Profilbild von Vabbyshabby
Vabbyshabbyvor 3 Monaten

255 tok/s on webgpu with gemma 4 is the milestone that separates a keynote from something i can rerun. i live in mlx/llama.cpp daily; open kernels beat another trust me bro demo every time.

Profilbild von Firworks
Firworksvor 3 Monaten

Well, it's definitely fast... that is all.

Profilbild von Vabbyshabby
Vabbyshabbyvor 3 Monaten

255 tok/s on webgpu with gemma 4 is the milestone that separates a keynote from something i can rerun. i live in mlx/llama.cpp daily; open kernels beat another "trust me bro" demo every time.

Profilbild von Frank ✈️ (📜,📜)
Frank ✈️ (📜,📜)vor 3 Monaten

255 tok/s in browser is absurd.

Profilbild von ZenithAi
ZenithAivor 3 Monaten

255 tok/s in your browser. Fable 5 proved it, now you can run it. Agentic kernels = local AI unchained

Profilbild von rakesh
rakeshvor 3 Monaten

still, my small mac can't handle it

Profilbild von Unni
Unnivor 3 Monaten

Will try but I hope it did not just optimise for your GPU 😅

Profilbild von usul365
usul365vor 3 Monaten

Fable 5 WebGPU'da Gemma 4'ü 255 token/s'ye taşıdı — tarayıcıda, yerel olarak. 🚀 İnanmayanlar için kod açık kaynak yapıldı. Kendiniz deneyin. Cihaz üzerinde çıkarım artık teori değil, gerçek. Bulut bağımlılığı bitiyor mu? 👇

Profilbild von bitslix
bitslixvor 3 Monaten

need some work i think...

Profilbild von Fred Terzi
Fred Terzivor 3 Monaten

@_akhaliq Incredible work! Maybe I missed it but I couldn’t find the model size? How many parameters for these results?

Profilbild von WuBu ⪋ WaefreBeorn 🇺🇸 👑
WuBu ⪋ WaefreBeorn 🇺🇸 👑vor 3 Monaten

@crosstensor thank you for releasing the work for peer review i respect your efforts now

Profilbild von Abdulmuiz Adeyemo
Abdulmuiz Adeyemovor 3 Monaten

Amazing one

Profilbild von Shubh Thorat
Shubh Thoratvor 3 Monaten

255 tok/s on WebGPU is the number that should make everyone nervous about server-side inference. the margin on cloud inference narrows every time a benchmark like this ships

Profilbild von Rizwan
Rizwanvor 3 Monaten

255 tok/s on WebGPU is wild, but the bigger story is the agent writing its own kernels. Curious: did Fable converge on shapes a human would write, or did it find unintuitive tilings you wouldn't have shipped by hand?

Profilbild von specimba
specimbavor 3 Monaten

check your fancy landing page discussion board

Ähnliche Videos