Video wird geladen...
Video konnte nicht geladen werden
Before Fable 5 was shut down, it pushed Gemma 4 to 255 tok/s on WebGPU. Some didn't believe it was real. Today we're releasing the demo and kernels it wrote for you to see yourself. Run it locally in your browser. Agentic kernel optimization is the future of on-device inference
485,900 Aufrufe • vor 3 Monaten •via X (Twitter)
35 Kommentare

In case you hadn't noticed, we're working on something big. Stay tuned. 🔗 Link to the demo:

it doesn't know...

I used GPT 5.5 to take 4x Intel B70s running Minimax m2.7 from 14 to over 100 tok/s (decode rate). It took two weeks of 24/7 auto research however. That's a lot of tokens!! Fable and GPT Pro can do it much faster. Even GLM 5.2 can do it I'm finding though; just slowly.

Failed to load: No supported WebGPU variant for com.xenova.gemma4.DecodeOprojNorm; rejected fused_rows: when guard resolved to false; fused: when guard resolved to false

How many GB will my browser load if i access the page?

That one went to 500 tok/sec

I'm much more interested if the output is still correct, quality normally deteriorates when speeds increases

As far as I can tell it's fast and not very high quality. So interesting technical work, but I wouldn't use the model for anything.

Not working at all for me...

I'd love to work on a challenge like this ala

@huggingface Really well done. Have to exchange notes with you. E2B?

bro retired too early!

WebGPU: Hardware accelerated Adapter selected with `powerPreference: "high-performance"`: ```js { vendor: "nvidia", architecture: "ampere", subgroupMinSize: 32, subgroupMaxSize: 128, features: [ "shader-f16", "subgroups", "timestamp-query", ... ] } ``` GPU: NVIDIA GeForce RTX 2050 Likely cause The embedded `DecodeOprojNorm` variant guard appears to require an exact fixed subgroup range: ```js device.features.has("subgroups") && device.adapterInfo.subgroupMinSize == 32 && device.adapterInfo.subgroupMaxSize == 32 ``` On this NVIDIA Ampere/D3D12 adapter, Chrome reports: ```js subgroupMinSize: 32 subgroupMaxSize: 128 So both `fused_rows` and `fused` variants are rejected before compilation, even though the adapter supports `subgroups` and `shader-f16`. Suggested fix Please add a compatible fallback or relax/add a variant for adapters where subgroup size includes 32 but `subgroupMaxSize > 32`, e.g. NVIDIA/D3D12. If the WGSL is safe for 32-lane subgroup assumptions, the guard might be closer to: ```js subgroupMinSize <= 32 && subgroupMaxSize >= 32 Otherwise, a separate NVIDIA/D3D12 variant or non-fixed-subgroup fallback would allow the demo to run on hardware-backed WebGPU adapters that expose a subgroup range rather than fixed 32. ---

255 tok/s on gemma 4 in browser is wild if it holds up. what's the model size they're running and is this with prefill + decode or just decode

Hmmm

Claude Mythos leak update:

index.html:1210 Error: No supported WebGPU variant for com.xenova.gemma4.DecodeOprojNorm; rejected fused_rows: when guard resolved to false; fused: when guard resolved to false at Mn.selectVariant (gemma-4-e2b.js:18:3407) at Mn.selectVariantAndScope (gemma-4-e2b.js:18:3171) at Mn.prepare (gemma-4-e2b.js:17:22996) at Xe.add (gemma-4-e2b.js:5161:83352) at Xe.oprojNorm (gemma-4-e2b.js:5161:89523) at (gemma-4-e2b.js:5161:103613) at async #s (gemma-4-e2b.js:5161:116283) at async #u (gemma-4-e2b.js:5161:117692) at async e.streamTokenIdsFromCache (gemma-4-e2b.js:5161:117018) at async e.warmup (gemma-4-e2b.js:5161:121223) loadModel @ index.html:1210

Fable 5 hit 255 tok/s Gemma 4 on WebGPU before shutdown. Demo + kernels are live run it in your browser

Root-caused the Windows/NVIDIA gibberish (#1/#8): QatMatMul uses a bare subgroupAdd where the rest use the sgExact32 butterfly guard. D3D12 miscompiles it after a divergent store. 217 tok/s clean on a 5070. Fix + 2s reproducer:

does not work... firefox or chrome, webgpu variant error.

255 tok/s on webgpu with gemma 4 is the milestone that separates a keynote from something i can rerun. i live in mlx/llama.cpp daily; open kernels beat another trust me bro demo every time.

Well, it's definitely fast... that is all.

255 tok/s on webgpu with gemma 4 is the milestone that separates a keynote from something i can rerun. i live in mlx/llama.cpp daily; open kernels beat another "trust me bro" demo every time.

255 tok/s in browser is absurd.

255 tok/s in your browser. Fable 5 proved it, now you can run it. Agentic kernels = local AI unchained

still, my small mac can't handle it

Will try but I hope it did not just optimise for your GPU 😅

Fable 5 WebGPU'da Gemma 4'ü 255 token/s'ye taşıdı — tarayıcıda, yerel olarak. 🚀 İnanmayanlar için kod açık kaynak yapıldı. Kendiniz deneyin. Cihaz üzerinde çıkarım artık teori değil, gerçek. Bulut bağımlılığı bitiyor mu? 👇

need some work i think...

@_akhaliq Incredible work! Maybe I missed it but I couldn’t find the model size? How many parameters for these results?

@crosstensor thank you for releasing the work for peer review i respect your efforts now

Amazing one

255 tok/s on WebGPU is the number that should make everyone nervous about server-side inference. the margin on cloud inference narrows every time a benchmark like this ships

255 tok/s on WebGPU is wild, but the bigger story is the agent writing its own kernels. Curious: did Fable converge on shapes a human would write, or did it find unintuitive tilings you wouldn't have shipped by hand?

check your fancy landing page discussion board

