Video yükleniyor...
Video Yüklenemedi
GOOGLE JUST MADE EVERY CHATBOT FEEL SLOW Diffusion Gemma 26b doesn’t predict word by word, it generates 256 tokens in parallel using bi-directional attention, like stable diffusion but for language it’s MoE so only 3.8B params activate during inference, fits on a single RTX 4090 with 18GB VRAM and... show more
27,385 görüntüleme • 3 ay önce •via X (Twitter)
26 Yorum

>and you can run it right now via llama.cpp Show us llama-server then :P Still waiting for full support. But you can run it full on vllm now:

There’s more to a chatbot than speed. Faster slop is still slop.

Will this run on a 64GB Macbook Pro Max M3

256 tokens in parallel is the part to watch; latency wins only matter if the output stays steerable.

Yes, it’s faster, but it’s also dumber than the base 26b a4b qat right now. Let’s hope they can close this gap.

great post, thanks!

Generating 256 tokens in parallel really shifts the latency game for apps. The MoE aspect helps with the practical cost of running this on consumer hardware too.

When a model this capable fits on consumer hardware, adoption accelerates fast.

The parallel token generation shifts the bottleneck from inference speed to VRAM bandwidth, making local deployment more viable for complex workflows.

People underestimate how much faster responses change the entire experience of using AI

chatbots waiting word by word may feel old soon

I like how the Gemma model line is evolving

Top

generates 256 tokens in parallel? It's awesome

Only valuable if the quality is actually usable…which it’s not. Google admits this. Once enough people test this I anticipate they will put out a better quality version

huh🤔

His speed is amazing

208ms per step over 97 steps is still not instant curious how quality holds when you push past the entropy bound on longer outputs

Banger

parallel token generation feels like the next obvious step

everything really seems slow now

local inference speed just made cloud providers nervous

18gb vram is the new baseline everyone skipped past

the crazy efficiency

So the idea is to be dumb but as fast as possible so you can figure out it’s no use of it faster on move to next on the hype

256 tokens in parallel is wild. but here's the dirty benchmark nobody talks about: how does it actually feel when your system is under load? we run local models on a 4090 that's also gaming, a mac studio that's also running home automation, and a mesh of pis doing mqtt. clean benchmarks look great until you're streaming while inferring. the moe architecture (3.8b active on a 26b model) should handle that load splitting beautifully — but real-world numbers under system stress are what actually matter for daily use.

