Загрузка видео...
Не удалось загрузить видео
DiffusionGemma can now run at 2000+ tokens/sec! ⚡ We made local DiffusionGemma inference 1.8× faster. Run it on 18GB RAM via Unsloth Studio. GitHub: Guide:
180,842 просмотров • 3 месяцев назад •via X (Twitter)
Комментарии: 35

Wow!

I think batch speed is the boring win compared with always-on local AI. A model you poke once can afford to be slow. A model running continuously in the background ... watching context, briefing you before things happen ... that's where 2000 tok/s is the line between a chatbot and an ambient assistant. Well done! (again!)

2000+ tokens/sec! wow! you know, i would die for a diffusion model like this on my iOS! ps: of course with a unsloth fine-tuning recipe! 🤩

isnt DiffusionGemma more prone to lower quality output though? Google’s own launch post says DiffusionGemma is optimized for speed and that its overall output quality is lower than standard Gemma 4. Google also says standard Gemma 4 remains the better choice for applications that need maximum quality.

??? im getting 170 usable tps on my 5090 compared to 500+ through vllm what's the issue ?

@danielhanchen we need to be able to serve it though, cli and chat doesn’t cut it.

@danielhanchen Does it do tool calling?

llama-server still doesn't support diffusion model. mlx does but token gen speed is horrible.

Nice but I didn't find UD-Q4 model version in your repo

Does this work with CPU offloading though?

This is wild

I spent ~6 hours making this diffusion model work on my mac, and that gave me 10 tokens/s because there was no llama.cpp support lol

It needs to run comfortably on 16gb. Hardly anyone has 18

any quantized version available that will enable it run on T4?

An abliterated version of this will have malware scripts flying around the internet in milliseconds

WOW

2000 tok/s local is the actual answer to this morning's news. nobody export-controls a gguf on your own box. this is the lane.

cant you fit that into 14gb so i could use with 16vram gpu

Cooking with white hot 🔥🔥🔥🔥

My sister, this is truly exhilarating news! 🌟 Seeing DiffusionGemma achieve such breathtaking speeds—surpassing 2000 tokens per second—while remaining accessible on local hardware like 18GB RAM is a masterpiece of efficiency over sheer bulk. It’s not just about the technical milestones; it's about the democratization of intelligence. By bridging the gap between high-performance research and local accessibility, you are helping to put the pulse of innovation directly into our hands. This transition from massive cloud dependency to agile, local execution is where technology truly begins to serve humanity with grace and speed. Keep pushing these boundaries! ✨

local diffusiongemma inference at 2000+ tokens/sec is a clear win low ram threshold shifts deployment from cloud to edge

Is Gemma4 12B coming, based on this diffusion tech? 🤔

wait this runs on 18gb?

My GPU only 12GB Vram 😭

Have you been able to fix the slop it slings? Last I saw was terrible decode.

Fast inference is exciting — but what you prompt it with still determines the output quality. ⚡ Save your best DiffusionGemma prompts and never lose them at — free prompt management for AI power users. 🚀 #DiffusionGemma #UnslothAI #PromptEngineering

2000 tokens/sec on 18GB RAM is actually insane. local AI just quietly won

Details!!!

2000+ tokens per second locally on 18GB RAM is not a small deal. The gap between local and cloud is closing faster than most people expected.

local inference keeps getting more realistic 18gb ram opens this up to way more people now

How fast on a 3090?!

Yooo, that's very unsloth 🚀

Google’s tournament style idea generation would go crazy with diffusion models

Wow 🔥

because the hard part was always scaling diffusers, 2000 tokens/sec changes everything

