Загрузка видео...

Не удалось загрузить видео

На главную

DiffusionGemma can now run at 2000+ tokens/sec! ⚡ We made local DiffusionGemma inference 1.8× faster. Run it on 18GB RAM via Unsloth Studio. GitHub: Guide:

180,842 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 35

Фото профиля Philipp Schmid
Philipp Schmid3 месяцев назад

Wow!

Фото профиля Tery Emilson
Tery Emilson3 месяцев назад

I think batch speed is the boring win compared with always-on local AI. A model you poke once can afford to be slow. A model running continuously in the background ... watching context, briefing you before things happen ... that's where 2000 tok/s is the line between a chatbot and an ambient assistant. Well done! (again!)

Фото профиля Maziyar PANAHI
Maziyar PANAHI3 месяцев назад

2000+ tokens/sec! wow! you know, i would die for a diffusion model like this on my iOS! ps: of course with a unsloth fine-tuning recipe! 🤩

Фото профиля Apollo
Apollo3 месяцев назад

isnt DiffusionGemma more prone to lower quality output though? Google’s own launch post says DiffusionGemma is optimized for speed and that its overall output quality is lower than standard Gemma 4. Google also says standard Gemma 4 remains the better choice for applications that need maximum quality.

Фото профиля Terp
Terp3 месяцев назад

??? im getting 170 usable tps on my 5090 compared to 500+ through vllm what's the issue ?

Фото профиля Le TechLead🔰
Le TechLead🔰3 месяцев назад

@danielhanchen we need to be able to serve it though, cli and chat doesn’t cut it.

Фото профиля Xaden Ryan
Xaden Ryan3 месяцев назад

@danielhanchen Does it do tool calling?

Фото профиля Ankit Prateek
Ankit Prateek3 месяцев назад

llama-server still doesn't support diffusion model. mlx does but token gen speed is horrible.

Фото профиля Tarrito.rocks
Tarrito.rocks3 месяцев назад

Nice but I didn't find UD-Q4 model version in your repo

Фото профиля Dariton
Dariton3 месяцев назад

Does this work with CPU offloading though?

Фото профиля Ankit Prateek
Ankit Prateek3 месяцев назад

This is wild

Фото профиля Ankit Prateek
Ankit Prateek3 месяцев назад

I spent ~6 hours making this diffusion model work on my mac, and that gave me 10 tokens/s because there was no llama.cpp support lol

Фото профиля ibrand
ibrand3 месяцев назад

It needs to run comfortably on 16gb. Hardly anyone has 18

Фото профиля Piyush
Piyush3 месяцев назад

any quantized version available that will enable it run on T4?

Фото профиля AACeeert
AACeeert3 месяцев назад

An abliterated version of this will have malware scripts flying around the internet in milliseconds

Фото профиля Eric ⚡️ Building...
Eric ⚡️ Building...3 месяцев назад

WOW

Фото профиля Vabbyshabby
Vabbyshabby3 месяцев назад

2000 tok/s local is the actual answer to this morning's news. nobody export-controls a gguf on your own box. this is the lane.

Фото профиля Emircan ERKUL
Emircan ERKUL3 месяцев назад

cant you fit that into 14gb so i could use with 16vram gpu

Фото профиля mr_r0b0t
mr_r0b0t3 месяцев назад

Cooking with white hot 🔥🔥🔥🔥

Фото профиля Anis🐬Al
Anis🐬Al3 месяцев назад

My sister, this is truly exhilarating news! 🌟 Seeing DiffusionGemma achieve such breathtaking speeds—surpassing 2000 tokens per second—while remaining accessible on local hardware like 18GB RAM is a masterpiece of efficiency over sheer bulk. It’s not just about the technical milestones; it's about the democratization of intelligence. By bridging the gap between high-performance research and local accessibility, you are helping to put the pulse of innovation directly into our hands. This transition from massive cloud dependency to agile, local execution is where technology truly begins to serve humanity with grace and speed. Keep pushing these boundaries! ✨

Фото профиля Secta
Secta3 месяцев назад

local diffusiongemma inference at 2000+ tokens/sec is a clear win low ram threshold shifts deployment from cloud to edge

Фото профиля Pranav
Pranav3 месяцев назад

Is Gemma4 12B coming, based on this diffusion tech? 🤔

Фото профиля netrunner
netrunner3 месяцев назад

wait this runs on 18gb?

Фото профиля ArdanZ
ArdanZ3 месяцев назад

My GPU only 12GB Vram 😭

Фото профиля Robert Keyes
Robert Keyes3 месяцев назад

Have you been able to fix the slop it slings? Last I saw was terrible decode.

Фото профиля Kaustubh Joshi
Kaustubh Joshi3 месяцев назад

Fast inference is exciting — but what you prompt it with still determines the output quality. ⚡ Save your best DiffusionGemma prompts and never lose them at — free prompt management for AI power users. 🚀 #DiffusionGemma #UnslothAI #PromptEngineering

Фото профиля Sanjay
Sanjay3 месяцев назад

2000 tokens/sec on 18GB RAM is actually insane. local AI just quietly won

Фото профиля oriel haim
oriel haim3 месяцев назад

Details!!!

Фото профиля AI Mastery Guide
AI Mastery Guide3 месяцев назад

2000+ tokens per second locally on 18GB RAM is not a small deal. The gap between local and cloud is closing faster than most people expected.

Фото профиля Gerladina
Gerladina3 месяцев назад

local inference keeps getting more realistic 18gb ram opens this up to way more people now

Фото профиля Twon.
Twon.3 месяцев назад

How fast on a 3090?!

Фото профиля Thor 雷神 ⚡️
Thor 雷神 ⚡️3 месяцев назад

Yooo, that's very unsloth 🚀

Фото профиля Thomas Linden
Thomas Linden3 месяцев назад

Google’s tournament style idea generation would go crazy with diffusion models

Фото профиля Verma
Verma3 месяцев назад

Wow 🔥

Фото профиля Adel Bucetta
Adel Bucetta3 месяцев назад

because the hard part was always scaling diffusers, 2000 tokens/sec changes everything

Похожие видео

Auto regressive LLMs are officially on notice. run Gemma 4 26B diffusion gguf with llama.cpp Google just dropped DiffusionGemma-26B, and it completely flips how we generate text. instead of predicting words one by one, it generates 256 tokens in parallel using bi-directional attention. its like stable diffusion, but for language. the model starts with random text "noise" and iteratively refines and self-corrects the entire block in real-time to fix formatting and reasoning errors on the fly. since it’s a Mixture of Experts (MoE) that only activates 3.8B parameters during inference, it fits perfectly on consumer hardware. You can run the Q4_K_M quant with an 18GB VRAM budget on a single RTX 3090 or RTX 4090 with exceptional throughput. Tested on Ubuntu 22 with CUDA 13.1 using the cutting edge experimental llama.cpp branch. Here is how to compile and run it with the live terminal denoising visualizer: # 1. Clone & check out the experimental PR (#24423) - 1) git clone && cd llama.cpp -git fetch origin 2) pull/24423/head:diffusiongemma && --git checkout diffusiongemma # 2. Build with CUDA support 1) cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native 2) cmake --build build -j $(nproc) --config Release --target llama-diffusion-cli # 3. Run with live visual denoising (llama.cpp flags) ./build/bin/llama-diffusion-cli \ -m /path/to/diffusiongemma-26B-A4B-it-Q4_K_M.gguf \ -ngl 99 -cnv -n 2048 --diffusion-visual Watch the video below to see the live --diffusion-visual canvas iteratively de noising the prompt output in real time. guide and unsloth's hugging face GGUF model links are in the comments below! Is auto regressive generation officially legacy tech? Let me know what you think.

Alok

52,656 просмотров • 3 месяцев назад