Loading video...
Video Failed to Load
Gemma 4 Diffusion landed in vLLM last week. Day 0. First diffusion LLM natively supported in vLLM. Instead of one token at a time, it predicts 256 tokens at once and iteratively denoises them in parallel. Result: 1,000+ tokens per second at batch size 1 on a single H100.... show more
17,637 views • 1 month ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
