
Red Hat AI
@RedHat_AI • 11,373 subscribers
Accelerating AI innovation with open platforms and community. The future of AI is open.
Videos

Gemma 4 Diffusion landed in vLLM last week. Day 0. First diffusion LLM natively supported in vLLM. Instead of one token at a time, it predicts 256 tokens at once and iteratively denoises them in parallel. Result: 1,000+ tokens per second at batch size 1 on a single H100. Built on Model Runner V2. Google Gemma
Red Hat AI17,637 次观看 • 1 个月前
没有更多内容可加载