
Red Hat AI
@RedHat_AI • 11,373 subscribers
Accelerating AI innovation with open platforms and community. The future of AI is open.
Videos

Gemma 4 Diffusion landed in vLLM last week. Day 0. First diffusion LLM natively supported in vLLM. Instead of one token at a time, it predicts 256 tokens at once and iteratively denoises them in parallel. Result: 1,000+ tokens per second at batch size 1 on a single H100. Built on Model Runner V2. Google Gemma
Red Hat AI17,637 görüntüleme • 1 ay önce

What compression looks like on vLLM. Same Gemma 4 31B. Red Hat AI's quantized version runs at nearly 2x tokens/sec, half the memory, 99%+ accuracy retained. Open source. Quantized with LLM Compressor. Links in comments. 🙏 Sawyer Bowerman for the 2-minute demo.
Red Hat AI34,199 görüntüleme • 3 ay önce

Michael Goin (Michael Goin) walks through what's new in vLLM v0.17, v0.18, and v0.19 in ~8 minutes. Flash Attention 4, new performance modes, zero-bubble async scheduling, online MXFP4 quantization, Gemma 4, and a lot more. 1,592 commits. 682 contributors (163 new). 🎉 🚀
Red Hat AI23,115 görüntüleme • 3 ay önce

A full year of vLLM in 30 minutes by vLLM Lead from UC Berkeley, Simon Mo. Model and hardware usage trends, model architectures, API evolution, V1 engine rebuild, multimodal progress, expanding hardware support, and more. Plus how we are thinking about 2026. Enjoy!
Red Hat AI15,713 görüntüleme • 7 ay önce
Daha fazla içerik yok.