
Red Hat AI
@RedHat_AI • 12,621 subscribers
Accelerating AI innovation with open platforms and community. The future of AI is open.
Videos

Red Hat AI just shipped DFlash speculator checkpoints for two of NVIDIA AI's most powerful open models: → Nemotron Ultra 550B → Nemotron Super 120B On math and reasoning: ~5 out of 7 draft tokens accepted on average. On code (HumanEval): ~3.4 out of 7. Both checkpoints trained with the open source Speculators library from vLLM. Apache 2.0. Validated on NVIDIA B200. One flag to enable in vLLM: --spec-model RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash --spec-tokens 7 --spec-method dflash 🔗 Ultra 550B: 🔗 Super 120B:
Red Hat AI15,077 次观看 • 2 个月前

Gemma 4 Diffusion landed in vLLM last week. Day 0. First diffusion LLM natively supported in vLLM. Instead of one token at a time, it predicts 256 tokens at once and iteratively denoises them in parallel. Result: 1,000+ tokens per second at batch size 1 on a single H100. Built on Model Runner V2. Google Gemma
Red Hat AI17,670 次观看 • 3 个月前
没有更多内容可加载