vLLM's banner
vLLM's profile picture

vLLM

@vllm_project46,592 subscribers

A high-throughput and memory-efficient inference and serving engine for LLMs. Join https://t.co/lxJ0SfX5pJ to discuss together with the community!

Shorts

🎉 Congrats to MiniMax (official) on releasing the open weights for MiniMax H3! Day-0 support in vLLM-Omni! One model reads text, images, video, and audio as a single context and returns video with native stereo audio. Text-to-video, first/last-frame, and multi-reference generation, 4 to 15 seconds at up to 2K and 24 FPS. The MP4 comes back with H.264 video and a synchronized stereo track already muxed in. It serves over the OpenAI-compatible /v1/videos endpoint, synchronous or async with job polling. 🔊 The video ⬇️ was made with H3, served on vLLM-Omni.

🎉 Congrats to MiniMax (official) on releasing the open weights for MiniMax H3! Day-0 support in vLLM-Omni! One model reads text, images, video, and audio as a single context and returns video with native stereo audio. Text-to-video, first/last-frame, and multi-reference generation, 4 to 15 seconds at up to 2K and 24 FPS. The MP4 comes back with H.264 video and a synchronized stereo track already muxed in. It serves over the OpenAI-compatible /v1/videos endpoint, synchronous or async with job polling. 🔊 The video ⬇️ was made with H3, served on vLLM-Omni.

127,714 Aufrufe

🎉 Congrats to Thinking Machines on TML Inkling—a 1T-parameter open-weight model supported in vLLM from Day 0. Highlights: • Natively multimodal across text, image, and audio • Up to 1M-token context • New architecture with relative attention, short convolutions, and MoE expert sinks • 8 MTP heads for speculative decoding vLLM supports both NVFP4 and BF16 checkpoints, optimized for NVIDIA Blackwell and Hopper, reaching up to 380 tok/s/user on 4× GB200 with MTP. Huge thanks to the Thinking Machines Lab team for the close collaboration 🙏 Read about implementation below 👇

🎉 Congrats to Thinking Machines on TML Inkling—a 1T-parameter open-weight model supported in vLLM from Day 0. Highlights: • Natively multimodal across text, image, and audio • Up to 1M-token context • New architecture with relative attention, short convolutions, and MoE expert sinks • 8 MTP heads for speculative decoding vLLM supports both NVFP4 and BF16 checkpoints, optimized for NVIDIA Blackwell and Hopper, reaching up to 380 tok/s/user on 4× GB200 with MTP. Huge thanks to the Thinking Machines Lab team for the close collaboration 🙏 Read about implementation below 👇

49,212 Aufrufe

🎉Announcing Gemma4 on vLLM model launch blog at Explore our detailed blogpost covering Gemma 4's capabilities, first-ever day-0 support across diverse hardware platforms, and ready-to-go deployment recipes! #Gemma4 #vLLM

🎉Announcing Gemma4 on vLLM model launch blog at Explore our detailed blogpost covering Gemma 4's capabilities, first-ever day-0 support across diverse hardware platforms, and ready-to-go deployment recipes! #Gemma4 #vLLM

12,090 Aufrufe

Videos

Keine weiteren Inhalte verfügbar