
LMSYS Org
@lmsysorg • 17,112 subscribers
Large Model Systems Organization: We developed SGLang @sgl_project (https://t.co/OjwQadINKU), Chatbot Arena (now @arena), and Vicuna!
Shorts
Videos

SGLang day-0 speed on Kimi K3: 423 tok/s (measured on gsm8k), plus RL support ready in Miles RadixArk! How the largest open-source model runs this fast: we natively implemented and deeply optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. We've passed Kimi Vendor Verifier and are ready for production! Thanks to Kimi.ai, NVIDIA, AMD, KVCache.AI, Modal, and Baseten for building this with us, and to Google Cloud, Nebius Token Factory, fal, DigitalOcean, Runpod, DeepInfra and GMI Cloud for serving K3 on SGLang. Blog, cookbook, benchmarks in the comments. P.S. This demo video? Kimi K3 made it itself. Play the game 👇
LMSYS Org203,615 просмотров • 1 месяц назад

MiniMax (official) H3 is live in SGLang Diffusion, with day-0 serving support 🎬 This open model matches Seedance 2.0 at 1/3 the cost, or $0 if you run it locally on 2x 5090 or 1 RTX 6000. With SGLang Diffusion, you can build visual concepts, motion design, e-commerce creatives, precise video edits, animation and stylized visuals, all locally on your own machine. No API bill, no waitlist. H3 takes text, image, video and audio in a single context, and generates 5-15s clips at native 2K, 24fps with native stereo audio. SGLang Diffusion also runs on NVIDIA AI Blackwell and Hopper, and AMD MI355X / MI300X. Let's create something with MiniMax H3 👇
LMSYS Org98,556 просмотров • 24 дней назад

Inkling, Thinking Machines' first open model, dropped today: 975B total / 41B active MoE, up to 1M context, reasoning natively over text, images, and audio. Serving and RL support are already live: you can run and shape it on an open stack, starting now. Day 0 support on SGLang SGLang and Miles RadixArk👇 - Inkling's new architecture (ShortConv, attention with relative positional embedding, shared expert sink MoE) is natively implemented and deeply optimized, with prefill full CUDA graph and MXFP8 KV cache - Full parameter and LoRA RL in a customized Megatron backend, train inference consistency via customized kernels, routing replay, and cross-runtime parameter synchronization - DFlash speculative decoding from Modal for low-latency serving Launching now, blog and cookbook in the comments ⬇️
LMSYS Org145,979 просмотров • 1 месяц назад

SGLang now supports DSpark, enabling confidence-driven, variable-length verification for speculative decoding 🎉 DSpark addresses a key bottleneck under load: instead of verifying every draft token, it verifies only where the draft model is confident, so the gains hold even as batch size scales. We heavily optimized variable-length verification in SGLang. Across batch sizes 1 to 256, DSpark gives the best throughput/latency tradeoff on DeepSeek-V4-Flash, ahead of both MTP and non-spec. At high concurrency, dynamic scheduling provides up to ~20% higher throughput compared to a fixed budget, while maintaining high verification quality across workloads. With fused kernels and zero-overhead scheduling, DeepSeek-V4-Pro reaches 383.7 tok/s at B=1 on B300. DSpark is now available in SGLang with support for Qwen3 and DeepSeek-V4. Thanks DeepSeek for open-sourcing! Blog with full technical details and commands to run below 👇
LMSYS Org170,098 просмотров • 1 месяц назад

Serving GLM5.2 NVFP4 Agentic Workload with SGLang: How We Reached 500 TPS on 8xB300 at bs=1 In this deep dive, we explain how SGLang reaches 500+ tok/s/user at bs=1 on 8xB300, with 18 to 34% higher single-user interactivity within two weeks since day-0, and 6 to 11% better peak throughput at high concurrency, benchmarked on a real multi-turn agentic coding workload. Our new TopK-V2 kernel is 2.33x faster at 80K ISL, scaling to 10.17x at 1M ISL, keeping interactivity essentially flat out to 1M tokens. Part of the story is the architecture itself. GLM-5.2 applies IndexShare to its DSA layers and ships a stronger MTP head reusing IndexShare and KVShare. The rest comes from our serving optimizations. Special thanks to NVIDIA AI for the help in day-0 support of GLM-5.2 NVFP4, and to Z.ai for IndexShare in SGLang!
LMSYS Org107,125 просмотров • 1 месяц назад

🎉 Congrats to Fish Audio on launching Fish Audio S2, a frontier TTS model with fine-grained prosody & emotion control via natural-language inline tags. SGLang Day-0 support is now live! 🏆 Best WER on Seed-TTS Eval; 81.88% win rate on EmergentTTS-Eval 🎙️ Voice cloning with 86.4% prefix-cache hit rate via RadixAttention ⚡️ RTF 0.34, 63.3 tok/s on single H200 (single batch) 🌍 Trained on 10M+ hours of audio across ~100 languages, GRPO-aligned 🔧 Dual-AR (Slow + Fast AR) is LLM-isomorphic: continuous batching, paged KV cache & CUDA graphs inherited natively 🗣️ Native multi-speaker: turn-taking, interruptions & cross-speaker emotion in a single pass 👉Cookbook: 👉Blog: 🎬 Curious how to run with SGLang? Check out this voice cloning demo from Chayenne Zhao with Fishaudio-S2-Pro:
LMSYS Org39,939 просмотров • 5 месяцев назад

🎉 Meet Krea 2 from Krea, an aesthetic open-source image model ranked #1 text-to-image from an independent lab on Artificial Analysis. Day-0 support is now live in SGLang! Krea 2 ships as two models built to work together: 1️⃣ RAW: undistilled base checkpoint: diverse & malleable, made for fine-tuning, post-training & LoRA 2️⃣ Turbo: 8-step distilled checkpoint: fast, high-quality text-to-image Train LoRAs on RAW, run them on Turbo: base for training, Turbo for fast inference on your own hardware Run it now with SGLang!
LMSYS Org11,216 просмотров • 2 месяцев назад
Больше нет контента для загрузки