
SGLang
@sgl_project • 8,049 subscribers
Run LLMs fast at any scale 🔗 https://t.co/F3u6wYESL0 Join our community https://t.co/fmlOfTOEec For AI tech blogs & deep-dives 👉 @lmsysorg
Shorts
Videos

GLM-5.3 weights from Z.ai are live, with SGLang powering day-0 serving support! GLM-5.3 inherits every optimization and feature we battle-tested for GLM-5.2 over the past months. On real-world multi-turn agentic workloads, we measured 537.6 tok/s/user on NVFP4 and 413 tok/s/user on FP8, at BS=1 with TP8 on 8x B300. It's fast, efficient, and production-ready today on NVIDIA AI Blackwell and Hopper, and AI at AMD MI300X/325X/355X. We believe this is a big step forward for GLM-Series in agentic tasks, with a path toward mythos-class cyber capability. Can't wait to see what people build with GLM-5.3 and SGLang🚀 Cookbook👇
SGLang96,129 görüntüleme • 4 gün önce

📢 Qwen's Qwen3.8-2.4T-A95B weights just dropped, and SGLang already has day-0 support with DSpark. HF: Cookbook: 🖖 What kind of model can help optimize an inference engine? Qwen3.8-2.4T-A95B can take on complex engineering tasks, plan and iterate on its own, and keep going until the job is done. We even asked it to optimize its own serving on SGLang, and it ran unattended for 4.5 hours with verified performance improvements. Thanks to the Qwen team Qwen, NVIDIA AI, and AI at AMD for the close collaboration.
SGLang27,145 görüntüleme • 20 gün önce
Daha fazla içerik yok.