正在加载视频...

视频加载失败

With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone 🚀 At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public. What K3 brings: 🧠 2.8T-parameter...

103,950 次观看 • 1 个月前 •via X (Twitter)

19 条评论

vLLM 的头像
vLLM1 个月前

Fitting 2.8T parameters was not the hard part. Caching them was. Most of K3's layers hold fixed-size KDA state instead of a growing KV cache, so there is no per-token KV to hash. Prefix caching was rebuilt around that, and every hybrid model after K3 inherits the result. 2/6

vLLM 的头像
vLLM1 个月前

✨ What you can turn on: 🔷P/D disaggregation, TP8 prefill to DP16/EP16 decode over NIXL 🔷MoE backend per topology, mega_moe for EP, trtllm for TP>1 🔷KV offloading, so the agent turns skip re-prefill 🔷Tool calling, reasoning, structured output 🔗 3/6

vLLM 的头像
vLLM1 个月前

📈For ultra-low latency on a 2.8T model without accuracy loss, speculative decoding is the natural choice. @inferact trained and open-sourced a DSpark speculator for K3 that drafts multiple tokens in a single parallel pass. Perf improvement: 118 → 370 tok/s single-stream, ~3.14× on real reasoning workload datasets. 🔗 4/6

vLLM 的头像
vLLM1 个月前

Kernels, recipes and the full write-up: 🔗 5/6

vLLM 的头像
vLLM1 个月前

And celebrating a huge milestone, we have merged Kimi K3 as PR #50000! 🎊🎊

Tuhin Srivastava 的头像
Tuhin Srivastava1 个月前

Was great to collaborate!

Robin Cheng 的头像
Robin Cheng1 个月前

The best part of this is that whenever vLLM posts numbers, they are real, not inflated, and not cherry-picked. That's actually quite rare in the industry. If vLLM gives us a number, we can trust it's reproducible and real.

zeeksa 的头像
zeeksa1 个月前

Gotta spin up my b300 for this

Omar Da'ajneh 的头像
Omar Da'ajneh1 个月前

why would vllm vibe code a video?

OrcaRouter 🐳 的头像
OrcaRouter 🐳1 个月前

Some free K3 credits to test:

Alois Vaclav 的头像
Alois Vaclav1 个月前

Hehe, this must be your happiest model release, cause nobody can actually run it and put in bug reports 😅

Sudip 的头像
Sudip1 个月前

Here is the Kimi K3 architecture breakdown such that anyone can digest in the form of an explainer video, one-shotted by Simi by Lamina Labs.

Kanishk Patel 的头像
Kanishk Patel1 个月前

Day-0 support is quietly the best argument for composing engines unforked. NVIDIA's Molt carries zero vLLM patches, so an engine upgrade is a container pin and K3-class models reach RL training the day the weights land.

Anmol 的头像
Anmol1 个月前

curious how many nodes it takes to even boot 2.8t

Nick Venturi 的头像
Nick Venturi1 个月前

my local power grid is going to hate me the second those weights drop

Agades Technologies 的头像
Agades Technologies1 个月前

🔥🔥

Sebastian Buzdugan 的头像
Sebastian Buzdugan1 个月前

1m context is nice, but prefill cost will decide most startup deployments

Nicholas Hyperion 的头像
Nicholas Hyperion1 个月前

Moonshot dropping 2.8T open weights with day-0 vLLM support. Closed labs must be thrilled

AI Mastery Guide 的头像
AI Mastery Guide1 个月前

A 1M token context that's actually affordable is the real win here.

相关视频

The Kimi AI Model And What If This Carnegie Mellon, PhD student was incentivized to open source AI in the US? The answer is Kimi would have been an open source US model. Yang Zhilin Moonshot’s founder and maker of the Kimi series turned his deep research expertise into one of the world’s most capable open AI systems. After earning his PhD at Carnegie Mellon under leading researchers and interning at Google Brain and Meta, he returned to China and co-founded Moonshot AI in March 2023 with Tsinghua classmates Zhou Xinyu and Wu Yuxin. Why? The US VC world and tax structure did not favor Zhilin’s proposal to open source the AI models as a strategy. So he left. He named the company after his favorite Pink Floyd album, reflecting his ambitious vision for scalable. Yang assembled a core technical team of inventors behind breakthroughs like Transformer-XL and RoPE. Together they focused on turning massive compute into efficient intelligence through innovative architectures. The journey began with the Kimi chatbot in October 2023, rapidly scaling context from 200,000 to millions of characters. This evolved into the Kimi series: K1.5 matched top reasoning models, K2 introduced a 1-trillion-parameter Mixture-of-Experts design trained on 15.5 trillion tokens and released openly, and K2 Thinking added advanced agentic capabilities. Kimi K3 represents the pinnacle, this 2.8-trillion-parameter model uses a sparse MoE architecture with 896 experts (only 16 active per token), new Kimi Delta Attention and attention residuals for efficiency, and a 1-million-token context window. It delivers frontier performance in long-horizon coding, reasoning, and multimodal tasks at competitive cost, with weights set for open release. Yang’s approach emphasizes openness, efficiency, and continuous self-improvement — enabling solo developers and teams to achieve what once required massive resources. By sharing technical insights in public talks, he has accelerated global progress toward more accessible, powerful AI. Imagine if we held open source higher than the fear theater games of Anthropic? We are chasing out some of the best minds. This is how you lose…

Brian Roemmele

37,185 次观看 • 1 个月前

Chinese AI models are wiping billions off Big Tech right now. Google just lost $200 billion in a single day, and the model it needed to fight back still isn't ready. Gemini 3.5 Pro, Google's most powerful model, is months behind schedule. Alphabet stock dropped 4.4% that same day. The Deepseek moment is happening again, and the new model is FAR bigger. On the same day Google's delay leaked, a Beijing lab called Moonshot released Kimi K3. It is the largest open model ever built, with 2.8 trillion parameters. It took the number one spot on the Frontend Code Arena, a live coding leaderboard, passing Anthropic's best model. And Moonshot is giving it away for free on July 27. The genius part: Anyone with enough computers can download it and run a frontier level AI without paying a cent to a US company. A single task on Kimi K3 costs about 94 cents. The same work on some American models costs nearly double. So why would a company keep paying premium prices for a model it can now get for free? The entire US AI business is built on selling access to models that cost billions to train. If a free Chinese version does most of the same work, that pricing power starts to crack. And Kimi is close to the best. On one closely watched intelligence ranking it scored 57, just behind the top American models GPT-5.6 Sol and Fable 5, and ahead of Claude Opus 4.8. Bank of America told clients that Kimi proves Chinese labs can keep making big leaps even with limited chips. And the founder of Moonshot, Yang Zhilin, learned to build AI as a researcher INSIDE Google. Google literally wrote the 2017 paper that made all of these models possible. Now the people who studied its work are using it to destroy Google, and handing it out for free. What happens next: Kimi K3's weights go public on July 27. Google reports earnings on July 22, and everyone will be asking the same question about Gemini. If free models keep topping the charts, every valuation built on paid AI access has to be rewritten. What do you think?

Ricardo

47,790 次观看 • 1 个月前