Loading video...

Video Failed to Load

Go Home

The bottleneck in LLM inference isn't compute. It's how fast you can move the weights. Our CTO Mathias Lechner, Mathias Lechner, joins Piotr Mazurek, Piotr Mazurek (in SF 🌉), from our inference team, to discuss what actually limits token throughput and how we're optimizing for it.

21,393 views • 3 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos