Loading video...
Video Failed to Load
The bottleneck in LLM inference isn't compute. It's how fast you can move the weights. Our CTO Mathias Lechner, Mathias Lechner, joins Piotr Mazurek, Piotr Mazurek (in SF 🌉), from our inference team, to discuss what actually limits token throughput and how we're optimizing for it.
21,393 views • 3 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
