Video yükleniyor...
Video Yüklenemedi
I built an interactive simulator that shows why LLM inference capacity is hard. Batching, queues, KV cache, GPU decode, priority tiers - all coupled. Change one knob and the bottleneck moves.
24,726 görüntüleme • 2 ay önce •via X (Twitter)
0 Yorum
Yorum bulunmuyor
Orijinal gönderinin yorumları burada görünecek
