Video yükleniyor...
Video Yüklenemedi
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out... show more
148,203 görüntüleme • 1 ay önce •via X (Twitter)
3 Yorum

John1 ay önce
speculative decoding was the one that surprised me most in practice. the draft model rejection rate changes a lot based on prompt style, so batching similar prompt types together before running speculative decode gave a much bigger throughput lift than just tuning the draft model itself

Aayaan Naqvi1 ay önce
@baseten @philipkiely @waterloo_intern woahhhhh @waterloo_intern sick

Jong Hyun Park1 ay önce
@baseten @philipkiely @waterloo_intern Is a durable moat possible in inference tech on a multi year timescale?

