Загрузка видео...
Не удалось загрузить видео
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out... show more
148,203 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 3

John1 месяц назад
speculative decoding was the one that surprised me most in practice. the draft model rejection rate changes a lot based on prompt style, so batching similar prompt types together before running speculative decode gave a much bigger throughput lift than just tuning the draft model itself

Aayaan Naqvi1 месяц назад
@baseten @philipkiely @waterloo_intern woahhhhh @waterloo_intern sick

Jong Hyun Park1 месяц назад
@baseten @philipkiely @waterloo_intern Is a durable moat possible in inference tech on a multi year timescale?

