正在加载视频...
视频加载失败
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out... show more
3 条评论

John1 个月前
speculative decoding was the one that surprised me most in practice. the draft model rejection rate changes a lot based on prompt style, so batching similar prompt types together before running speculative decode gave a much bigger throughput lift than just tuning the draft model itself

Aayaan Naqvi1 个月前
@baseten @philipkiely @waterloo_intern woahhhhh @waterloo_intern sick

Jong Hyun Park1 个月前
@baseten @philipkiely @waterloo_intern Is a durable moat possible in inference tech on a multi year timescale?

