Video wird geladen...
Video konnte nicht geladen werden
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out... show more
148,203 Aufrufe • vor 1 Monat •via X (Twitter)
3 Kommentare

Johnvor 1 Monat
speculative decoding was the one that surprised me most in practice. the draft model rejection rate changes a lot based on prompt style, so batching similar prompt types together before running speculative decode gave a much bigger throughput lift than just tuning the draft model itself

Aayaan Naqvivor 1 Monat
@baseten @philipkiely @waterloo_intern woahhhhh @waterloo_intern sick

Jong Hyun Parkvor 1 Monat
@baseten @philipkiely @waterloo_intern Is a durable moat possible in inference tech on a multi year timescale?

