Loading video...
Video Failed to Load
The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out... show more
137,910 views • 3 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

