Loading video...
Video Failed to Load
we sped up distributed inference by up to 5x with decentralized speculative decoding. many don't realize that AI models normally generate text one single word at a time, waiting for the network after every word. speculative decoding changes this by using a "guess & confirm" system, similar to autocomplete.... show more
45,584 views • 6 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

