Loading video...
Video Failed to Load
Ex-vLLM core contributor explained how to make LLM inference 10x cheaper in 34 minutes - better than $3000 inference optimization bootcamps. request comes in -> check LMCache -> hit? load KV cache from CPU/SSD/remote -> skip prefill -> serve. That loop is why Bloomberg and other production stacks now... show more
61,202 views • 1 month ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
