正在加载视频...
视频加载失败
Ex-vLLM core contributor explained how to make LLM inference 10x cheaper in 34 minutes - better than $3000 inference optimization bootcamps. request comes in -> check LMCache -> hit? load KV cache from CPU/SSD/remote -> skip prefill -> serve. That loop is why Bloomberg and other production stacks now... show more
0 条评论
暂无评论
原始帖子的评论将显示在这里
