正在加载视频...
视频加载失败
A good technical LLM interview question: Your LLM chatbot takes 12s before it generates the first token, and the users are complaining. So you move the model onto a GPU with 3x the computing power. The time to first token barely improves. Why did this happen? (answer below) Latency... show more
21,007 次观看 • 5 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
