Loading video...
Video Failed to Load
A good technical LLM interview question: Your LLM chatbot takes 12s before it generates the first token, and the users are complaining. So you move the model onto a GPU with 3x the computing power. The time to first token barely improves. Why did this happen? (answer below) Latency... show more
21,163 views • 5 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
