Loading video...
Video Failed to Load
LLM 的速度,有两种体验:等多久开始回答,回答时是否流畅。 两种体验,与 Prefill 和 Decode 密切相关。 从发出请求到收到第一个 token,这段等待叫 TTFT,首 token 反应时间。 模型在其中进行 prefill,处理输入内容 开始回答后,模型通过 decode 逐步生成后续 token,相邻两个 token 到达的间隔叫 ITL,token间延迟。 TTFT 越短,反应越快;ITL 越短且越稳定,回答就越流畅。
11,768 views • 18 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
