正在加载视频...
视频加载失败
An H100 GPU can run 10^15 FLOPs/s. But somehow, a model like Gemini Flash-Lite is stuck at 200 tokens/s. These numbers show just how much of a bottleneck memory bandwidth is. In this video, we look at how LLMs are evolving to work around this limitation. Thanks Inception for... show more
26,110 次观看 • 6 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
