正在加载视频...
视频加载失败
4/ to achieve maximum memory efficiency, we quantize model weights to ~4-bit, getting the language model under 20GB with room for the kv cache, perception encoder, and drafter alongside it. a dflash drafter proposes blocks of tokens the main model verifies in parallel, so it stays responsive.
67,151 次观看 • 13 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
