Loading video...
Video Failed to Load
4/ to achieve maximum memory efficiency, we quantize model weights to ~4-bit, getting the language model under 20GB with room for the kv cache, perception encoder, and drafter alongside it. a dflash drafter proposes blocks of tokens the main model verifies in parallel, so it stays responsive.
67,151 views • 13 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
