Video wird geladen...
Video konnte nicht geladen werden
4/ to achieve maximum memory efficiency, we quantize model weights to ~4-bit, getting the language model under 20GB with room for the kv cache, perception encoder, and drafter alongside it. a dflash drafter proposes blocks of tokens the main model verifies in parallel, so it stays responsive.
67,151 Aufrufe • vor 13 Tagen •via X (Twitter)
0 Kommentare
Keine Kommentare verfügbar
Kommentare vom Original-Post werden hier angezeigt
