
Guangxuan Xiao
@Guangxuan_Xiao • 5,255 subscribers
MTS @thinkymachines | Ph.D. @MITEECS
Videos

Introducing DuoAttention: Our new framework slashes both memory and latency for long-context LLMs without sacrificing performance! By applying full KV cache only to critical heads, we achieve: ⚡ 2.55x memory reduction ⚡ 2.18x decoding speedup ⚡ 3.3M tokens on a single A100 GPU
Guangxuan Xiao31,073 просмотров • 1 год назад
Больше нет контента для загрузки