Loading video...

Video Failed to Load

Go Home

i just beat Google DeepMind's turboquant introducing Shard. 10x KV cache compression on Llama-3.1-8B. zero quality loss - 10x @ 8K context, 11.2x @ 32K - NIAH recall 1.000 across 4K-32K - LongBench Δ ≈ 0 vs FP16 turboquant tops out at 4-6x at the same quality. we doubled...

153,712 views • 18 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos