Loading video...
Video Failed to Load
i just beat Google DeepMind's turboquant introducing Shard. 10x KV cache compression on Llama-3.1-8B. zero quality loss - 10x @ 8K context, 11.2x @ 32K - NIAH recall 1.000 across 4K-32K - LongBench Δ ≈ 0 vs FP16 turboquant tops out at 4-6x at the same quality. we doubled... show more
153,712 views • 18 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
