
Krish
@krishgarg • 1,828 subscribers
cs @uwaterloo
Videos

i just beat Google DeepMind's turboquant introducing Shard. 10x KV cache compression on Llama-3.1-8B. zero quality loss - 10x @ 8K context, 11.2x @ 32K - NIAH recall 1.000 across 4K-32K - LongBench Δ ≈ 0 vs FP16 turboquant tops out at 4-6x at the same quality. we doubled it. read more: Kirri
Krish152,452 views • 9 days ago
1:11
Sensitive content
This media may contain sensitive content.
No more content to load