Video wird geladen...
Video konnte nicht geladen werden
Sentra just killed Google Research's TurboQuant. SpectralQuant — 5.95× KV cache compression on Mistral 7B at +7.5% perplexity overhead. TurboQuant at the same compression: +22%. 3× less degradation. 15-second calibration. One per-model, then drop-in for any HuggingFace LLM, ViT, ESM, AlphaFold Evoformer, or VideoMAE. Check out the findings and... show more
62,624 Aufrufe • vor 4 Monaten •via X (Twitter)
23 Kommentare

1. The problem — Modern transformers spend most of their inference memory on the KV cache. Mistral 7B at 8k context = 4 GB just for keys and values. Llama 70B at 32k = 80 GB. Quantizing this is the highest-leverage compression target in inference today. TurboQuant (ICLR 2026) was state-of-the-art at it.

2. The insight Google missed — KV vectors are not random. We measured the eigenspectrum of every attention head in Mistral, Qwen, ESM-2 and ViT. The pattern is universal: ~80% of the variance lives in 4 of 128 dimensions. This is the empirical fact every compressor should exploit. TurboQuant doesn't.

3. Why random rotation wastes bits — TurboQuant rotates KV vectors by a random orthogonal matrix Π before quantizing. That's mathematically optimal *if you don't know anything about your data*. But we do know — and after Π, the variance is roughly evenly spread across all 128 coords. Same total bits, sprayed across noise.

4. What we do instead — SpectralQuant rotates by the eigenvectors of the K covariance, measured during a 15-second calibration pass. After rotation: dim 0 has variance λ₀ (largest). Dim 127 has variance λ₁₂₇ (~0). We know exactly where the information lives, so we put the bits there. This is just the Karhunen-Loève Transform from 1947, applied to attention.

5. Then the second trick — selective correction TurboQuant applies its JL bias correction to all 128 dims. Costs 128 sign bits per token. Correction noise scales with √128. We apply it to only the 4 signal dims. 4 sign bits. Noise scales with √4. 30× fewer dot products at attention time. Same accuracy. Less softmax noise.

6. And water-filling for the dominant dims — Even within the top 4 signal dims, eigenvalues vary. λ₀ = 0.50, λ₃ = 0.04. Uniform allocation: [3, 3, 3, 3] bits per dim. Water-filling (Shannon-Berger, 1948): [5, 3, 2, 2]. The dominant dim gets 64 levels of precision. Same total budget, optimal split.

7. This isn't an LLM trick — The eigenvalue cliff shows up wherever attention does. We tested on: · Mistral 7B → 5.95× KV compression, +7.5% perplexity · DINOv2 / Depth Anything → compression improves accuracy · ESM-2 → matched fold quality on CASP15 · VideoMAE / AlphaFold Evoformer → same story 15-second calibration. Pure PyTorch. pip install spectralquant next week.

We've open-sourced the whole thing, code, data, and results. Take a look:

@sentra_app @GoogleResearch

@sentra_app @GoogleResearch OKAY NOW WE ARE FUCKING TALKING

@sentra_app @GoogleResearch holy

@sentra_app @GoogleResearch Cool stuff guys, we should collab

@sentra_app @GoogleResearch DM me. We can figure something out.

@sentra_app @GoogleResearch Lmao,This logo looks like my previous team's Anyway, congratulations on the great release!

@sentra_app @GoogleResearch Does it add latency to prefill?

@sentra_app @GoogleResearch soo interesting! i lovee the way you explain @ashwingop

@sentra_app @GoogleResearch really cool!

@sentra_app @GoogleResearch Such a brilliant explanation and a great innovation!

@sentra_app @GoogleResearch Let’s talk!

@sentra_app @GoogleResearch DMed

@sentra_app @GoogleResearch This smells like poop bs

@sentra_app @GoogleResearch @MrAhmadAwais

@sentra_app @GoogleResearch Great stuff!!
