Loading video...
Video Failed to Load
Muse Glimmer, A 30B parameter dense model swallowing a 130,000 token context window using only 19.3 GB of VRAM (extreme efficiency). No KV cache quantization required. I just benched the new Muse Glimmer 30B (dense) on a single RTX 4090. We are pulling 3,100+ t/s prefill and 75 tokens/second... show more
64,857 views • 10 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
