Video wird geladen...
Video konnte nicht geladen werden
What compression looks like on vLLM. Same Gemma 4 31B. Red Hat AI's quantized version runs at nearly 2x tokens/sec, half the memory, 99%+ accuracy retained. Open source. Quantized with LLM Compressor. Links in comments. 🙏 Sawyer Bowerman for the 2-minute demo.
34,260 Aufrufe • vor 5 Monaten •via X (Twitter)
9 Kommentare

Quantize your own model with LLM Compressor: Already quantized Gemma 4 models: -RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic: -RedHatAI/gemma-4-31B-it-FP8-Dynamic: -RedHatAI/gemma-4-31B-it-FP8-block: -RedHatAI/gemma-4-31B-it-NVFP4: -RedHatAI/gemma-4-26B-A4B-it-NVFP4: We also created a speculator model, RedHatAI/gemma-4-31B-it-speculator.eagle3: Happy efficient inferencing!

@vllm_project @_soyr_ cracked hacker quantizes your favorite model and uploads it for everyone to use

@vllm_project @_soyr_ This is awesome! Is @RedHat_AI watching for and pulling in chat template fixes from @GoogleDeepMind and @UnslothAI for these? I think there were some upstream changes for better tool calling, agentic use, etc. made in the past few days.

@vllm_project @_soyr_ Great work!

@vllm_project @_soyr_ Quantized a model last month for my finance app and the speed difference was wild. Worth the accuracy tradeoff for most use cases honestly.

, @DAlistarh, and team did 500,000 evals on quantized models and found that accuracy impact is very minimal. They wrote a paper on the research called "Give Me BF16 or Give Me Death? Accuracy-Performance Trade-Offs in LLM Quantization": All quantized models in the Red Hat AI Hugging Face repo, for example, recover to 99%+ of baseline accuracy:

@vllm_project @_soyr_ Local Models are becoming more and more appealing with each day that goes by. This is awesome

@vllm_project @_soyr_ Any chance of a minimax-2.7 sub 96gb? 🙏

@vllm_project @_soyr_
