Loading video...
Video Failed to Load
We've partnered to bring more Gemma 3 quantized models to you! ๐ We worked with Georgi Gerganov llama.cpp, LM Studio, MLX, ollama to make sure you can run it using your favorite tool! Gemma models optimized with QAT, reduce memory requirements while keeping quality! All models checkpoints are available... show more
15,996 views โข 1 year ago โขvia X (Twitter)
11 Comments

Read more about it, explore quantized Gemma and start building with your favorite tool: Blog: Llama cpp: LMStudio: Ollama: MLX: Google AI Edge (For Phones):

Upgrade your Tesla with UP-03 Forged Wheels from Unplugged Performance! Unmatched strength, lightweight design, and track-proven durabilityโperfect for Model S, 3, X, and Y. Ready to ship with a lifetime warranty. #Tesla #UP03 #UnpluggedPerformance

@ggerganov @lmstudio @ollama Community enthusiasts have further reduced the GGUF size of the Gemma-3-27b-it-qat model to just 15.6GB. 24GB VRAM X 16GB VRAM โ

@ggerganov @lmstudio @ollama Any comparison between the regular Q4_K_M ones and these QAT Q4_0?

May be you should also partner with @NeuralMagic to get handy quants out of their llm-compressor for vLLM junkies like me. ;-)

@ggerganov @lmstudio @ollama @awnihannun Maybe it's time for MLX to have it's own X account, so it can be tagged? ๐

@ggerganov @lmstudio @ollama Do they truly keep quality? Because when I tested them, they didnโt work the same at Hebrew language.

@ggerganov @lmstudio @ollama 1) How does your 27B QAT int4 version compare to traditional Q6_K in precision/accuracy relative to the regular model? 2) Will Vision capabilities work on ollama? There have been issues with some Gemma versions there Great work! Now, we need Gemma 3 Thinking ๐

@ggerganov @lmstudio @ollama Fantastic!

@ggerganov @lmstudio @ollama ยกInnovaciรณn tangible! Gemma 3 con QAT acerca los LLMs a mรกs desarrolladores. ยฟCรณmo impactarรก en chatbots o anรกlisis en-edge? ๐ La optimizaciรณn para RTX 4060/3090 es un logro notable.

@ggerganov @lmstudio @ollama Big models. Small GPUs. Still fast, still good. This isnโt just about tooling anymore llama.cpp, MLX, and Ollama are laying the groundwork for something bigger. Training or inference, local is back.
