正在加载视频...
视频加载失败
Super excited to introduce Gemma 4 12B! 💎 - Multimodal: audio, image, video, and text input - Novel architecture: we removed the multimodal encoders for a unified, streamlined arch - New MacOS desktop app powered by LiteRT - MTP support Excited to see what you build with it!
39 条评论

We collaborated with Hugging Face, llama.cpp, Ollama, VLLM, SGLang, Unsloth, MLX, LM Studio, and the rest of the ecosystem to land day 0 support. Enjoy! Read our developer guide:

The real win with encoder-free isn't just saving VRAM, it's the TTFT impact. Bypassing the heavy encoder forward pass means prefill starts instantly, criminally underrated for local edge runs.

I just made a video about the new Gemma model what it is, what's changed, and everything user/dev should know from Google's blog. Credits: Created using DistilBook(.)com

the benchmarks are wild – 12b almost matches the 26b

Gemma 4:124b when?

Macos and ios app not available in europe yet?

Looking into it

Should be working now

Still getting this for ios, havent retried mac app yet.

@osanseviero Hi, the macOS app should be available in the EU now! Please let me know if it still doesn't work for you. iOS is still restricted unfortunately, will keep you posted!

@osanseviero Thanks for the followup! Yes, macos app now works😀

@osanseviero exploring the eloquent macos app now. great job! one thing though. which model is this? havent seen any official Gemma 4 model at 2B? is it the as the one in the edge gallery app?

Incredible Omar 👏👏

@samsheffer sooo good 🔥🔥

thank you. this sounds very promising. is there a mlx? can't find on hf

It is very slow, at least transcribe fast and then refine it, but both are very slow on M1 mac, cold start is brutal as well takes 10 sec to start transcribing.

the encoder removal is the headline for me. model releases happen every week but changing a core part of how multimodal systems are built is a much bigger deal.

the "separately trained components bolted together" problem is real. ran cross-modal tasks where the vision encoder and the LLM clearly disagreed on what to attend to. unified architecture doesn't just simplify deployment. it changes what the model can reason about natively. different ceiling, not just different cost.

the era of running powerful AI locally just got real.

Google studio?

excited to see what all going to build with powerful local audio + vision capabilities using Gemma 4 12B! 💻

@ivanfioravanti cool, now go 10x

Gemma 4 12B , is real gem for local llm, while being multimodal its Token per second speed is best,

sadly it needs more than like 16 gb of vram to be able to run locally. i am running distilled qwen locally and some other open models in kilo

Local LLMs are finally thriving! Thanks for saving my MacBook💻

Thank you for making this. I think we’re gonna be putting this in our app. @PerspectIntel I may also see how it runs on my MacBook Neo. I do like your new approach smaller and cheaper. Open models are the future of AI.

🔥🔥

Multi modal out?

Nice, but I still prefere Qwen3.6 35b a3b

@grok Too small to truly be multimodal?

unified arch without separate encoders means 12B multimodal fits where it wouldn't before. MacOS LiteRT app on day 0 is a good sign they're building for people who want to run things, not just screenshot benchmarks.

The instruction tuned model suffers from bad hallucinations when understanding audio inputs. The base model doesn’t seem to have that problem. So I suspect something about instruction tuning regressed audio

GOATs

Congrats on the launch! Awesome model

i think we just need gemma 4.1 26b

This is awesome. Incredible work! Such a huge fan of the open multimodal work

This is awesome stuff

Super excited to start building with this!

🔥

