正在加载视频...

视频加载失败

Super excited to introduce Gemma 4 12B! 💎 - Multimodal: audio, image, video, and text input - Novel architecture: we removed the multimodal encoders for a unified, streamlined arch - New MacOS desktop app powered by LiteRT - MTP support Excited to see what you build with it!

125,795 次观看 • 3 个月前 •via X (Twitter)

39 条评论

Omar Sanseviero 的头像
Omar Sanseviero3 个月前

We collaborated with Hugging Face, llama.cpp, Ollama, VLLM, SGLang, Unsloth, MLX, LM Studio, and the rest of the ecosystem to land day 0 support. Enjoy! Read our developer guide:

Sakura Yuki 的头像
Sakura Yuki3 个月前

The real win with encoder-free isn't just saving VRAM, it's the TTFT impact. Bypassing the heavy encoder forward pass means prefill starts instantly, criminally underrated for local edge runs.

Ajith 的头像
Ajith3 个月前

I just made a video about the new Gemma model what it is, what's changed, and everything user/dev should know from Google's blog. Credits: Created using DistilBook(.)com

Alina Fomina 的头像
Alina Fomina3 个月前

the benchmarks are wild – 12b almost matches the 26b

Arunabh hazarika 的头像
Arunabh hazarika3 个月前

Gemma 4:124b when?

Revo Laition 的头像
Revo Laition3 个月前

Macos and ios app not available in europe yet?

Omar Sanseviero 的头像
Omar Sanseviero3 个月前

Looking into it

Omar Sanseviero 的头像
Omar Sanseviero3 个月前

Should be working now

Revo Laition 的头像
Revo Laition3 个月前

Still getting this for ios, havent retried mac app yet.

Suril Shah 的头像
Suril Shah3 个月前

@osanseviero Hi, the macOS app should be available in the EU now! Please let me know if it still doesn't work for you. iOS is still restricted unfortunately, will keep you posted!

Revo Laition 的头像
Revo Laition3 个月前

@osanseviero Thanks for the followup! Yes, macos app now works😀

Revo Laition 的头像
Revo Laition3 个月前

@osanseviero exploring the eloquent macos app now. great job! one thing though. which model is this? havent seen any official Gemma 4 model at 2B? is it the as the one in the edge gallery app?

AshutoshShrivastava 的头像
AshutoshShrivastava3 个月前

Incredible Omar 👏👏

Haider. 的头像
Haider.3 个月前

@samsheffer sooo good 🔥🔥

bharani manoharan 的头像
bharani manoharan3 个月前

thank you. this sounds very promising. is there a mlx? can't find on hf

Omkar Deshmukh 的头像
Omkar Deshmukh3 个月前

It is very slow, at least transcribe fast and then refine it, but both are very slow on M1 mac, cold start is brutal as well takes 10 sec to start transcribing.

weople 的头像
weople3 个月前

the encoder removal is the headline for me. model releases happen every week but changing a core part of how multimodal systems are built is a much bigger deal.

Alex Rogov 的头像
Alex Rogov3 个月前

the "separately trained components bolted together" problem is real. ran cross-modal tasks where the vision encoder and the LLM clearly disagreed on what to attend to. unified architecture doesn't just simplify deployment. it changes what the model can reason about natively. different ceiling, not just different cost.

Poonam Soni 的头像
Poonam Soni3 个月前

the era of running powerful AI locally just got real.

Harris 的头像
Harris3 个月前

Google studio?

adhd.dev 的头像
adhd.dev3 个月前

excited to see what all going to build with powerful local audio + vision capabilities using Gemma 4 12B! 💻

Vaclav Cerny 的头像
Vaclav Cerny3 个月前

@ivanfioravanti cool, now go 10x

‎‎ . 的头像
‎‎ .3 个月前

Gemma 4 12B , is real gem for local llm, while being multimodal its Token per second speed is best,

Sanchit 的头像
Sanchit3 个月前

sadly it needs more than like 16 gb of vram to be able to run locally. i am running distilled qwen locally and some other open models in kilo

SolomanNB 的头像
SolomanNB3 个月前

Local LLMs are finally thriving! Thanks for saving my MacBook💻

Taylor Arndt 的头像
Taylor Arndt3 个月前

Thank you for making this. I think we’re gonna be putting this in our app. @PerspectIntel I may also see how it runs on my MacBook Neo. I do like your new approach smaller and cheaper. Open models are the future of AI.

Vanar 的头像
Vanar3 个月前

🔥🔥

Rohit Jha 的头像
Rohit Jha3 个月前

Multi modal out?

Techonsapevole priv/acc d/acc 的头像
Techonsapevole priv/acc d/acc3 个月前

Nice, but I still prefere Qwen3.6 35b a3b

Benjamin Atkin 的头像
Benjamin Atkin3 个月前

@grok Too small to truly be multimodal?

synabun.ai 的头像
synabun.ai3 个月前

unified arch without separate encoders means 12B multimodal fits where it wouldn't before. MacOS LiteRT app on day 0 is a good sign they're building for people who want to run things, not just screenshot benchmarks.

Matt Henderson 的头像
Matt Henderson3 个月前

The instruction tuned model suffers from bad hallucinations when understanding audio inputs. The base model doesn’t seem to have that problem. So I suspect something about instruction tuning regressed audio

Taha ⵣ 的头像
Taha ⵣ3 个月前

GOATs

Miguel Guerrero 的头像
Miguel Guerrero3 个月前

Congrats on the launch! Awesome model

ρ:ɡeon 的头像
ρ:ɡeon3 个月前

i think we just need gemma 4.1 26b

vright 的头像
vright3 个月前

This is awesome. Incredible work! Such a huge fan of the open multimodal work

The Canaanite 的头像
The Canaanite3 个月前

This is awesome stuff

Jananadi W 的头像
Jananadi W3 个月前

Super excited to start building with this!

Craig Merry 的头像
Craig Merry3 个月前

🔥

相关视频