Loading video...

Video Failed to Load

Go Home

Now you can use Gemma directly in the Gemini CLI! 🚀 v0.40.0 introduces experimental support for local Gemma models, starting with intelligent model routing (with full local execution on the roadmap!).

89,093 views • 5 months ago •via X (Twitter)

32 Comments

Google Gemma's profile picture
Google Gemma5 months ago

Full release notes in the official article here:

Brad H's profile picture
Brad H5 months ago

I asked Gemini CLI to just set this up for me in a way that it just works and I don't have to think about it, put it in YOLO mode, and 30 seconds later, I'm all set up. This is absolutely amazing!

Mike's profile picture
Mike5 months ago

Why not just have Gemma team make their own GemmaCLI that works with the GeminiCLI team but specializes in Gemma optimizations. It's really silly needing 3rd party tools to use these models. I should be able to use Gemma on my computer and phone using the same 1st party app.

Microtexter's profile picture
Microtexter5 months ago

Today I asked Gemini Cli “what is mcp?” It went to and didn’t reply in 10 minutes. Now I wander what Gemini Cli is for?

Roronoa D. Zoro's profile picture
Roronoa D. Zoro5 months ago

Why is Gemma not available inside Antigravity

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack5 months ago

@googlegemma can't wait to try this out, local execution sounds like a game changer!

tamimbuilds's profile picture
tamimbuilds5 months ago

I made a tweet yesterday and today that is live thanks

Cheicolate's profile picture
Cheicolate5 months ago

Hugeeee

Soroush Fadaeimanesh's profile picture
Soroush Fadaeimanesh5 months ago

intelligent routing is the right unlock. local for repetitive bounded tasks, cloud for the long tail. cuts latency on 80% of the dev loop without giving up the 20% that needs the big model.

Josh Herzberg's profile picture
Josh Herzberg5 months ago

Cool. Can run all day long on ur local machine. I just want my new Mac to come. This is another way that Google is partnering with Apple for AI 😂

Hermes Agent Tips's profile picture
Hermes Agent Tips5 months ago

This is a GAMECHANGER

JC Castaneda's profile picture
JC Castaneda5 months ago

can we fine-tune Gemma with Rust programming language related data?

Soroush Fadaeimanesh's profile picture
Soroush Fadaeimanesh5 months ago

The model routing piece is the unlock. Local for cheap fast turns, cloud for heavier reasoning. Curious if the routing decision is token-budget driven or complexity heuristic. The harness layer just got more interesting.

Pierpaolo Wurzburger's profile picture
Pierpaolo Wurzburger5 months ago

And Gemma via api is possible?

ZQ's profile picture
ZQ5 months ago

That‘s great. It helps to reduce the anxiety of tokens.

Avais Aziz's profile picture
Avais Aziz5 months ago

Interesting addition in 0.40.0. Local Gemma for routing decisions via gemini gemma commands should help with latency and cost on simple tasks. Setup flow with LiteRT-LM looks straightforward.

Micha's profile picture
Micha5 months ago

Is the local execution path actually running on the host CPU/GPU, or is there a required backend service?

Roshan Ramani's profile picture
Roshan Ramani5 months ago

finally, my cli can hallucinate faster than me

Kimi's profile picture
Kimi5 months ago

"This is huge for edge AI. What if you also integrated Gemma with local data storage, allowing users to train models on device with their own data? That would take personalization to the next level"

Gregor's profile picture
Gregor5 months ago

What's the local execution timeline looking like, and will it actually run offline or still phone home for routing?

Chat Data's profile picture
Chat Data5 months ago

Local routing first is a smart wedge. Once full local execution lands, the interesting part will be how clearly the CLI shows when a task stayed on device versus when it escalated to a hosted model.

Jethro Lising's profile picture
Jethro Lising5 months ago

um based?

Keepass Keep's profile picture
Keepass Keep5 months ago

Just stop to tweeting and add API expose to "Android edge gallery" app. Nothing more.

Bessi's profile picture
Bessi5 months ago

Local for the easy stuff, cloud for when it matters.

Soroush Fadaeimanesh's profile picture
Soroush Fadaeimanesh5 months ago

Local routing changes the latency math for tool-use loops. Curious if the router decides per-call or per-conversation. Per-call gets you cheap reflection without losing the bigger model on planning.

Trucos IA | Ahorra Tiempo's profile picture
Trucos IA | Ahorra Tiempo5 months ago

Esto es genial! buena noticia sin duda

Optinuss's profile picture
Optinuss5 months ago

😊 great thanks! Which Gemma 4 models are optimized for the CLI?

Alexey Volkov 🇸🇮's profile picture
Alexey Volkov 🇸🇮5 months ago

Any chance to see local Gemma in Antigravity?

adelvo software's profile picture
adelvo software5 months ago

Love this move by Google with intelligent local ↔ cloud routing for Gemma in the Gemini CLI! 📷This is exactly what our users praise most about Local Live Translator — the ability to effortlessly switch between fully local AI (on-device Whisper + Mistral 7B on Apple Silicon) and cloud models for real-time live translation. Blazing fast, private, and flexible. Native Mac app with a free iOS companion too! Keep pushing the local + hybrid AI frontier! 📷 🚀

Petra's profile picture
Petra5 months ago

Great, I can make a Snake Game.

AJ's profile picture
AJ5 months ago

What are the current bottlenecks in the models we have right now? In terms of power, search and speed? Can we make them generate results fast?

Davit Yan's profile picture
Davit Yan5 months ago

finally, but how to download and use Gemma 4 models???

Related Videos