Video wird geladen...
Video konnte nicht geladen werden
Now you can use Gemma directly in the Gemini CLI! 🚀 v0.40.0 introduces experimental support for local Gemma models, starting with intelligent model routing (with full local execution on the roadmap!).
89,093 Aufrufe • vor 5 Monaten •via X (Twitter)
32 Kommentare

Full release notes in the official article here:

I asked Gemini CLI to just set this up for me in a way that it just works and I don't have to think about it, put it in YOLO mode, and 30 seconds later, I'm all set up. This is absolutely amazing!

Why not just have Gemma team make their own GemmaCLI that works with the GeminiCLI team but specializes in Gemma optimizations. It's really silly needing 3rd party tools to use these models. I should be able to use Gemma on my computer and phone using the same 1st party app.

Today I asked Gemini Cli “what is mcp?” It went to and didn’t reply in 10 minutes. Now I wander what Gemini Cli is for?

Why is Gemma not available inside Antigravity

@googlegemma can't wait to try this out, local execution sounds like a game changer!

I made a tweet yesterday and today that is live thanks

Hugeeee

intelligent routing is the right unlock. local for repetitive bounded tasks, cloud for the long tail. cuts latency on 80% of the dev loop without giving up the 20% that needs the big model.

Cool. Can run all day long on ur local machine. I just want my new Mac to come. This is another way that Google is partnering with Apple for AI 😂

This is a GAMECHANGER

can we fine-tune Gemma with Rust programming language related data?

The model routing piece is the unlock. Local for cheap fast turns, cloud for heavier reasoning. Curious if the routing decision is token-budget driven or complexity heuristic. The harness layer just got more interesting.

And Gemma via api is possible?

That‘s great. It helps to reduce the anxiety of tokens.

Interesting addition in 0.40.0. Local Gemma for routing decisions via gemini gemma commands should help with latency and cost on simple tasks. Setup flow with LiteRT-LM looks straightforward.

Is the local execution path actually running on the host CPU/GPU, or is there a required backend service?

finally, my cli can hallucinate faster than me

"This is huge for edge AI. What if you also integrated Gemma with local data storage, allowing users to train models on device with their own data? That would take personalization to the next level"

What's the local execution timeline looking like, and will it actually run offline or still phone home for routing?

Local routing first is a smart wedge. Once full local execution lands, the interesting part will be how clearly the CLI shows when a task stayed on device versus when it escalated to a hosted model.

um based?

Just stop to tweeting and add API expose to "Android edge gallery" app. Nothing more.

Local for the easy stuff, cloud for when it matters.

Local routing changes the latency math for tool-use loops. Curious if the router decides per-call or per-conversation. Per-call gets you cheap reflection without losing the bigger model on planning.

Esto es genial! buena noticia sin duda

😊 great thanks! Which Gemma 4 models are optimized for the CLI?

Any chance to see local Gemma in Antigravity?

Love this move by Google with intelligent local ↔ cloud routing for Gemma in the Gemini CLI! 📷This is exactly what our users praise most about Local Live Translator — the ability to effortlessly switch between fully local AI (on-device Whisper + Mistral 7B on Apple Silicon) and cloud models for real-time live translation. Blazing fast, private, and flexible. Native Mac app with a free iOS companion too! Keep pushing the local + hybrid AI frontier! 📷 🚀

Great, I can make a Snake Game.

What are the current bottlenecks in the models we have right now? In terms of power, search and speed? Can we make them generate results fast?

finally, but how to download and use Gemma 4 models???

