Video yükleniyor...
Video Yüklenemedi
We didn't just make local models faster. We made them dependable at tool calling. Gemma 4 12B now picks the right tool and runs it correctly, on-device and ~60% faster. Small models just got serious.
49,030 görüntüleme • 3 ay önce •via X (Twitter)
33 Yorum

Free forever. MIT licensed. Native macOS, no browser engine.

Getting a 12B model to reliably pick the right tool on-device is honestly more impressive than just being faster

We've improved tool calling and model loading for all local models (specifically those that are <=20B)

that <=20b range is exactly where most people run local, glad to see that getting real attention

local tool calling has always been possible, here is a 4B model doing it (running on my macbook):

Of course. But osaurus runs it on a sandbox (isolated VM using containerization framework). And we have a diverse set of tools & plugins (mail, notes, calendar etc) which was previously unusable with local models. That's what we've improved & the same is the context of this post!

Local models is the future

And the future belongs to open source!

😃

Quitareis una versión para linux en el futuro?

Yes it's in our roadmap

So this a fine tune ? For better function calling ?

No. This is just a usual quantized version of Gemma 4 with no fine tuning. The real improvement is in the harness itself which makes tool calling reliable for all local models

right tool is the easy half. dependable means it recovers when the call returns garbage, not just when it picks correctly

We already have an agent loop that does exactly this! The models won't give up after a single tool call failure. Failures are fed back to the model as structured error envelopes so the model can correct course on the next turn rather than blindly repeat itself

When ornith?

Soon! It's already in the works

Gonna try that now!

macOS only 😔

好像不能安装mcp

You can connect to remote MCP providers (Settings -> Provider -> Add Provider) and also run Osaurus as an MCP server

This is only for mac, right?

Yes. For now

I wish the Ui was designed in a more Codex-esque approach... ie... more project leaning instead of chat-centric. But... ngl... this interface might be one of the most reliable (open-source) interfaces for tool-calling that I've used thus far. Fast/reliable. Great work!

Appreciate the feedback!

how do you improve the accuracy when it comes to tool calls and hallucinations

getting tool selection right at 12B params is the hard part — once that's reliable, the interesting problem shifts to what happens when you chain multiple correct calls and the side effects accumulate

We've already solved this problem with agent loops!

Great Gemma 4 12B small, on device, 60% faster, and actually reliable at tool calling.

It is tempting to try Osaurus myself. But i am using my mac headless. Lots of good apps look like they are made with desktop use primarily. Which is fine. Guess i am “weird” again. I will read the docs and try it.

Osaurus also exposes a local server with a drop-in OpenAI compatible endpoint. You can point any OpenAI client at it and it still works great when running headless

Precision meets speed. Gemma 4 12B means business.

on-device wins until one bad schema update turns correct calls into silent failures
