Loading video...

Video Failed to Load

Go Home

We didn't just make local models faster. We made them dependable at tool calling. Gemma 4 12B now picks the right tool and runs it correctly, on-device and ~60% faster. Small models just got serious.

49,030 views • 3 months ago •via X (Twitter)

33 Comments

Osaurus's profile picture
Osaurus3 months ago

Free forever. MIT licensed. Native macOS, no browser engine.

AI Mastery Guide's profile picture
AI Mastery Guide3 months ago

Getting a 12B model to reliably pick the right tool on-device is honestly more impressive than just being faster

Osaurus's profile picture
Osaurus3 months ago

We've improved tool calling and model loading for all local models (specifically those that are <=20B)

AI Mastery Guide's profile picture
AI Mastery Guide3 months ago

that <=20b range is exactly where most people run local, glad to see that getting real attention

Chiming Wang's profile picture
Chiming Wang3 months ago

local tool calling has always been possible, here is a 4B model doing it (running on my macbook):

Osaurus's profile picture
Osaurus3 months ago

Of course. But osaurus runs it on a sandbox (isolated VM using containerization framework). And we have a diverse set of tools & plugins (mail, notes, calendar etc) which was previously unusable with local models. That's what we've improved & the same is the context of this post!

VibeRaven's profile picture
VibeRaven3 months ago

Local models is the future

Osaurus's profile picture
Osaurus3 months ago

And the future belongs to open source!

VibeRaven's profile picture
VibeRaven3 months ago

😃

ᐱ ᑎ ᑐ ᒋ ᕮ's profile picture
ᐱ ᑎ ᑐ ᒋ ᕮ3 months ago

Quitareis una versión para linux en el futuro?

Osaurus's profile picture
Osaurus3 months ago

Yes it's in our roadmap

SourceCodeplz's profile picture
SourceCodeplz3 months ago

So this a fine tune ? For better function calling ?

Osaurus's profile picture
Osaurus3 months ago

No. This is just a usual quantized version of Gemma 4 with no fine tuning. The real improvement is in the harness itself which makes tool calling reliable for all local models

Michał Piszczek's profile picture
Michał Piszczek3 months ago

right tool is the easy half. dependable means it recovers when the call returns garbage, not just when it picks correctly

Osaurus's profile picture
Osaurus3 months ago

We already have an agent loop that does exactly this! The models won't give up after a single tool call failure. Failures are fed back to the model as structured error envelopes so the model can correct course on the next turn rather than blindly repeat itself

Alican's profile picture
Alican3 months ago

When ornith?

Osaurus's profile picture
Osaurus3 months ago

Soon! It's already in the works

Santosh | n0c0de.com's profile picture
Santosh | n0c0de.com3 months ago

Gonna try that now!

Dido Realm's profile picture
Dido Realm3 months ago

macOS only 😔

anarkhgatsby's profile picture
anarkhgatsby3 months ago

好像不能安装mcp

Osaurus's profile picture
Osaurus3 months ago

You can connect to remote MCP providers (Settings -> Provider -> Add Provider) and also run Osaurus as an MCP server

Prabha's profile picture
Prabha3 months ago

This is only for mac, right?

Osaurus's profile picture
Osaurus3 months ago

Yes. For now

Stackin' ฿its's profile picture
Stackin' ฿its3 months ago

I wish the Ui was designed in a more Codex-esque approach... ie... more project leaning instead of chat-centric. But... ngl... this interface might be one of the most reliable (open-source) interfaces for tool-calling that I've used thus far. Fast/reliable. Great work!

Osaurus's profile picture
Osaurus3 months ago

Appreciate the feedback!

Shivay Lamba's profile picture
Shivay Lamba3 months ago

how do you improve the accuracy when it comes to tool calls and hallucinations

Milton Yan's profile picture
Milton Yan3 months ago

getting tool selection right at 12B params is the hard part — once that's reliable, the interesting problem shifts to what happens when you chain multiple correct calls and the side effects accumulate

Osaurus's profile picture
Osaurus3 months ago

We've already solved this problem with agent loops!

Alice The Ai Expert's profile picture
Alice The Ai Expert3 months ago

Great Gemma 4 12B small, on device, 60% faster, and actually reliable at tool calling.

The one's profile picture
The one3 months ago

It is tempting to try Osaurus myself. But i am using my mac headless. Lots of good apps look like they are made with desktop use primarily. Which is fine. Guess i am “weird” again. I will read the docs and try it.

Osaurus's profile picture
Osaurus3 months ago

Osaurus also exposes a local server with a drop-in OpenAI compatible endpoint. You can point any OpenAI client at it and it still works great when running headless

RAZA | AI EXPLORER's profile picture
RAZA | AI EXPLORER3 months ago

Precision meets speed. Gemma 4 12B means business.

Sebastian Buzdugan's profile picture
Sebastian Buzdugan3 months ago

on-device wins until one bad schema update turns correct calls into silent failures

Related Videos