Loading video...

Video Failed to Load

Go Home

It’s happened. Mac Studio is here. Gemma 4 31b Google DeepMind installed, chatting with my main OpenClaw🦞 for $0 in token expenses now... I've burned $5-6k on tokens on my crazy ideas over past few months, so this mac studio should pencil out for me within 3 months or...

851,624 views • 5 months ago •via X (Twitter)

64 Comments

TPA_Patriot's profile picture
TPA_Patriot5 months ago

@GoogleDeepMind @openclaw Hey Jessie! We just received our 512 gb ram studio for our work. I downloaded all 4 - Gemma 4:31b Qwen 3:235b Qwen 3.5:122b Qwen 3.5:35b And said “run a benchmark based on what we do/have set up at our company” and this is what it found. I would have your agent do the same!

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Great suggestion thanks!

Konstantin Gladych's profile picture
Konstantin Gladych5 months ago

@GoogleDeepMind @openclaw Nice! Jessy, You are very welcome to try @atomic_chat_hq as your local models provider. They support google turbo quant that gives your openclaw even larger context window and tool calling for complicated requests.

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw @atomic_chat_hq Cool! Will take a look. I just got the studio today lol so I’m figuring it all out.

Konstantin Gladych's profile picture
Konstantin Gladych5 months ago

@GoogleDeepMind @openclaw @atomic_chat_hq If you struggle setup everything from a command line, we also have openclaw app with build in local models. Open source and free @atomicbot_ai

⭕ AI & Design (Marco)'s profile picture
⭕ AI & Design (Marco)5 months ago

@GoogleDeepMind @openclaw Serious question: Can you list the top 3 things you got out of that $5-6k you spent on tokens that make that expense worth it?

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw 1. Homeschool manager planning all curriculum and logging lessons for our four young children 2. Chief of staff agent that also does accounting work 3. Passion projects like building and launching

⭕ AI & Design (Marco)'s profile picture
⭕ AI & Design (Marco)5 months ago

@GoogleDeepMind @openclaw This all sounds like stuff you could have done with Claude Code for a fraction of the cost.

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw You asked and I’m trying to answer, I only gave three examples because you asked or three. Some people want to be position online others don’t… 🤷‍♀️

⭕ AI & Design (Marco)'s profile picture
⭕ AI & Design (Marco)5 months ago

@GoogleDeepMind @openclaw Ok well I’ll concede that if it works for you at the token cost it incurs, more power to you!

⭕ AI & Design (Marco)'s profile picture
⭕ AI & Design (Marco)5 months ago

@GoogleDeepMind @openclaw I mean can you say in all honesty that this was worth $6k and you think there’s no way you could have done it in a month or two on a $200 Claude Max plan? Because I think you could have. I promise I’m not trying to be a d—k here. I just don’t believe in the ClawdBot hype.

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Rotated two max plans until I got canceled by anthropic

⭕ AI & Design (Marco)'s profile picture
⭕ AI & Design (Marco)5 months ago

@GoogleDeepMind @openclaw You got canceled for using Claude Code? Or for using ClawdBot with it?

Robert Scoble's profile picture
Robert Scoble5 months ago

@GoogleDeepMind @openclaw Here's the model analysis you need to figure out which local model to run on it.

Aiafter50withPops's profile picture
Aiafter50withPops5 months ago

@GoogleDeepMind @openclaw Running @openclaw on my Mac mini M4 for 47 days now. Game changer for a 58 year old truck salesman who didn't know what AI was 2 months ago. The future is local 🦞 #AIafter50withPops

jezza kezza's profile picture
jezza kezza5 months ago

@GoogleDeepMind @openclaw It is a good model, but it is slower on that setup and will go down way more dead ends. Still worth a try, I am a huge fan of local models, but it just not in the league same as a frontier model for some tasks and I would recommend doing some steps on cloud - go hybrid.

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Yes hybrid for sure! My plan is multiple agents working in tandem with some on frontier models still

jezza kezza's profile picture
jezza kezza5 months ago

good to hear :) Works well for us. The spectrum from tactical stuff, which works well locally, to planning, where frontier is great. Will be interesting to see how your mix lands, as local models have had a huge leap lately and there are a lot to choose from - I am still a huge Qwen fan.

Ryan Carson's profile picture
Ryan Carson5 months ago

@peteskomoroch @GoogleDeepMind @openclaw I want to try this

Bigoted No𝕏's profile picture
Bigoted No𝕏5 months ago

@GoogleDeepMind @openclaw I didn't know Amish people used computers

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw 😂💕

Remco's profile picture
Remco5 months ago

@GoogleDeepMind @openclaw There are serious risks for using a small model in Openclaw

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Explain please!

Remco's profile picture
Remco5 months ago

@GoogleDeepMind @openclaw For example: Prompt injection can be done by smarter (larger) models that are able to fool smaller less intelligent models.

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ5 months ago

@jessegenet @GoogleDeepMind @openclaw Prompt injection can be done to any model independently from size, same for jailbreak.

Remco's profile picture
Remco5 months ago

@jessegenet @GoogleDeepMind @openclaw Yes but smarter mods are more protected against that

Remco's profile picture
Remco5 months ago

@jessegenet @GoogleDeepMind @openclaw Also using chatgpt or Claude code service are more safe as those conpanies have extra safety measures build in around the model which a local model does not have

Jesse Genet's profile picture
Jesse Genet5 months ago

@ivanfioravanti @GoogleDeepMind @openclaw I think this is a good debate, but part of that debate should include that along with that safety comes no privacy… anthropic and openAI have an entire record of your chats and a mandate to turn those over if asked by courts etc… so every approach has its own ‘risks’

Remco's profile picture
Remco5 months ago

@ivanfioravanti @GoogleDeepMind @openclaw True but having you model send all your data to an external server in the end is also a privacy issue. 😅

Julien Chaumond's profile picture
Julien Chaumond5 months ago

Starlink <> @huggingface connection is quite fast

Sean Clarke's profile picture
Sean Clarke5 months ago

@GoogleDeepMind @openclaw Hi Jesse would you like to join me on my podcast, love what you have been sharing

Uri Eliabayev's profile picture
Uri Eliabayev5 months ago

@GoogleDeepMind @openclaw I have done the same. I made Claude code operating local LLMs and use them when needed.

Adzo Nb's profile picture
Adzo Nb5 months ago

@GoogleDeepMind @openclaw Qwen 3.5 35B is better than Genma 4… And it’s free 😀

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Yes I have it now!

Crazy Freakin Planet's profile picture
Crazy Freakin Planet5 months ago

Very cool!! Looking forward to seeing those savings and your set up on task based model usage. Timely considering this came from Claude today. One note: starting April 4, third-party harnesses like OpenClaw connected to your Claude account will draw from extra usage instead of from your subscription. If you don’t use them, nothing changes. If you do, the credit and bundles above have you covered.

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Ya I saw that too lol

Kai's profile picture
Kai5 months ago

@GoogleDeepMind @openclaw What configuration is your Mac Studio ?

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw 512 m3

Tyler's profile picture
Tyler5 months ago

@GoogleDeepMind @openclaw What quant are you using? And are you using llama w/ metal or mlx? Using kv cache? I think posts like this are super dangerous because I haven’t seen any real benchmarks of anyone running a 31b model with any real context length work on a Mac. How many t/sec u get?

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw I’m curious what you mean by this, what’s the dangerous part?

Tyler's profile picture
Tyler5 months ago

@GoogleDeepMind @openclaw promoting their setup w/ local models as if it’s tried and true. Many seeing posts like this are dumpin money in buying studios thinking it’s gonna work similar to cloud providers. Reality is 31b on a Mac with any sort of context isn’t there yet Happy to be proven wrong

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw Gotcha, I admit I got the computer today. So I don’t think anyone will think this is a comprehensive review… but don’t you think it’s fair to be bullish on the fast improvement of these models? Like whatever I struggle with using local for now is bound to improve?

Tyler's profile picture
Tyler5 months ago

@GoogleDeepMind @openclaw Hopefully! I’m doin a fair bit of testing on models in the coming week to see what’s actually viable from a model/param/quant/tools perspective on various Mac hardware but haven’t seen any Gemma 4 31b models working smooth yet with any real context length

Greg Mushen's profile picture
Greg Mushen5 months ago

@GoogleDeepMind @openclaw Nice. I'm in the process of getting it setup. Did you use Qwen 3.5 before?

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw I’m going to try the qwen models. Just didn’t get to it today!

KITE AI's profile picture
KITE AI5 months ago

@ClementDelangue @GoogleDeepMind @openclaw Zero token costs reveal the infrastructure arbitrage opportunity. We're seeing the early stages of compute decentralization where local hardware competes directly with cloud inference pricing.

Jyothi Venkat's profile picture
Jyothi Venkat5 months ago

@Scobleizer @GoogleDeepMind @openclaw Jesse this is soo good!! I am planning to do the same w/ Gemma, can’t wait to hear more about schema, thanks

Scott's profile picture
Scott5 months ago

@BrianRoemmele @GoogleDeepMind @openclaw This is the way.

Gareth █████'s profile picture
Gareth █████5 months ago

@GoogleDeepMind @openclaw The 2B model is incredible (4B has issues (fails the strawberry test), even when running Q8). 2B has genuine utility; it can run on anything, and is fast AF. I can see it being the main one I interact with, delegating when needed.

John Serrao's profile picture
John Serrao5 months ago

@GoogleDeepMind @openclaw Try kimi 2.5, surprisingly great

David Rojas's profile picture
David Rojas5 months ago

@GoogleDeepMind @openclaw Are you using the MLX version?

Philip Clark's profile picture
Philip Clark5 months ago

@GoogleDeepMind @openclaw Yesss... love this. Waiting for the new studio to come out this summer before taking the plunge.

WildPinesAI's profile picture
WildPinesAI5 months ago

@Scobleizer @GoogleDeepMind @openclaw 31B dense ranked #3 on Arena running at 20GB quantized on unified memory. The "local models are a toy" era ended quietly yesterday. Cloud APIs aren't dead but the floor for what you can run for $0 just jumped dramatically.

ashen's profile picture
ashen5 months ago

@GoogleDeepMind @openclaw you're low-key cooking here. what is the ram on that studio you have? i think the most important thing you said here is that in a few months ai should get to the point where you really can run intelligent models locally. so why wouldn't you?

Tiller's profile picture
Tiller5 months ago

@GoogleDeepMind @openclaw I’m thinking a MacBook Pro M5 with 64GB would work for most models and it’s actually something one can get in a few weeks not months!

Milan's profile picture
Milan5 months ago

@GoogleDeepMind @openclaw Mac Studio M4 Ultra with 192GB unified memory could run 70B+ models at 4-bit. At 3-bit with better codebooks even more headroom for context. How's the inference speed on Gemma 4 31B?

Sam Ward's profile picture
Sam Ward5 months ago

@GoogleDeepMind @openclaw Same setup here. Mac Minis running our legal agents. The ROI math is obvious once you're burning real money on tokens. But the real win is data never leaving your machine. In regulated industries that's not a nice to have, it's the whole point.

pSacramento's profile picture
pSacramento5 months ago

@ivanfioravanti @GoogleDeepMind @openclaw Thank you so much for sharing! Can’t wait to learn more about your experience with the local models.

Eusebio Resende's profile picture
Eusebio Resende5 months ago

@GoogleDeepMind @openclaw Totally agree. Local inference is the future. In the meantime we need to use all the strategies and tools to save us token costs. Shameless plug on my MCP tool:

Verso's profile picture
Verso5 months ago

@GoogleDeepMind @openclaw Gemma 4 is not even close comparing to Qwen 3.5 / 3.6 for local usage And it's very slow unless you're using 5090 / H100 lol

NoMo • Download Now!'s profile picture
NoMo • Download Now!5 months ago

@GoogleDeepMind @openclaw the real flex isnt the mac studio its the $0 token bill. meanwhile theres founders out here paying $200/mo for claude max and still copy pasting into chatgpt on the side lmao local models are unironically the biggest unlock for indie builders rn

Freddie Morra's profile picture
Freddie Morra5 months ago

@GoogleDeepMind @openclaw If you haven't accomplished anything with $5K of tokens on a frontier model, how is a local model going to get you anywhere?

Jesse Genet's profile picture
Jesse Genet5 months ago

@GoogleDeepMind @openclaw 🧌

Minjune Song's profile picture
Minjune Song5 months ago

@DamiDina @GoogleDeepMind @openclaw Whats mac studio

Related Videos