Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

It’s happened. Mac Studio is here. Gemma 4 31b Google DeepMind installed, chatting with my main OpenClaw🦞 for $0 in token expenses now... I've burned $5-6k on tokens on my crazy ideas over past few months, so this mac studio should pencil out for me within 3 months or...

851,624 Aufrufe • vor 5 Monaten •via X (Twitter)

64 Kommentare

Profilbild von TPA_Patriot
TPA_Patriotvor 5 Monaten

@GoogleDeepMind @openclaw Hey Jessie! We just received our 512 gb ram studio for our work. I downloaded all 4 - Gemma 4:31b Qwen 3:235b Qwen 3.5:122b Qwen 3.5:35b And said “run a benchmark based on what we do/have set up at our company” and this is what it found. I would have your agent do the same!

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Great suggestion thanks!

Profilbild von Konstantin Gladych
Konstantin Gladychvor 5 Monaten

@GoogleDeepMind @openclaw Nice! Jessy, You are very welcome to try @atomic_chat_hq as your local models provider. They support google turbo quant that gives your openclaw even larger context window and tool calling for complicated requests.

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw @atomic_chat_hq Cool! Will take a look. I just got the studio today lol so I’m figuring it all out.

Profilbild von Konstantin Gladych
Konstantin Gladychvor 5 Monaten

@GoogleDeepMind @openclaw @atomic_chat_hq If you struggle setup everything from a command line, we also have openclaw app with build in local models. Open source and free @atomicbot_ai

Profilbild von ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)vor 5 Monaten

@GoogleDeepMind @openclaw Serious question: Can you list the top 3 things you got out of that $5-6k you spent on tokens that make that expense worth it?

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw 1. Homeschool manager planning all curriculum and logging lessons for our four young children 2. Chief of staff agent that also does accounting work 3. Passion projects like building and launching

Profilbild von ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)vor 5 Monaten

@GoogleDeepMind @openclaw This all sounds like stuff you could have done with Claude Code for a fraction of the cost.

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw You asked and I’m trying to answer, I only gave three examples because you asked or three. Some people want to be position online others don’t… 🤷‍♀️

Profilbild von ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)vor 5 Monaten

@GoogleDeepMind @openclaw Ok well I’ll concede that if it works for you at the token cost it incurs, more power to you!

Profilbild von ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)vor 5 Monaten

@GoogleDeepMind @openclaw I mean can you say in all honesty that this was worth $6k and you think there’s no way you could have done it in a month or two on a $200 Claude Max plan? Because I think you could have. I promise I’m not trying to be a d—k here. I just don’t believe in the ClawdBot hype.

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Rotated two max plans until I got canceled by anthropic

Profilbild von ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)vor 5 Monaten

@GoogleDeepMind @openclaw You got canceled for using Claude Code? Or for using ClawdBot with it?

Profilbild von Robert Scoble
Robert Scoblevor 5 Monaten

@GoogleDeepMind @openclaw Here's the model analysis you need to figure out which local model to run on it.

Profilbild von Aiafter50withPops
Aiafter50withPopsvor 5 Monaten

@GoogleDeepMind @openclaw Running @openclaw on my Mac mini M4 for 47 days now. Game changer for a 58 year old truck salesman who didn't know what AI was 2 months ago. The future is local 🦞 #AIafter50withPops

Profilbild von jezza kezza
jezza kezzavor 5 Monaten

@GoogleDeepMind @openclaw It is a good model, but it is slower on that setup and will go down way more dead ends. Still worth a try, I am a huge fan of local models, but it just not in the league same as a frontier model for some tasks and I would recommend doing some steps on cloud - go hybrid.

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Yes hybrid for sure! My plan is multiple agents working in tandem with some on frontier models still

Profilbild von jezza kezza
jezza kezzavor 5 Monaten

good to hear :) Works well for us. The spectrum from tactical stuff, which works well locally, to planning, where frontier is great. Will be interesting to see how your mix lands, as local models have had a huge leap lately and there are a lot to choose from - I am still a huge Qwen fan.

Profilbild von Ryan Carson
Ryan Carsonvor 5 Monaten

@peteskomoroch @GoogleDeepMind @openclaw I want to try this

Profilbild von Bigoted No𝕏
Bigoted No𝕏vor 5 Monaten

@GoogleDeepMind @openclaw I didn't know Amish people used computers

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw 😂💕

Profilbild von Remco
Remcovor 5 Monaten

@GoogleDeepMind @openclaw There are serious risks for using a small model in Openclaw

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Explain please!

Profilbild von Remco
Remcovor 5 Monaten

@GoogleDeepMind @openclaw For example: Prompt injection can be done by smarter (larger) models that are able to fool smaller less intelligent models.

Profilbild von Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅvor 5 Monaten

@jessegenet @GoogleDeepMind @openclaw Prompt injection can be done to any model independently from size, same for jailbreak.

Profilbild von Remco
Remcovor 5 Monaten

@jessegenet @GoogleDeepMind @openclaw Yes but smarter mods are more protected against that

Profilbild von Remco
Remcovor 5 Monaten

@jessegenet @GoogleDeepMind @openclaw Also using chatgpt or Claude code service are more safe as those conpanies have extra safety measures build in around the model which a local model does not have

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@ivanfioravanti @GoogleDeepMind @openclaw I think this is a good debate, but part of that debate should include that along with that safety comes no privacy… anthropic and openAI have an entire record of your chats and a mandate to turn those over if asked by courts etc… so every approach has its own ‘risks’

Profilbild von Remco
Remcovor 5 Monaten

@ivanfioravanti @GoogleDeepMind @openclaw True but having you model send all your data to an external server in the end is also a privacy issue. 😅

Profilbild von Julien Chaumond
Julien Chaumondvor 5 Monaten

Starlink <> @huggingface connection is quite fast

Profilbild von Sean Clarke
Sean Clarkevor 5 Monaten

@GoogleDeepMind @openclaw Hi Jesse would you like to join me on my podcast, love what you have been sharing

Profilbild von Uri Eliabayev
Uri Eliabayevvor 5 Monaten

@GoogleDeepMind @openclaw I have done the same. I made Claude code operating local LLMs and use them when needed.

Profilbild von Adzo Nb
Adzo Nbvor 5 Monaten

@GoogleDeepMind @openclaw Qwen 3.5 35B is better than Genma 4… And it’s free 😀

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Yes I have it now!

Profilbild von Crazy Freakin Planet
Crazy Freakin Planetvor 5 Monaten

Very cool!! Looking forward to seeing those savings and your set up on task based model usage. Timely considering this came from Claude today. One note: starting April 4, third-party harnesses like OpenClaw connected to your Claude account will draw from extra usage instead of from your subscription. If you don’t use them, nothing changes. If you do, the credit and bundles above have you covered.

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Ya I saw that too lol

Profilbild von Kai
Kaivor 5 Monaten

@GoogleDeepMind @openclaw What configuration is your Mac Studio ?

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw 512 m3

Profilbild von Tyler
Tylervor 5 Monaten

@GoogleDeepMind @openclaw What quant are you using? And are you using llama w/ metal or mlx? Using kv cache? I think posts like this are super dangerous because I haven’t seen any real benchmarks of anyone running a 31b model with any real context length work on a Mac. How many t/sec u get?

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw I’m curious what you mean by this, what’s the dangerous part?

Profilbild von Tyler
Tylervor 5 Monaten

@GoogleDeepMind @openclaw promoting their setup w/ local models as if it’s tried and true. Many seeing posts like this are dumpin money in buying studios thinking it’s gonna work similar to cloud providers. Reality is 31b on a Mac with any sort of context isn’t there yet Happy to be proven wrong

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw Gotcha, I admit I got the computer today. So I don’t think anyone will think this is a comprehensive review… but don’t you think it’s fair to be bullish on the fast improvement of these models? Like whatever I struggle with using local for now is bound to improve?

Profilbild von Tyler
Tylervor 5 Monaten

@GoogleDeepMind @openclaw Hopefully! I’m doin a fair bit of testing on models in the coming week to see what’s actually viable from a model/param/quant/tools perspective on various Mac hardware but haven’t seen any Gemma 4 31b models working smooth yet with any real context length

Profilbild von Greg Mushen
Greg Mushenvor 5 Monaten

@GoogleDeepMind @openclaw Nice. I'm in the process of getting it setup. Did you use Qwen 3.5 before?

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw I’m going to try the qwen models. Just didn’t get to it today!

Profilbild von KITE AI
KITE AIvor 5 Monaten

@ClementDelangue @GoogleDeepMind @openclaw Zero token costs reveal the infrastructure arbitrage opportunity. We're seeing the early stages of compute decentralization where local hardware competes directly with cloud inference pricing.

Profilbild von Jyothi Venkat
Jyothi Venkatvor 5 Monaten

@Scobleizer @GoogleDeepMind @openclaw Jesse this is soo good!! I am planning to do the same w/ Gemma, can’t wait to hear more about schema, thanks

Profilbild von Scott
Scottvor 5 Monaten

@BrianRoemmele @GoogleDeepMind @openclaw This is the way.

Profilbild von Gareth █████
Gareth █████vor 5 Monaten

@GoogleDeepMind @openclaw The 2B model is incredible (4B has issues (fails the strawberry test), even when running Q8). 2B has genuine utility; it can run on anything, and is fast AF. I can see it being the main one I interact with, delegating when needed.

Profilbild von John Serrao
John Serraovor 5 Monaten

@GoogleDeepMind @openclaw Try kimi 2.5, surprisingly great

Profilbild von David Rojas
David Rojasvor 5 Monaten

@GoogleDeepMind @openclaw Are you using the MLX version?

Profilbild von Philip Clark
Philip Clarkvor 5 Monaten

@GoogleDeepMind @openclaw Yesss... love this. Waiting for the new studio to come out this summer before taking the plunge.

Profilbild von WildPinesAI
WildPinesAIvor 5 Monaten

@Scobleizer @GoogleDeepMind @openclaw 31B dense ranked #3 on Arena running at 20GB quantized on unified memory. The "local models are a toy" era ended quietly yesterday. Cloud APIs aren't dead but the floor for what you can run for $0 just jumped dramatically.

Profilbild von ashen
ashenvor 5 Monaten

@GoogleDeepMind @openclaw you're low-key cooking here. what is the ram on that studio you have? i think the most important thing you said here is that in a few months ai should get to the point where you really can run intelligent models locally. so why wouldn't you?

Profilbild von Tiller
Tillervor 5 Monaten

@GoogleDeepMind @openclaw I’m thinking a MacBook Pro M5 with 64GB would work for most models and it’s actually something one can get in a few weeks not months!

Profilbild von Milan
Milanvor 5 Monaten

@GoogleDeepMind @openclaw Mac Studio M4 Ultra with 192GB unified memory could run 70B+ models at 4-bit. At 3-bit with better codebooks even more headroom for context. How's the inference speed on Gemma 4 31B?

Profilbild von Sam Ward
Sam Wardvor 5 Monaten

@GoogleDeepMind @openclaw Same setup here. Mac Minis running our legal agents. The ROI math is obvious once you're burning real money on tokens. But the real win is data never leaving your machine. In regulated industries that's not a nice to have, it's the whole point.

Profilbild von pSacramento
pSacramentovor 5 Monaten

@ivanfioravanti @GoogleDeepMind @openclaw Thank you so much for sharing! Can’t wait to learn more about your experience with the local models.

Profilbild von Eusebio Resende
Eusebio Resendevor 5 Monaten

@GoogleDeepMind @openclaw Totally agree. Local inference is the future. In the meantime we need to use all the strategies and tools to save us token costs. Shameless plug on my MCP tool:

Profilbild von Verso
Versovor 5 Monaten

@GoogleDeepMind @openclaw Gemma 4 is not even close comparing to Qwen 3.5 / 3.6 for local usage And it's very slow unless you're using 5090 / H100 lol

Profilbild von NoMo • Download Now!
NoMo • Download Now!vor 5 Monaten

@GoogleDeepMind @openclaw the real flex isnt the mac studio its the $0 token bill. meanwhile theres founders out here paying $200/mo for claude max and still copy pasting into chatgpt on the side lmao local models are unironically the biggest unlock for indie builders rn

Profilbild von Freddie Morra
Freddie Morravor 5 Monaten

@GoogleDeepMind @openclaw If you haven't accomplished anything with $5K of tokens on a frontier model, how is a local model going to get you anywhere?

Profilbild von Jesse Genet
Jesse Genetvor 5 Monaten

@GoogleDeepMind @openclaw 🧌

Profilbild von Minjune Song
Minjune Songvor 5 Monaten

@DamiDina @GoogleDeepMind @openclaw Whats mac studio

Ähnliche Videos