Загрузка видео...

Не удалось загрузить видео

На главную

It’s happened. Mac Studio is here. Gemma 4 31b Google DeepMind installed, chatting with my main OpenClaw🦞 for $0 in token expenses now... I've burned $5-6k on tokens on my crazy ideas over past few months, so this mac studio should pencil out for me within 3 months or...

851,624 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 64

Фото профиля TPA_Patriot
TPA_Patriot5 месяцев назад

@GoogleDeepMind @openclaw Hey Jessie! We just received our 512 gb ram studio for our work. I downloaded all 4 - Gemma 4:31b Qwen 3:235b Qwen 3.5:122b Qwen 3.5:35b And said “run a benchmark based on what we do/have set up at our company” and this is what it found. I would have your agent do the same!

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Great suggestion thanks!

Фото профиля Konstantin Gladych
Konstantin Gladych5 месяцев назад

@GoogleDeepMind @openclaw Nice! Jessy, You are very welcome to try @atomic_chat_hq as your local models provider. They support google turbo quant that gives your openclaw even larger context window and tool calling for complicated requests.

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw @atomic_chat_hq Cool! Will take a look. I just got the studio today lol so I’m figuring it all out.

Фото профиля Konstantin Gladych
Konstantin Gladych5 месяцев назад

@GoogleDeepMind @openclaw @atomic_chat_hq If you struggle setup everything from a command line, we also have openclaw app with build in local models. Open source and free @atomicbot_ai

Фото профиля ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)5 месяцев назад

@GoogleDeepMind @openclaw Serious question: Can you list the top 3 things you got out of that $5-6k you spent on tokens that make that expense worth it?

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw 1. Homeschool manager planning all curriculum and logging lessons for our four young children 2. Chief of staff agent that also does accounting work 3. Passion projects like building and launching

Фото профиля ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)5 месяцев назад

@GoogleDeepMind @openclaw This all sounds like stuff you could have done with Claude Code for a fraction of the cost.

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw You asked and I’m trying to answer, I only gave three examples because you asked or three. Some people want to be position online others don’t… 🤷‍♀️

Фото профиля ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)5 месяцев назад

@GoogleDeepMind @openclaw Ok well I’ll concede that if it works for you at the token cost it incurs, more power to you!

Фото профиля ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)5 месяцев назад

@GoogleDeepMind @openclaw I mean can you say in all honesty that this was worth $6k and you think there’s no way you could have done it in a month or two on a $200 Claude Max plan? Because I think you could have. I promise I’m not trying to be a d—k here. I just don’t believe in the ClawdBot hype.

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Rotated two max plans until I got canceled by anthropic

Фото профиля ⭕ AI & Design (Marco)
⭕ AI & Design (Marco)5 месяцев назад

@GoogleDeepMind @openclaw You got canceled for using Claude Code? Or for using ClawdBot with it?

Фото профиля Robert Scoble
Robert Scoble5 месяцев назад

@GoogleDeepMind @openclaw Here's the model analysis you need to figure out which local model to run on it.

Фото профиля Aiafter50withPops
Aiafter50withPops5 месяцев назад

@GoogleDeepMind @openclaw Running @openclaw on my Mac mini M4 for 47 days now. Game changer for a 58 year old truck salesman who didn't know what AI was 2 months ago. The future is local 🦞 #AIafter50withPops

Фото профиля jezza kezza
jezza kezza5 месяцев назад

@GoogleDeepMind @openclaw It is a good model, but it is slower on that setup and will go down way more dead ends. Still worth a try, I am a huge fan of local models, but it just not in the league same as a frontier model for some tasks and I would recommend doing some steps on cloud - go hybrid.

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Yes hybrid for sure! My plan is multiple agents working in tandem with some on frontier models still

Фото профиля jezza kezza
jezza kezza5 месяцев назад

good to hear :) Works well for us. The spectrum from tactical stuff, which works well locally, to planning, where frontier is great. Will be interesting to see how your mix lands, as local models have had a huge leap lately and there are a lot to choose from - I am still a huge Qwen fan.

Фото профиля Ryan Carson
Ryan Carson5 месяцев назад

@peteskomoroch @GoogleDeepMind @openclaw I want to try this

Фото профиля Bigoted No𝕏
Bigoted No𝕏5 месяцев назад

@GoogleDeepMind @openclaw I didn't know Amish people used computers

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw 😂💕

Фото профиля Remco
Remco5 месяцев назад

@GoogleDeepMind @openclaw There are serious risks for using a small model in Openclaw

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Explain please!

Фото профиля Remco
Remco5 месяцев назад

@GoogleDeepMind @openclaw For example: Prompt injection can be done by smarter (larger) models that are able to fool smaller less intelligent models.

Фото профиля Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅ5 месяцев назад

@jessegenet @GoogleDeepMind @openclaw Prompt injection can be done to any model independently from size, same for jailbreak.

Фото профиля Remco
Remco5 месяцев назад

@jessegenet @GoogleDeepMind @openclaw Yes but smarter mods are more protected against that

Фото профиля Remco
Remco5 месяцев назад

@jessegenet @GoogleDeepMind @openclaw Also using chatgpt or Claude code service are more safe as those conpanies have extra safety measures build in around the model which a local model does not have

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@ivanfioravanti @GoogleDeepMind @openclaw I think this is a good debate, but part of that debate should include that along with that safety comes no privacy… anthropic and openAI have an entire record of your chats and a mandate to turn those over if asked by courts etc… so every approach has its own ‘risks’

Фото профиля Remco
Remco5 месяцев назад

@ivanfioravanti @GoogleDeepMind @openclaw True but having you model send all your data to an external server in the end is also a privacy issue. 😅

Фото профиля Julien Chaumond
Julien Chaumond5 месяцев назад

Starlink <> @huggingface connection is quite fast

Фото профиля Sean Clarke
Sean Clarke5 месяцев назад

@GoogleDeepMind @openclaw Hi Jesse would you like to join me on my podcast, love what you have been sharing

Фото профиля Uri Eliabayev
Uri Eliabayev5 месяцев назад

@GoogleDeepMind @openclaw I have done the same. I made Claude code operating local LLMs and use them when needed.

Фото профиля Adzo Nb
Adzo Nb5 месяцев назад

@GoogleDeepMind @openclaw Qwen 3.5 35B is better than Genma 4… And it’s free 😀

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Yes I have it now!

Фото профиля Crazy Freakin Planet
Crazy Freakin Planet5 месяцев назад

Very cool!! Looking forward to seeing those savings and your set up on task based model usage. Timely considering this came from Claude today. One note: starting April 4, third-party harnesses like OpenClaw connected to your Claude account will draw from extra usage instead of from your subscription. If you don’t use them, nothing changes. If you do, the credit and bundles above have you covered.

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Ya I saw that too lol

Фото профиля Kai
Kai5 месяцев назад

@GoogleDeepMind @openclaw What configuration is your Mac Studio ?

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw 512 m3

Фото профиля Tyler
Tyler5 месяцев назад

@GoogleDeepMind @openclaw What quant are you using? And are you using llama w/ metal or mlx? Using kv cache? I think posts like this are super dangerous because I haven’t seen any real benchmarks of anyone running a 31b model with any real context length work on a Mac. How many t/sec u get?

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw I’m curious what you mean by this, what’s the dangerous part?

Фото профиля Tyler
Tyler5 месяцев назад

@GoogleDeepMind @openclaw promoting their setup w/ local models as if it’s tried and true. Many seeing posts like this are dumpin money in buying studios thinking it’s gonna work similar to cloud providers. Reality is 31b on a Mac with any sort of context isn’t there yet Happy to be proven wrong

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw Gotcha, I admit I got the computer today. So I don’t think anyone will think this is a comprehensive review… but don’t you think it’s fair to be bullish on the fast improvement of these models? Like whatever I struggle with using local for now is bound to improve?

Фото профиля Tyler
Tyler5 месяцев назад

@GoogleDeepMind @openclaw Hopefully! I’m doin a fair bit of testing on models in the coming week to see what’s actually viable from a model/param/quant/tools perspective on various Mac hardware but haven’t seen any Gemma 4 31b models working smooth yet with any real context length

Фото профиля Greg Mushen
Greg Mushen5 месяцев назад

@GoogleDeepMind @openclaw Nice. I'm in the process of getting it setup. Did you use Qwen 3.5 before?

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw I’m going to try the qwen models. Just didn’t get to it today!

Фото профиля KITE AI
KITE AI5 месяцев назад

@ClementDelangue @GoogleDeepMind @openclaw Zero token costs reveal the infrastructure arbitrage opportunity. We're seeing the early stages of compute decentralization where local hardware competes directly with cloud inference pricing.

Фото профиля Jyothi Venkat
Jyothi Venkat5 месяцев назад

@Scobleizer @GoogleDeepMind @openclaw Jesse this is soo good!! I am planning to do the same w/ Gemma, can’t wait to hear more about schema, thanks

Фото профиля Scott
Scott5 месяцев назад

@BrianRoemmele @GoogleDeepMind @openclaw This is the way.

Фото профиля Gareth █████
Gareth █████5 месяцев назад

@GoogleDeepMind @openclaw The 2B model is incredible (4B has issues (fails the strawberry test), even when running Q8). 2B has genuine utility; it can run on anything, and is fast AF. I can see it being the main one I interact with, delegating when needed.

Фото профиля John Serrao
John Serrao5 месяцев назад

@GoogleDeepMind @openclaw Try kimi 2.5, surprisingly great

Фото профиля David Rojas
David Rojas5 месяцев назад

@GoogleDeepMind @openclaw Are you using the MLX version?

Фото профиля Philip Clark
Philip Clark5 месяцев назад

@GoogleDeepMind @openclaw Yesss... love this. Waiting for the new studio to come out this summer before taking the plunge.

Фото профиля WildPinesAI
WildPinesAI5 месяцев назад

@Scobleizer @GoogleDeepMind @openclaw 31B dense ranked #3 on Arena running at 20GB quantized on unified memory. The "local models are a toy" era ended quietly yesterday. Cloud APIs aren't dead but the floor for what you can run for $0 just jumped dramatically.

Фото профиля ashen
ashen5 месяцев назад

@GoogleDeepMind @openclaw you're low-key cooking here. what is the ram on that studio you have? i think the most important thing you said here is that in a few months ai should get to the point where you really can run intelligent models locally. so why wouldn't you?

Фото профиля Tiller
Tiller5 месяцев назад

@GoogleDeepMind @openclaw I’m thinking a MacBook Pro M5 with 64GB would work for most models and it’s actually something one can get in a few weeks not months!

Фото профиля Milan
Milan5 месяцев назад

@GoogleDeepMind @openclaw Mac Studio M4 Ultra with 192GB unified memory could run 70B+ models at 4-bit. At 3-bit with better codebooks even more headroom for context. How's the inference speed on Gemma 4 31B?

Фото профиля Sam Ward
Sam Ward5 месяцев назад

@GoogleDeepMind @openclaw Same setup here. Mac Minis running our legal agents. The ROI math is obvious once you're burning real money on tokens. But the real win is data never leaving your machine. In regulated industries that's not a nice to have, it's the whole point.

Фото профиля pSacramento
pSacramento5 месяцев назад

@ivanfioravanti @GoogleDeepMind @openclaw Thank you so much for sharing! Can’t wait to learn more about your experience with the local models.

Фото профиля Eusebio Resende
Eusebio Resende5 месяцев назад

@GoogleDeepMind @openclaw Totally agree. Local inference is the future. In the meantime we need to use all the strategies and tools to save us token costs. Shameless plug on my MCP tool:

Фото профиля Verso
Verso5 месяцев назад

@GoogleDeepMind @openclaw Gemma 4 is not even close comparing to Qwen 3.5 / 3.6 for local usage And it's very slow unless you're using 5090 / H100 lol

Фото профиля NoMo • Download Now!
NoMo • Download Now!5 месяцев назад

@GoogleDeepMind @openclaw the real flex isnt the mac studio its the $0 token bill. meanwhile theres founders out here paying $200/mo for claude max and still copy pasting into chatgpt on the side lmao local models are unironically the biggest unlock for indie builders rn

Фото профиля Freddie Morra
Freddie Morra5 месяцев назад

@GoogleDeepMind @openclaw If you haven't accomplished anything with $5K of tokens on a frontier model, how is a local model going to get you anywhere?

Фото профиля Jesse Genet
Jesse Genet5 месяцев назад

@GoogleDeepMind @openclaw 🧌

Фото профиля Minjune Song
Minjune Song5 месяцев назад

@DamiDina @GoogleDeepMind @openclaw Whats mac studio

Похожие видео