Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

This is MiniMax-M2.5 MLX running in LM Studio on an Apple Mac Studio M3 Ultra 512GB. Fast enough out of the box for hosting OpenClaw, n8n workflows, and Open WebUI for the team.

76,702 Aufrufe • vor 7 Monaten •via X (Twitter)

43 Kommentare

Profilbild von Wei Jiang
Wei Jiangvor 7 Monaten

try inferencer, it has faster pp speed and prompt caching.

Profilbild von Patrick J Kennedy
Patrick J Kennedyvor 7 Monaten

This was more of a quick setup and get it working.

Profilbild von Newyork Pandian
Newyork Pandianvor 7 Monaten

For the novice, that’s $10k for the Mac Studio and a 45–60-day waiting period.

Profilbild von WhiteFlour
WhiteFlourvor 7 Monaten

When I setup my N8N workflows my API costs almost buried me and the token use ran out quickly. It was easy to setup, now I have agents assigned to those tasks that they post directly to Reddit, LinkedIN, moltbook, etc. I have a new server coming, not a Mac, but hosting my own like you is something I will be doing soon. Thank you for posting your information and experiences here for others like me to learn from as it helps all of us. I hope to assist back when I can as well! Any way, good on you! :)

Profilbild von Ryan R. Hughes
Ryan R. Hughesvor 7 Monaten

What does the time to first token look like when you throw it a real prompt? This has been a consistent holdup for me with my M4 MacBook Pro. It's great at processing simple prompts but the moment you throw it a 100k request; you're waiting 1min for prompt processing. I'm wondering how much better the Ultra chip is for that.

Profilbild von Matthew J. Maughan, PharmD, MHCDS, CPEL
Matthew J. Maughan, PharmD, MHCDS, CPELvor 7 Monaten

This is what past the event horizon looks like. The model ships on a desk. The cost drops to electricity. What matters after that is the governance layer, not the weights.

Profilbild von John Dawson
John Dawsonvor 7 Monaten

That's awesome. Do you know the tokens per second off the top of your head?

Profilbild von Patrick J Kennedy
Patrick J Kennedyvor 7 Monaten

8-bit was 28-35Tps there

Profilbild von John Dawson
John Dawsonvor 7 Monaten

Wow that’s awesome, thank you so much Patrick!

Profilbild von ExileAI
ExileAIvor 7 Monaten

It’s a badass model. Just finished benchmarking a few models with an agent to agent language we created yesterday. I run my swarms local, air-gapped with a no dependency runtime and i need all the token savings i can squeeze out of it.

Profilbild von Jakub Adamowicz
Jakub Adamowiczvor 7 Monaten

Better than opus?

Profilbild von Patrick J Kennedy
Patrick J Kennedyvor 7 Monaten

Opus 4.6, no. It is also much smarter than gpt-oss-120b from what we have seen thus far. So we can replace gpt-oss-120b not just for OpenClaw, but also in our n8n workflows that require accuracy

Profilbild von Matteo
Matteovor 7 Monaten

Prompt processing is shit. OpenClaw would take 3 days for a single task.

Profilbild von M Asif Rahman
M Asif Rahmanvor 7 Monaten

This is the dream setup for teams that want full local control. M3 Ultra handles OpenClaw + local LLM inference without breaking a sweat. Interesting to see the split forming: power users going local hardware like this, while non-technical teams are gravitating toward managed cloud hosting to skip the ops work entirely. Both paths are valid — depends on whether you want to own the stack or own the outcome.

Profilbild von Jer
Jervor 7 Monaten

Only cost you a small 10k 💸

Profilbild von Michiel V
Michiel Vvor 6 Monaten

the m3 ultra is serious overkill but i get it - you want headroom for running multiple openclaw instances + local models without throttling. once you go there you don't look back

Profilbild von dora li
dora livor 7 Monaten

35tps?cannot reach 50?

Profilbild von dirkpostma
dirkpostmavor 7 Monaten

512GB RAM.. that’s €11000+ Did you make calculation? Compare with hosted LLMs?

Profilbild von 09jul.eth
09jul.ethvor 7 Monaten

Can use for open code? And have you try kimi k2.5?

Profilbild von CryptoBanana
CryptoBananavor 7 Monaten

You can run step 3.5 on 128GB

Profilbild von RefurbSitter
RefurbSittervor 1 Monat

Nice MLX stack. Tip for builders: high-RAM Certified Refurbished Studios still appear, but the scarce ones are gone in minutes.

Profilbild von Kevin Nexus
Kevin Nexusvor 7 Monaten

Local horsepower unlocks serious agent orchestration.

Profilbild von Michiel V
Michiel Vvor 7 Monaten

mac studio m3 ultra for openclaw + the full team stack is a flex ngl. what's your power draw looking like?

Profilbild von fly
flyvor 7 Monaten

I’ll try that tomorrow as had nothing but positive feedback on that model. 🪰

Profilbild von Josh
Joshvor 7 Monaten

Full precision?

Profilbild von Patrick J Kennedy
Patrick J Kennedyvor 7 Monaten

@dynemetis 8

Profilbild von Alexandre Forget
Alexandre Forgetvor 7 Monaten

@grok can a rtx3090 run this? compare the cost/performance of nvidia tech with apple

Profilbild von Sergio
Sergiovor 7 Monaten

Not bad

Profilbild von Donny Li
Donny Livor 7 Monaten

Nice setup. If you benchmark, it’d be great to include tok/s with batch size, quantization level, and context length since MLX throughput shifts a lot with seq length and unified memory pressure.

Profilbild von 09jul.eth
09jul.ethvor 7 Monaten

And 1 more question, what is the SSD hdd size you choose for running the local LLM?

Profilbild von Vova
Vovavor 6 Monaten

Context size?

Profilbild von Patrick J Kennedy
Patrick J Kennedyvor 6 Monaten

That was 196K

Profilbild von moskstraumen
moskstraumenvor 7 Monaten

Please tell which quant (if quantized).

Profilbild von Patrick J Kennedy
Patrick J Kennedyvor 7 Monaten

This is MLX 8-bit

Profilbild von moskstraumen
moskstraumenvor 7 Monaten

FYI, this is the MLX 6-bit quant on my M3U/60C/256Gb. Takes ~185Gb.

Profilbild von Michiel V
Michiel Vvor 7 Monaten

solid setup. the local model angle makes sense for teams that want everything on-prem. we went the other direction with full k8s managed hosting but for raw local inference throughput a M3 Ultra is hard to beat ngl

Profilbild von Vishnu Dhayal Giridhari
Vishnu Dhayal Giridharivor 7 Monaten

Looks performing well.

Profilbild von Billy
Billyvor 7 Monaten

Also, my nightly tests for datasurface burn 7m tokens….

Profilbild von ReverseAI
ReverseAIvor 7 Monaten

That's rather fast. For a small investment you can go 100% local.

Profilbild von Barrak
Barrakvor 7 Monaten

MiniMax-M2.5 running smoothly on Mac Studio M3 Ultra demonstrates impressive local model performance. Fast enough for OpenClaw hosting plus n8n workflows shows practical edge deployment viability.

Profilbild von Marchen
Marchenvor 7 Monaten

512gb unified memory is the sweet spot for local llms right now

Profilbild von sharaff |🦞
sharaff |🦞vor 7 Monaten

thats on an M3. wow still good

Profilbild von Michiel V
Michiel Vvor 7 Monaten

nice setup. the m3 ultra is overkill in a good way. more than enough headroom to run a few agents in parallel

Ähnliche Videos