Video wird geladen...
Video konnte nicht geladen werden
This is MiniMax-M2.5 MLX running in LM Studio on an Apple Mac Studio M3 Ultra 512GB. Fast enough out of the box for hosting OpenClaw, n8n workflows, and Open WebUI for the team.
76,702 Aufrufe • vor 7 Monaten •via X (Twitter)
43 Kommentare

try inferencer, it has faster pp speed and prompt caching.

This was more of a quick setup and get it working.

For the novice, that’s $10k for the Mac Studio and a 45–60-day waiting period.

When I setup my N8N workflows my API costs almost buried me and the token use ran out quickly. It was easy to setup, now I have agents assigned to those tasks that they post directly to Reddit, LinkedIN, moltbook, etc. I have a new server coming, not a Mac, but hosting my own like you is something I will be doing soon. Thank you for posting your information and experiences here for others like me to learn from as it helps all of us. I hope to assist back when I can as well! Any way, good on you! :)

What does the time to first token look like when you throw it a real prompt? This has been a consistent holdup for me with my M4 MacBook Pro. It's great at processing simple prompts but the moment you throw it a 100k request; you're waiting 1min for prompt processing. I'm wondering how much better the Ultra chip is for that.

This is what past the event horizon looks like. The model ships on a desk. The cost drops to electricity. What matters after that is the governance layer, not the weights.

That's awesome. Do you know the tokens per second off the top of your head?

8-bit was 28-35Tps there

Wow that’s awesome, thank you so much Patrick!

It’s a badass model. Just finished benchmarking a few models with an agent to agent language we created yesterday. I run my swarms local, air-gapped with a no dependency runtime and i need all the token savings i can squeeze out of it.

Better than opus?

Opus 4.6, no. It is also much smarter than gpt-oss-120b from what we have seen thus far. So we can replace gpt-oss-120b not just for OpenClaw, but also in our n8n workflows that require accuracy

Prompt processing is shit. OpenClaw would take 3 days for a single task.

This is the dream setup for teams that want full local control. M3 Ultra handles OpenClaw + local LLM inference without breaking a sweat. Interesting to see the split forming: power users going local hardware like this, while non-technical teams are gravitating toward managed cloud hosting to skip the ops work entirely. Both paths are valid — depends on whether you want to own the stack or own the outcome.

Only cost you a small 10k 💸

the m3 ultra is serious overkill but i get it - you want headroom for running multiple openclaw instances + local models without throttling. once you go there you don't look back

35tps?cannot reach 50?

512GB RAM.. that’s €11000+ Did you make calculation? Compare with hosted LLMs?

Can use for open code? And have you try kimi k2.5?

You can run step 3.5 on 128GB

Nice MLX stack. Tip for builders: high-RAM Certified Refurbished Studios still appear, but the scarce ones are gone in minutes.

Local horsepower unlocks serious agent orchestration.

mac studio m3 ultra for openclaw + the full team stack is a flex ngl. what's your power draw looking like?

I’ll try that tomorrow as had nothing but positive feedback on that model. 🪰

Full precision?

@dynemetis 8

@grok can a rtx3090 run this? compare the cost/performance of nvidia tech with apple

Not bad

Nice setup. If you benchmark, it’d be great to include tok/s with batch size, quantization level, and context length since MLX throughput shifts a lot with seq length and unified memory pressure.

And 1 more question, what is the SSD hdd size you choose for running the local LLM?

Context size?

That was 196K

Please tell which quant (if quantized).

This is MLX 8-bit

FYI, this is the MLX 6-bit quant on my M3U/60C/256Gb. Takes ~185Gb.

solid setup. the local model angle makes sense for teams that want everything on-prem. we went the other direction with full k8s managed hosting but for raw local inference throughput a M3 Ultra is hard to beat ngl

Looks performing well.

Also, my nightly tests for datasurface burn 7m tokens….

That's rather fast. For a small investment you can go 100% local.

MiniMax-M2.5 running smoothly on Mac Studio M3 Ultra demonstrates impressive local model performance. Fast enough for OpenClaw hosting plus n8n workflows shows practical edge deployment viability.

512gb unified memory is the sweet spot for local llms right now

thats on an M3. wow still good

nice setup. the m3 ultra is overkill in a good way. more than enough headroom to run a few agents in parallel
