Loading video...

Video Failed to Load

Go Home

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

319,142 views • 3 months ago •via X (Twitter)

38 Comments

antirez's profile picture
antirez3 months ago

P.s. DwarfStar

antirez's profile picture
antirez3 months ago

In full residency in an M3 ultra it goes to 15 t/s. With indexed attention the performance is retained for long contexts work as well, the slope is acceptable. Will post a video of that as well.

Ivan Fioravanti's profile picture
Ivan Fioravanti3 months ago

Great job! 🙌🏻 Is DwarfStar ready to run GLM 5.2 on M3 Ultra? Can I benchmark it? Or better wait a more stable release? Thanks 🙏

Matteo Collina's profile picture
Matteo Collina3 months ago

wow! Can I test this on the spark?

Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩's profile picture
Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩3 months ago

I am glad I am not the only one torturing LLMs with COBOL: GitHub - ricudis/cobocaml: A bidirectional COBOL <> OCaml FFI bridge · GitHub

CymraegKid's profile picture
CymraegKid3 months ago

Cool, but way too slow. Using an agent at this speed sounds painful, honestly.

Ljubomir Josifovski's profile picture
Ljubomir Josifovski3 months ago

Nice! Have you seen this? - M3 Ultra (819 GB/s, but only 26 TFLOPS) + DGX Spark (only 273 GB/s, but 100 TFLOPS) + 10 GbE connect.

Timur Yessenov's profile picture
Timur Yessenov3 months ago

This is the local-AI point I care about. Not “my laptop beats the cloud.” It doesn’t. But a 128GB Mac running GLM 5.2 with SSD streaming is a real fallback for ugly long tasks when cloud access, price, or rate limits get weird.

ree's profile picture
ree3 months ago

I get 2.5t on old 64 x gold xeon CPU only with 544gb. Try out the IK fork of llama-server (not sure how well it does on apple though). M5 Max should get 3x that.

antirez's profile picture
antirez3 months ago

It goes 15 t/s with full residency in memory. This is SSD streamed.

ree's profile picture
ree3 months ago

Oh, ssd offload with expert streaming? Cool. That sounds like an ssd killer lol. I run mine full context I wonder what full context speed gives you.

François Fleuret's profile picture
François Fleuret2 months ago

I do not understand. It's speculative decoding? It cannot stream the full model 20 times per s.

witcheer's profile picture
witcheer3 months ago

that is super cool

Rameswar's profile picture
Rameswar3 months ago

apple silicon's unified memory makes this kind of extreme offloading more practical than on discrete gpu setups the m5 max 128gb is still expensive, but it's becoming a realistic high end local workstation instead of requiring server class hardware

Sergio Suave's profile picture
Sergio Suave3 months ago

Very cool, but painful too watch at that tps...

GeekPark's profile picture
GeekPark3 months ago

Zhipu just shipped what the open-source community needed: frontier-class performance that fits in a home office.

Azure | 数据分析万物's profile picture
Azure | 数据分析万物3 months ago

This is a waste of M5 Max...

Jakove_HR's profile picture
Jakove_HR3 months ago

@ivanfioravanti The goat

Tiezhen WANG's profile picture
Tiezhen WANG3 months ago

Nice work! Are you running MTP?

miko's profile picture
miko3 months ago

Intanto grazie per quello che stai facendo. La stessa "velocità" che ho sul mio pc ma con modelli più scemi di me. Perdona se la domanda è idiota come Gemma: è possibile usare dei pesi e switcharli? Es. Per coding, non mi serve la storia d'Italia. Pesi diversi in base all'uso.

Marshall's profile picture
Marshall3 months ago

Quantized… that’s going to come with some major drawbacks

Peter Dedene's profile picture
Peter Dedene3 months ago

Can we do this on a DGX Spark as well?

mzba's profile picture
mzba3 months ago

Was playing around the ssd offload stuff recently and borrowed a lot of cool ideas from ds4 project, it works very well for MOE models..

Charlie Kuo's profile picture
Charlie Kuo3 months ago

At that tps it is a stretch calling it “running”…

Peter Morris's profile picture
Peter Morris3 months ago

Too slow for me :)

SourceCodeplz's profile picture
SourceCodeplz3 months ago

Amazing! I mean it’s running ..

Daksh Trehan's profile picture
Daksh Trehan3 months ago

the bit i'd watch: tok/s when a token routes to an expert that isn't in the 75GiB lockable cache yet. MoE + SSD streaming should fly on cache hits and stall on the cold ones. are you getting steady throughput or a sawtooth as the working set shifts?

Human Meteorite's profile picture
Human Meteorite3 months ago

@ivanfioravanti You make me want to quit my office job and go all in

Dragos Roua's profile picture
Dragos Roua3 months ago

This looks very promising! I assume a new session in a different terminal will slow this down a bit, right? So we’re limited to one flow / machine.

Petit Biscuit's profile picture
Petit Biscuit3 months ago

not usable

Vandos ❓'s profile picture
Vandos ❓3 months ago

The tradeoff here is real, not free: SSD streaming for experts means latency spikes whenever routing hits an expert not already cached. Worth asking what the actual tokens/sec looks like under load, since the headline isn’t ‘it fits,’ it’s ‘how usable is it once it fits.

Donato Molino 🦆's profile picture
Donato Molino 🦆3 months ago

What is the problem that causes 3t/s?

Petrus Pennanen's profile picture
Petrus Pennanen2 months ago

Wow I had missed this, want to try on my macbook! Where can I find instructions? Thanks :)

John Davis | Solar Strategies's profile picture
John Davis | Solar Strategies3 months ago

Asus GX10?

Andrew Dubin's profile picture
Andrew Dubin3 months ago

Has anyone tested this (or DS4 Pro) on M4 Max 128GB? I know the SSD is supposed to be 2x slower, but my tokens were closer to the 0.1-0.5 range...

Eustus Mwaniki's profile picture
Eustus Mwaniki3 months ago

Tell Dario to come outside.

云儿淡风也轻's profile picture
云儿淡风也轻3 months ago

太棒了,等不及测试了

Ferbin's profile picture
Ferbin3 months ago

Does the routing + 2.6-bit quantization actually batch well enough to hide SSD latency, or is it mostly "works on paper" territory?

Related Videos