正在加载视频...
视频加载失败
GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.
319,142 次观看 • 3 个月前 •via X (Twitter)
38 条评论

P.s. DwarfStar

In full residency in an M3 ultra it goes to 15 t/s. With indexed attention the performance is retained for long contexts work as well, the slope is acceptable. Will post a video of that as well.

Great job! 🙌🏻 Is DwarfStar ready to run GLM 5.2 on M3 Ultra? Can I benchmark it? Or better wait a more stable release? Thanks 🙏

wow! Can I test this on the spark?

I am glad I am not the only one torturing LLMs with COBOL: GitHub - ricudis/cobocaml: A bidirectional COBOL <> OCaml FFI bridge · GitHub

Cool, but way too slow. Using an agent at this speed sounds painful, honestly.

Nice! Have you seen this? - M3 Ultra (819 GB/s, but only 26 TFLOPS) + DGX Spark (only 273 GB/s, but 100 TFLOPS) + 10 GbE connect.

This is the local-AI point I care about. Not “my laptop beats the cloud.” It doesn’t. But a 128GB Mac running GLM 5.2 with SSD streaming is a real fallback for ugly long tasks when cloud access, price, or rate limits get weird.

I get 2.5t on old 64 x gold xeon CPU only with 544gb. Try out the IK fork of llama-server (not sure how well it does on apple though). M5 Max should get 3x that.

It goes 15 t/s with full residency in memory. This is SSD streamed.

Oh, ssd offload with expert streaming? Cool. That sounds like an ssd killer lol. I run mine full context I wonder what full context speed gives you.

I do not understand. It's speculative decoding? It cannot stream the full model 20 times per s.

that is super cool

apple silicon's unified memory makes this kind of extreme offloading more practical than on discrete gpu setups the m5 max 128gb is still expensive, but it's becoming a realistic high end local workstation instead of requiring server class hardware

Very cool, but painful too watch at that tps...

Zhipu just shipped what the open-source community needed: frontier-class performance that fits in a home office.

This is a waste of M5 Max...

@ivanfioravanti The goat

Nice work! Are you running MTP?

Intanto grazie per quello che stai facendo. La stessa "velocità" che ho sul mio pc ma con modelli più scemi di me. Perdona se la domanda è idiota come Gemma: è possibile usare dei pesi e switcharli? Es. Per coding, non mi serve la storia d'Italia. Pesi diversi in base all'uso.

Quantized… that’s going to come with some major drawbacks

Can we do this on a DGX Spark as well?

Was playing around the ssd offload stuff recently and borrowed a lot of cool ideas from ds4 project, it works very well for MOE models..

At that tps it is a stretch calling it “running”…

Too slow for me :)

Amazing! I mean it’s running ..

the bit i'd watch: tok/s when a token routes to an expert that isn't in the 75GiB lockable cache yet. MoE + SSD streaming should fly on cache hits and stall on the cold ones. are you getting steady throughput or a sawtooth as the working set shifts?

@ivanfioravanti You make me want to quit my office job and go all in

This looks very promising! I assume a new session in a different terminal will slow this down a bit, right? So we’re limited to one flow / machine.

not usable

The tradeoff here is real, not free: SSD streaming for experts means latency spikes whenever routing hits an expert not already cached. Worth asking what the actual tokens/sec looks like under load, since the headline isn’t ‘it fits,’ it’s ‘how usable is it once it fits.

What is the problem that causes 3t/s?

Wow I had missed this, want to try on my macbook! Where can I find instructions? Thanks :)

Asus GX10?

Has anyone tested this (or DS4 Pro) on M4 Max 128GB? I know the SSD is supposed to be 2x slower, but my tokens were closer to the 0.1-0.5 range...

Tell Dario to come outside.

太棒了,等不及测试了

Does the routing + 2.6-bit quantization actually batch well enough to hide SSD latency, or is it mostly "works on paper" territory?
