正在加载视频...

视频加载失败

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

319,142 次观看 • 3 个月前 •via X (Twitter)

38 条评论

antirez 的头像
antirez3 个月前

P.s. DwarfStar

antirez 的头像
antirez3 个月前

In full residency in an M3 ultra it goes to 15 t/s. With indexed attention the performance is retained for long contexts work as well, the slope is acceptable. Will post a video of that as well.

Ivan Fioravanti 的头像
Ivan Fioravanti3 个月前

Great job! 🙌🏻 Is DwarfStar ready to run GLM 5.2 on M3 Ultra? Can I benchmark it? Or better wait a more stable release? Thanks 🙏

Matteo Collina 的头像
Matteo Collina3 个月前

wow! Can I test this on the spark?

Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩 的头像
Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩3 个月前

I am glad I am not the only one torturing LLMs with COBOL: GitHub - ricudis/cobocaml: A bidirectional COBOL <> OCaml FFI bridge · GitHub

CymraegKid 的头像
CymraegKid3 个月前

Cool, but way too slow. Using an agent at this speed sounds painful, honestly.

Ljubomir Josifovski 的头像
Ljubomir Josifovski3 个月前

Nice! Have you seen this? - M3 Ultra (819 GB/s, but only 26 TFLOPS) + DGX Spark (only 273 GB/s, but 100 TFLOPS) + 10 GbE connect.

Timur Yessenov 的头像
Timur Yessenov3 个月前

This is the local-AI point I care about. Not “my laptop beats the cloud.” It doesn’t. But a 128GB Mac running GLM 5.2 with SSD streaming is a real fallback for ugly long tasks when cloud access, price, or rate limits get weird.

ree 的头像
ree3 个月前

I get 2.5t on old 64 x gold xeon CPU only with 544gb. Try out the IK fork of llama-server (not sure how well it does on apple though). M5 Max should get 3x that.

antirez 的头像
antirez3 个月前

It goes 15 t/s with full residency in memory. This is SSD streamed.

ree 的头像
ree3 个月前

Oh, ssd offload with expert streaming? Cool. That sounds like an ssd killer lol. I run mine full context I wonder what full context speed gives you.

François Fleuret 的头像
François Fleuret2 个月前

I do not understand. It's speculative decoding? It cannot stream the full model 20 times per s.

witcheer 的头像
witcheer3 个月前

that is super cool

Rameswar 的头像
Rameswar3 个月前

apple silicon's unified memory makes this kind of extreme offloading more practical than on discrete gpu setups the m5 max 128gb is still expensive, but it's becoming a realistic high end local workstation instead of requiring server class hardware

Sergio Suave 的头像
Sergio Suave3 个月前

Very cool, but painful too watch at that tps...

GeekPark 的头像
GeekPark3 个月前

Zhipu just shipped what the open-source community needed: frontier-class performance that fits in a home office.

Azure | 数据分析万物 的头像
Azure | 数据分析万物3 个月前

This is a waste of M5 Max...

Jakove_HR 的头像
Jakove_HR3 个月前

@ivanfioravanti The goat

Tiezhen WANG 的头像
Tiezhen WANG3 个月前

Nice work! Are you running MTP?

miko 的头像
miko3 个月前

Intanto grazie per quello che stai facendo. La stessa "velocità" che ho sul mio pc ma con modelli più scemi di me. Perdona se la domanda è idiota come Gemma: è possibile usare dei pesi e switcharli? Es. Per coding, non mi serve la storia d'Italia. Pesi diversi in base all'uso.

Marshall 的头像
Marshall3 个月前

Quantized… that’s going to come with some major drawbacks

Peter Dedene 的头像
Peter Dedene3 个月前

Can we do this on a DGX Spark as well?

mzba 的头像
mzba3 个月前

Was playing around the ssd offload stuff recently and borrowed a lot of cool ideas from ds4 project, it works very well for MOE models..

Charlie Kuo 的头像
Charlie Kuo3 个月前

At that tps it is a stretch calling it “running”…

Peter Morris 的头像
Peter Morris3 个月前

Too slow for me :)

SourceCodeplz 的头像
SourceCodeplz3 个月前

Amazing! I mean it’s running ..

Daksh Trehan 的头像
Daksh Trehan3 个月前

the bit i'd watch: tok/s when a token routes to an expert that isn't in the 75GiB lockable cache yet. MoE + SSD streaming should fly on cache hits and stall on the cold ones. are you getting steady throughput or a sawtooth as the working set shifts?

Human Meteorite 的头像
Human Meteorite3 个月前

@ivanfioravanti You make me want to quit my office job and go all in

Dragos Roua 的头像
Dragos Roua3 个月前

This looks very promising! I assume a new session in a different terminal will slow this down a bit, right? So we’re limited to one flow / machine.

Petit Biscuit 的头像
Petit Biscuit3 个月前

not usable

Vandos ❓ 的头像
Vandos ❓3 个月前

The tradeoff here is real, not free: SSD streaming for experts means latency spikes whenever routing hits an expert not already cached. Worth asking what the actual tokens/sec looks like under load, since the headline isn’t ‘it fits,’ it’s ‘how usable is it once it fits.

Donato Molino 🦆 的头像
Donato Molino 🦆3 个月前

What is the problem that causes 3t/s?

Petrus Pennanen 的头像
Petrus Pennanen2 个月前

Wow I had missed this, want to try on my macbook! Where can I find instructions? Thanks :)

John Davis | Solar Strategies 的头像
John Davis | Solar Strategies3 个月前

Asus GX10?

Andrew Dubin 的头像
Andrew Dubin3 个月前

Has anyone tested this (or DS4 Pro) on M4 Max 128GB? I know the SSD is supposed to be 2x slower, but my tokens were closer to the 0.1-0.5 range...

Eustus Mwaniki 的头像
Eustus Mwaniki3 个月前

Tell Dario to come outside.

云儿淡风也轻 的头像
云儿淡风也轻3 个月前

太棒了,等不及测试了

Ferbin 的头像
Ferbin3 个月前

Does the routing + 2.6-bit quantization actually batch well enough to hide SSD latency, or is it mostly "works on paper" territory?

相关视频