Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

319,142 Aufrufe • vor 3 Monaten •via X (Twitter)

38 Kommentare

Profilbild von antirez
antirezvor 3 Monaten

P.s. DwarfStar

Profilbild von antirez
antirezvor 3 Monaten

In full residency in an M3 ultra it goes to 15 t/s. With indexed attention the performance is retained for long contexts work as well, the slope is acceptable. Will post a video of that as well.

Profilbild von Ivan Fioravanti
Ivan Fioravantivor 3 Monaten

Great job! 🙌🏻 Is DwarfStar ready to run GLM 5.2 on M3 Ultra? Can I benchmark it? Or better wait a more stable release? Thanks 🙏

Profilbild von Matteo Collina
Matteo Collinavor 3 Monaten

wow! Can I test this on the spark?

Profilbild von Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩
Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩vor 3 Monaten

I am glad I am not the only one torturing LLMs with COBOL: GitHub - ricudis/cobocaml: A bidirectional COBOL <> OCaml FFI bridge · GitHub

Profilbild von CymraegKid
CymraegKidvor 3 Monaten

Cool, but way too slow. Using an agent at this speed sounds painful, honestly.

Profilbild von Ljubomir Josifovski
Ljubomir Josifovskivor 3 Monaten

Nice! Have you seen this? - M3 Ultra (819 GB/s, but only 26 TFLOPS) + DGX Spark (only 273 GB/s, but 100 TFLOPS) + 10 GbE connect.

Profilbild von Timur Yessenov
Timur Yessenovvor 3 Monaten

This is the local-AI point I care about. Not “my laptop beats the cloud.” It doesn’t. But a 128GB Mac running GLM 5.2 with SSD streaming is a real fallback for ugly long tasks when cloud access, price, or rate limits get weird.

Profilbild von ree
reevor 3 Monaten

I get 2.5t on old 64 x gold xeon CPU only with 544gb. Try out the IK fork of llama-server (not sure how well it does on apple though). M5 Max should get 3x that.

Profilbild von antirez
antirezvor 3 Monaten

It goes 15 t/s with full residency in memory. This is SSD streamed.

Profilbild von ree
reevor 3 Monaten

Oh, ssd offload with expert streaming? Cool. That sounds like an ssd killer lol. I run mine full context I wonder what full context speed gives you.

Profilbild von François Fleuret
François Fleuretvor 2 Monaten

I do not understand. It's speculative decoding? It cannot stream the full model 20 times per s.

Profilbild von witcheer
witcheervor 3 Monaten

that is super cool

Profilbild von Rameswar
Rameswarvor 3 Monaten

apple silicon's unified memory makes this kind of extreme offloading more practical than on discrete gpu setups the m5 max 128gb is still expensive, but it's becoming a realistic high end local workstation instead of requiring server class hardware

Profilbild von Sergio Suave
Sergio Suavevor 3 Monaten

Very cool, but painful too watch at that tps...

Profilbild von GeekPark
GeekParkvor 3 Monaten

Zhipu just shipped what the open-source community needed: frontier-class performance that fits in a home office.

Profilbild von Azure | 数据分析万物
Azure | 数据分析万物vor 3 Monaten

This is a waste of M5 Max...

Profilbild von Jakove_HR
Jakove_HRvor 3 Monaten

@ivanfioravanti The goat

Profilbild von Tiezhen WANG
Tiezhen WANGvor 3 Monaten

Nice work! Are you running MTP?

Profilbild von miko
mikovor 3 Monaten

Intanto grazie per quello che stai facendo. La stessa "velocità" che ho sul mio pc ma con modelli più scemi di me. Perdona se la domanda è idiota come Gemma: è possibile usare dei pesi e switcharli? Es. Per coding, non mi serve la storia d'Italia. Pesi diversi in base all'uso.

Profilbild von Marshall
Marshallvor 3 Monaten

Quantized… that’s going to come with some major drawbacks

Profilbild von Peter Dedene
Peter Dedenevor 3 Monaten

Can we do this on a DGX Spark as well?

Profilbild von mzba
mzbavor 3 Monaten

Was playing around the ssd offload stuff recently and borrowed a lot of cool ideas from ds4 project, it works very well for MOE models..

Profilbild von Charlie Kuo
Charlie Kuovor 3 Monaten

At that tps it is a stretch calling it “running”…

Profilbild von Peter Morris
Peter Morrisvor 3 Monaten

Too slow for me :)

Profilbild von SourceCodeplz
SourceCodeplzvor 3 Monaten

Amazing! I mean it’s running ..

Profilbild von Daksh Trehan
Daksh Trehanvor 3 Monaten

the bit i'd watch: tok/s when a token routes to an expert that isn't in the 75GiB lockable cache yet. MoE + SSD streaming should fly on cache hits and stall on the cold ones. are you getting steady throughput or a sawtooth as the working set shifts?

Profilbild von Human Meteorite
Human Meteoritevor 3 Monaten

@ivanfioravanti You make me want to quit my office job and go all in

Profilbild von Dragos Roua
Dragos Rouavor 3 Monaten

This looks very promising! I assume a new session in a different terminal will slow this down a bit, right? So we’re limited to one flow / machine.

Profilbild von Petit Biscuit
Petit Biscuitvor 3 Monaten

not usable

Profilbild von Vandos ❓
Vandos ❓vor 3 Monaten

The tradeoff here is real, not free: SSD streaming for experts means latency spikes whenever routing hits an expert not already cached. Worth asking what the actual tokens/sec looks like under load, since the headline isn’t ‘it fits,’ it’s ‘how usable is it once it fits.

Profilbild von Donato Molino 🦆
Donato Molino 🦆vor 3 Monaten

What is the problem that causes 3t/s?

Profilbild von Petrus Pennanen
Petrus Pennanenvor 2 Monaten

Wow I had missed this, want to try on my macbook! Where can I find instructions? Thanks :)

Profilbild von John Davis | Solar Strategies
John Davis | Solar Strategiesvor 3 Monaten

Asus GX10?

Profilbild von Andrew Dubin
Andrew Dubinvor 3 Monaten

Has anyone tested this (or DS4 Pro) on M4 Max 128GB? I know the SSD is supposed to be 2x slower, but my tokens were closer to the 0.1-0.5 range...

Profilbild von Eustus Mwaniki
Eustus Mwanikivor 3 Monaten

Tell Dario to come outside.

Profilbild von 云儿淡风也轻
云儿淡风也轻vor 3 Monaten

太棒了,等不及测试了

Profilbild von Ferbin
Ferbinvor 3 Monaten

Does the routing + 2.6-bit quantization actually batch well enough to hide SSD latency, or is it mostly "works on paper" territory?

Ähnliche Videos