Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

319,142 görüntüleme • 3 ay önce •via X (Twitter)

38 Yorum

antirez profil fotoğrafı
antirez3 ay önce

P.s. DwarfStar

antirez profil fotoğrafı
antirez3 ay önce

In full residency in an M3 ultra it goes to 15 t/s. With indexed attention the performance is retained for long contexts work as well, the slope is acceptable. Will post a video of that as well.

Ivan Fioravanti profil fotoğrafı
Ivan Fioravanti3 ay önce

Great job! 🙌🏻 Is DwarfStar ready to run GLM 5.2 on M3 Ultra? Can I benchmark it? Or better wait a more stable release? Thanks 🙏

Matteo Collina profil fotoğrafı
Matteo Collina3 ay önce

wow! Can I test this on the spark?

Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩 profil fotoğrafı
Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩3 ay önce

I am glad I am not the only one torturing LLMs with COBOL: GitHub - ricudis/cobocaml: A bidirectional COBOL <> OCaml FFI bridge · GitHub

CymraegKid profil fotoğrafı
CymraegKid3 ay önce

Cool, but way too slow. Using an agent at this speed sounds painful, honestly.

Ljubomir Josifovski profil fotoğrafı
Ljubomir Josifovski3 ay önce

Nice! Have you seen this? - M3 Ultra (819 GB/s, but only 26 TFLOPS) + DGX Spark (only 273 GB/s, but 100 TFLOPS) + 10 GbE connect.

Timur Yessenov profil fotoğrafı
Timur Yessenov3 ay önce

This is the local-AI point I care about. Not “my laptop beats the cloud.” It doesn’t. But a 128GB Mac running GLM 5.2 with SSD streaming is a real fallback for ugly long tasks when cloud access, price, or rate limits get weird.

ree profil fotoğrafı
ree3 ay önce

I get 2.5t on old 64 x gold xeon CPU only with 544gb. Try out the IK fork of llama-server (not sure how well it does on apple though). M5 Max should get 3x that.

antirez profil fotoğrafı
antirez3 ay önce

It goes 15 t/s with full residency in memory. This is SSD streamed.

ree profil fotoğrafı
ree3 ay önce

Oh, ssd offload with expert streaming? Cool. That sounds like an ssd killer lol. I run mine full context I wonder what full context speed gives you.

François Fleuret profil fotoğrafı
François Fleuret2 ay önce

I do not understand. It's speculative decoding? It cannot stream the full model 20 times per s.

witcheer profil fotoğrafı
witcheer3 ay önce

that is super cool

Rameswar profil fotoğrafı
Rameswar3 ay önce

apple silicon's unified memory makes this kind of extreme offloading more practical than on discrete gpu setups the m5 max 128gb is still expensive, but it's becoming a realistic high end local workstation instead of requiring server class hardware

Sergio Suave profil fotoğrafı
Sergio Suave3 ay önce

Very cool, but painful too watch at that tps...

GeekPark profil fotoğrafı
GeekPark3 ay önce

Zhipu just shipped what the open-source community needed: frontier-class performance that fits in a home office.

Azure | 数据分析万物 profil fotoğrafı
Azure | 数据分析万物3 ay önce

This is a waste of M5 Max...

Jakove_HR profil fotoğrafı
Jakove_HR3 ay önce

@ivanfioravanti The goat

Tiezhen WANG profil fotoğrafı
Tiezhen WANG3 ay önce

Nice work! Are you running MTP?

miko profil fotoğrafı
miko3 ay önce

Intanto grazie per quello che stai facendo. La stessa "velocità" che ho sul mio pc ma con modelli più scemi di me. Perdona se la domanda è idiota come Gemma: è possibile usare dei pesi e switcharli? Es. Per coding, non mi serve la storia d'Italia. Pesi diversi in base all'uso.

Marshall profil fotoğrafı
Marshall3 ay önce

Quantized… that’s going to come with some major drawbacks

Peter Dedene profil fotoğrafı
Peter Dedene3 ay önce

Can we do this on a DGX Spark as well?

mzba profil fotoğrafı
mzba3 ay önce

Was playing around the ssd offload stuff recently and borrowed a lot of cool ideas from ds4 project, it works very well for MOE models..

Charlie Kuo profil fotoğrafı
Charlie Kuo3 ay önce

At that tps it is a stretch calling it “running”…

Peter Morris profil fotoğrafı
Peter Morris3 ay önce

Too slow for me :)

SourceCodeplz profil fotoğrafı
SourceCodeplz3 ay önce

Amazing! I mean it’s running ..

Daksh Trehan profil fotoğrafı
Daksh Trehan3 ay önce

the bit i'd watch: tok/s when a token routes to an expert that isn't in the 75GiB lockable cache yet. MoE + SSD streaming should fly on cache hits and stall on the cold ones. are you getting steady throughput or a sawtooth as the working set shifts?

Human Meteorite profil fotoğrafı
Human Meteorite3 ay önce

@ivanfioravanti You make me want to quit my office job and go all in

Dragos Roua profil fotoğrafı
Dragos Roua3 ay önce

This looks very promising! I assume a new session in a different terminal will slow this down a bit, right? So we’re limited to one flow / machine.

Petit Biscuit profil fotoğrafı
Petit Biscuit3 ay önce

not usable

Vandos ❓ profil fotoğrafı
Vandos ❓3 ay önce

The tradeoff here is real, not free: SSD streaming for experts means latency spikes whenever routing hits an expert not already cached. Worth asking what the actual tokens/sec looks like under load, since the headline isn’t ‘it fits,’ it’s ‘how usable is it once it fits.

Donato Molino 🦆 profil fotoğrafı
Donato Molino 🦆3 ay önce

What is the problem that causes 3t/s?

Petrus Pennanen profil fotoğrafı
Petrus Pennanen2 ay önce

Wow I had missed this, want to try on my macbook! Where can I find instructions? Thanks :)

John Davis | Solar Strategies profil fotoğrafı
John Davis | Solar Strategies3 ay önce

Asus GX10?

Andrew Dubin profil fotoğrafı
Andrew Dubin3 ay önce

Has anyone tested this (or DS4 Pro) on M4 Max 128GB? I know the SSD is supposed to be 2x slower, but my tokens were closer to the 0.1-0.5 range...

Eustus Mwaniki profil fotoğrafı
Eustus Mwaniki3 ay önce

Tell Dario to come outside.

云儿淡风也轻 profil fotoğrafı
云儿淡风也轻3 ay önce

太棒了,等不及测试了

Ferbin profil fotoğrafı
Ferbin3 ay önce

Does the routing + 2.6-bit quantization actually batch well enough to hide SSD latency, or is it mostly "works on paper" territory?

Benzer Videolar