Загрузка видео...

Не удалось загрузить видео

На главную

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.

319,142 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 38

Фото профиля antirez
antirez3 месяцев назад

P.s. DwarfStar

Фото профиля antirez
antirez3 месяцев назад

In full residency in an M3 ultra it goes to 15 t/s. With indexed attention the performance is retained for long contexts work as well, the slope is acceptable. Will post a video of that as well.

Фото профиля Ivan Fioravanti
Ivan Fioravanti3 месяцев назад

Great job! 🙌🏻 Is DwarfStar ready to run GLM 5.2 on M3 Ultra? Can I benchmark it? Or better wait a more stable release? Thanks 🙏

Фото профиля Matteo Collina
Matteo Collina3 месяцев назад

wow! Can I test this on the spark?

Фото профиля Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩
Christos Ricudis 🇬🇷🇸🇬🇹🇭🇮🇩3 месяцев назад

I am glad I am not the only one torturing LLMs with COBOL: GitHub - ricudis/cobocaml: A bidirectional COBOL <> OCaml FFI bridge · GitHub

Фото профиля CymraegKid
CymraegKid3 месяцев назад

Cool, but way too slow. Using an agent at this speed sounds painful, honestly.

Фото профиля Ljubomir Josifovski
Ljubomir Josifovski3 месяцев назад

Nice! Have you seen this? - M3 Ultra (819 GB/s, but only 26 TFLOPS) + DGX Spark (only 273 GB/s, but 100 TFLOPS) + 10 GbE connect.

Фото профиля Timur Yessenov
Timur Yessenov3 месяцев назад

This is the local-AI point I care about. Not “my laptop beats the cloud.” It doesn’t. But a 128GB Mac running GLM 5.2 with SSD streaming is a real fallback for ugly long tasks when cloud access, price, or rate limits get weird.

Фото профиля ree
ree3 месяцев назад

I get 2.5t on old 64 x gold xeon CPU only with 544gb. Try out the IK fork of llama-server (not sure how well it does on apple though). M5 Max should get 3x that.

Фото профиля antirez
antirez3 месяцев назад

It goes 15 t/s with full residency in memory. This is SSD streamed.

Фото профиля ree
ree3 месяцев назад

Oh, ssd offload with expert streaming? Cool. That sounds like an ssd killer lol. I run mine full context I wonder what full context speed gives you.

Фото профиля François Fleuret
François Fleuret2 месяцев назад

I do not understand. It's speculative decoding? It cannot stream the full model 20 times per s.

Фото профиля witcheer
witcheer3 месяцев назад

that is super cool

Фото профиля Rameswar
Rameswar3 месяцев назад

apple silicon's unified memory makes this kind of extreme offloading more practical than on discrete gpu setups the m5 max 128gb is still expensive, but it's becoming a realistic high end local workstation instead of requiring server class hardware

Фото профиля Sergio Suave
Sergio Suave3 месяцев назад

Very cool, but painful too watch at that tps...

Фото профиля GeekPark
GeekPark3 месяцев назад

Zhipu just shipped what the open-source community needed: frontier-class performance that fits in a home office.

Фото профиля Azure | 数据分析万物
Azure | 数据分析万物3 месяцев назад

This is a waste of M5 Max...

Фото профиля Jakove_HR
Jakove_HR3 месяцев назад

@ivanfioravanti The goat

Фото профиля Tiezhen WANG
Tiezhen WANG3 месяцев назад

Nice work! Are you running MTP?

Фото профиля miko
miko3 месяцев назад

Intanto grazie per quello che stai facendo. La stessa "velocità" che ho sul mio pc ma con modelli più scemi di me. Perdona se la domanda è idiota come Gemma: è possibile usare dei pesi e switcharli? Es. Per coding, non mi serve la storia d'Italia. Pesi diversi in base all'uso.

Фото профиля Marshall
Marshall3 месяцев назад

Quantized… that’s going to come with some major drawbacks

Фото профиля Peter Dedene
Peter Dedene3 месяцев назад

Can we do this on a DGX Spark as well?

Фото профиля mzba
mzba3 месяцев назад

Was playing around the ssd offload stuff recently and borrowed a lot of cool ideas from ds4 project, it works very well for MOE models..

Фото профиля Charlie Kuo
Charlie Kuo3 месяцев назад

At that tps it is a stretch calling it “running”…

Фото профиля Peter Morris
Peter Morris3 месяцев назад

Too slow for me :)

Фото профиля SourceCodeplz
SourceCodeplz3 месяцев назад

Amazing! I mean it’s running ..

Фото профиля Daksh Trehan
Daksh Trehan3 месяцев назад

the bit i'd watch: tok/s when a token routes to an expert that isn't in the 75GiB lockable cache yet. MoE + SSD streaming should fly on cache hits and stall on the cold ones. are you getting steady throughput or a sawtooth as the working set shifts?

Фото профиля Human Meteorite
Human Meteorite3 месяцев назад

@ivanfioravanti You make me want to quit my office job and go all in

Фото профиля Dragos Roua
Dragos Roua3 месяцев назад

This looks very promising! I assume a new session in a different terminal will slow this down a bit, right? So we’re limited to one flow / machine.

Фото профиля Petit Biscuit
Petit Biscuit3 месяцев назад

not usable

Фото профиля Vandos ❓
Vandos ❓3 месяцев назад

The tradeoff here is real, not free: SSD streaming for experts means latency spikes whenever routing hits an expert not already cached. Worth asking what the actual tokens/sec looks like under load, since the headline isn’t ‘it fits,’ it’s ‘how usable is it once it fits.

Фото профиля Donato Molino 🦆
Donato Molino 🦆3 месяцев назад

What is the problem that causes 3t/s?

Фото профиля Petrus Pennanen
Petrus Pennanen2 месяцев назад

Wow I had missed this, want to try on my macbook! Where can I find instructions? Thanks :)

Фото профиля John Davis | Solar Strategies
John Davis | Solar Strategies3 месяцев назад

Asus GX10?

Фото профиля Andrew Dubin
Andrew Dubin3 месяцев назад

Has anyone tested this (or DS4 Pro) on M4 Max 128GB? I know the SSD is supposed to be 2x slower, but my tokens were closer to the 0.1-0.5 range...

Фото профиля Eustus Mwaniki
Eustus Mwaniki3 месяцев назад

Tell Dario to come outside.

Фото профиля 云儿淡风也轻
云儿淡风也轻3 месяцев назад

太棒了,等不及测试了

Фото профиля Ferbin
Ferbin3 месяцев назад

Does the routing + 2.6-bit quantization actually batch well enough to hide SSD latency, or is it mostly "works on paper" territory?

Похожие видео