Video yükleniyor...
Video Yüklenemedi
I put MCDMA through its paces today and tested Qwen 3.8 Flash Next on my dual Spark / Mac Studio cluster. My first test was disaggregated prefill across my two Sparks, then decode onto my Studio. PP 2,100 tok/s Decode 80 tok/s at 20k context, higher on short replies.... show more
13,121 görüntüleme • 21 gün önce •via X (Twitter)
44 Yorum

Unexpected side effect of splitting inference across the Mac Studio and the DGX Sparks over RDMA: it all runs cooler, by a big margin for the Sparks. Prefill on the Sparks alone pinned them at 80-85 °C. With decode moved to the Studio, 45 min into a task, the Sparks sit at 44 °C and the Studio at 50 °C.

Have you measured this but just with Ethernet? I got pretty good results just with that

I did it with USBC, and it was slower, but not via Ethernet. I will try that, as I do need to see where the benefits are.

Also I tried with Muse. Models have diff kv cache patterns.

What was the outcome?

You can see in the repos I shared. Muser is the sever implementation for muse. 4x ttft, prefil etc. Also if you want secure kv cache you should check our kvpack project.

Bro, we need to connect on TP=4 Sparks to 2x M3 Ultras. I want to see DSv4.1 Flash or GLM 5.3 move with pace in that config.

Currently trying to get DSv4.1 to pipeline across MCDMA into the two sparks and studio. Two studios would be naughty 😈

Question do I need two Mellanox cards/enclosures or can I run TB to each studio from one? I think I am going to need a bigger QSPFP Fabric either way, right now I only have a Microtik 504…

Yeah you will need two enclosures and Mellanox cards one per Mac. You could hang more than one enclosure off a Mac though and that’s what I’m going to test later down the line.

Damn.. good progress.

Honestly did not expect it to be this good.

cuda to metal sits between PP and decode, so ttft is one that tells you if the handoff paid for itself.

Here is my first test, dude!

@volatilemarkts nice one Ash, did you always expect to have to use the OWC and additional hardware or was this a decision that happened along the way? I ask because the set up is somewhat reminiscent of what Alex Ziskind showed once in one of his videos on disagg set ups.

@volatilemarkts I tried wirh usbc but couldn’t get the sparks to enter TB4 mode so this was the next logical step to make rdma possible.

@volatilemarkts have been resisting for the longest time but maybe just have to accept this is what it takes with the hardware we are given. I’ll most def keep an eye on your updates. Super effort mate.

How are you liking this dual machine setup?

I’m genuinely impressed. Only downside I can see is I would normally run qwen3.8 flash next on my studio and glm5.3 on the sparks simultaneously.

Good to know. I've got a 256gb Mac on pre order and have been wondering whether to keep it or instead move towards 4 sparks.

Casually solving the GPU vs RAM speed gap 🤯

AFD community seems hard working now, maybe 5090 prefill for mac studio and run glm-5.3-flash

the Studio alone prefills at about 1,200 tok/s——Can the M3 Ultra really achieve such high prefill throughput? Also, what if two M3 Ultra Mac Studios are connected via RDMA?

Measured, yes. 6B active of 125B = ~14 TFLOP/s of matmul. Works because 3 of 4 layers are DeltaNet with fixed size state, no quadratic attention. Two Studios is wrong for prefill/decode; both are bandwidth-rich and compute-poor. Right for running a model that won't fit on one.

Spark prefill into Studio decode is exactly why people keep unified memory on the desk.

Can you pin/post concurrency runs for prefill speeds . That’s the only main painpoint I foresee on M5U , testing next week though .

This is cool! ConnectX to TB5 cable would be great for this - hope we find a supplier soon

disaggregated prefill on sparks then decode on studio used to be a paper abstract. now it is a tuesday.

I am so so so excited man!! So impressed with momentum and progress. How can I try this with my Macs and Sparks ASAP?

Point your agent at the repo dude!

Just did! I need to get the hardware first right? My agent tells me this: First get Ash’s runnable inference code and confirm whether it can also use ordinary networking.

RDMA uses verbs so you need to ask your agent to set up your inference engine. I’m going to do a full pr on oMLX for this part.

M5 Ultra will erase the prefill advantage of Spark, what's a good combo there?

Maybe but what about all the people who have sparks already and M3U studios, etc? What happens when Nvidia release a new Spark that slaps the M5U's prefill again?

I'm the people 😂 Just that I have 512GB so it's a bit of a mismatch, but 256GB or even 96GB is a bingo.

Do you have any Sparks as well? I have 2 Sparks connected via CX7, feeding my 256GB Studio.

I have 5, but 4 in a cluster, so it's possible to combine? 4x will match or exceed M5U 512GB according to my calculations. I would probably buy now 2x 256GB M5U (more bandwidth/concurrency?)

Yeah, we can combine that. Do you have a MikroTik switch handling all the comms for the Sparks? You will need a TB5 enclosure and a Mellanox card but it should work.

Yes, CRS804. Can you tell which models exactly? (I'm in Europe not sure if there's a lot of choices)

I’m in the UK. I ordered these; Mellanox ConnectX®-5 Ex EN NIC, 100GbE dual-port QSFP28, PCIe MCX516A-CDAT Mellanox Passive Copper Cable 100GbE QSFP28 to QSFP28 1M MCP1600-C001E30N x2 OWC Mercury Helios 5S Thunderbolt 5 (80Gb/s) Single Slot PCIe Card Expansion Solution

Would it work for Macbook M5 Max and Spark?

Yes.

@volatilemarkts Amazing!

Prefill on mac mini m4 Pro is a pain... :D
