Video yükleniyor...
Video Yüklenemedi
How DeepSeek v4.1 Flash looks on 3x DGX Sparks with 4 concurrent streams
19,730 görüntüleme • 3 gün önce •via X (Twitter)
20 Yorum

You can run it too

Damnit damnit damnit. Yesterday I had 3 sparks and now I have only 2…

Thanks for sharing! How does it fit your different use cases compared with GLM-5.3-Flash and DeepSeek-V4-Flash? I haven’t seen many real-world comparisons yet. I’m not sure it’s worth adding a third Spark when GLM fits on two and has somewhat acceptable task-completion time

What a great model

running flash locally on 3 sparks with 4 streams like that is wild. cheap way to get a lot of parallel work done without renting a huge box, really shows local setups catching up

looks promising, you were the reason for my second spark and now third.

Amazing!

你这个工具是啥,跟claude code很像

Thanks. My own harness, unreleased yet

真好,老妹儿,你心灵手巧

Pretty cool. Love how fast you and other AI labs are crunching this model onto desktops. Next stop? TWO DGX Sparks!

2 dgx Spark recipe coming soon

How did you solve ttft?

Thinking off? Looks amazing btw!

Actually didn't pay attention to Thinking, it's on auto, I think it's off by auto. Either way even with thinking it works really well!

Thats awesome. Looks great. I am still struggling with concurrencies in GLM. I think is due to prefill, that is ñimited to one at a time.

Tensor parallel across three Sparks is the neat part. The 890 bytes KV per token figure is exactly what makes a 1M-context model fit on desk hardware now.

Actual Hermes performance on C6 agentic dev load. (using patched NCCL 4x Sparks)

Insanity

Do I go into debt to buy dgx I so what this !
Benzer Videolar
DEEPSEEK v4.1 Flash is SO FAST What are they feeding this thing?
Tak 🦞
67,635 görüntüleme • 4 gün önce
