Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

How DeepSeek v4.1 Flash looks on 3x DGX Sparks with 4 concurrent streams

19,730 görüntüleme • 3 gün önce •via X (Twitter)

20 Yorum

Mia profil fotoğrafı
Mia3 gün önce

You can run it too

Frank profil fotoğrafı
Frank3 gün önce

Damnit damnit damnit. Yesterday I had 3 sparks and now I have only 2…

Thomas profil fotoğrafı
Thomas3 gün önce

Thanks for sharing! How does it fit your different use cases compared with GLM-5.3-Flash and DeepSeek-V4-Flash? I haven’t seen many real-world comparisons yet. I’m not sure it’s worth adding a third Spark when GLM fits on two and has somewhat acceptable task-completion time

~/.shane profil fotoğrafı
~/.shane3 gün önce

What a great model

Ibesh profil fotoğrafı
Ibesh3 gün önce

running flash locally on 3 sparks with 4 streams like that is wild. cheap way to get a lot of parallel work done without renting a huge box, really shows local setups catching up

R N .. profil fotoğrafı
R N ..3 gün önce

looks promising, you were the reason for my second spark and now third.

Norsyx profil fotoğrafı
Norsyx3 gün önce

Amazing!

pd19 profil fotoğrafı
pd193 gün önce

你这个工具是啥,跟claude code很像

Mia profil fotoğrafı
Mia3 gün önce

Thanks. My own harness, unreleased yet

pd19 profil fotoğrafı
pd193 gün önce

真好,老妹儿,你心灵手巧

HealthRanger profil fotoğrafı
HealthRanger3 gün önce

Pretty cool. Love how fast you and other AI labs are crunching this model onto desktops. Next stop? TWO DGX Sparks!

Mia profil fotoğrafı
Mia3 gün önce

2 dgx Spark recipe coming soon

error501 profil fotoğrafı
error5013 gün önce

How did you solve ttft?

JCode profil fotoğrafı
JCode3 gün önce

Thinking off? Looks amazing btw!

Mia profil fotoğrafı
Mia3 gün önce

Actually didn't pay attention to Thinking, it's on auto, I think it's off by auto. Either way even with thinking it works really well!

JCode profil fotoğrafı
JCode3 gün önce

Thats awesome. Looks great. I am still struggling with concurrencies in GLM. I think is due to prefill, that is ñimited to one at a time.

Context Studios - AI Development Studio Berlin profil fotoğrafı
Context Studios - AI Development Studio Berlin3 gün önce

Tensor parallel across three Sparks is the neat part. The 890 bytes KV per token figure is exactly what makes a 1M-context model fit on desk hardware now.

Koldfrontier profil fotoğrafı
Koldfrontier3 gün önce

Actual Hermes performance on C6 agentic dev load. (using patched NCCL 4x Sparks)

zompdesigns profil fotoğrafı
zompdesigns3 gün önce

Insanity

Benjamin profil fotoğrafı
Benjamin3 gün önce

Do I go into debt to buy dgx I so what this !

Benzer Videolar