Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Check out how DeepSeek v4.1 Flash performs on 2x DGX Sparks after the update. Smooth experience!

39,395 görüntüleme • 14 gün önce •via X (Twitter)

36 Yorum

The Ultimate Accelerationist profil fotoğrafı
The Ultimate Accelerationist14 gün önce

At this stage, I still haven’t found a reason to replace your GLM-5.3-flash-exl3 with deepseek v4.1 flash. I don’t have any frontend requirements on my side.

Mia profil fotoğrafı
Mia14 gün önce

Some prefer DS. Some GLM.

pd19 profil fotoğrafı
pd1914 gün önce

我是真喜欢你这个工具,老妹

John Carl - e/acc profil fotoğrafı
John Carl - e/acc14 gün önce

🤯

basedcapital profil fotoğrafı
basedcapital14 gün önce

273 gb/s is bad for dense decode since every token reads all weights. moe reads only active experts, so a big moe split across two sparks is where this box stops looking silly.

Mike Gannotti profil fotoğrafı
Mike Gannotti14 gün önce

Dang!!

Pepo Jiménez profil fotoğrafı
Pepo Jiménez14 gün önce

Which framework do you use for that test? Thanks!

Dan D'Ascenzo profil fotoğrafı
Dan D'Ascenzo14 gün önce

it's indeed almost as fast as TP4!

Javier ⚛ priv/acc profil fotoğrafı
Javier ⚛ priv/acc14 gün önce

this is beautiful 🤩, thanks for making it work on 2 sparks!

recursive_wave profil fotoğrafı
recursive_wave14 gün önce

Here's to models never treating that prompt as a refusal

Kamakura profil fotoğrafı
Kamakura14 gün önce

bruh i need to upgrade my internet to download these models 😭 just downloaded qwen 3.8 lol

VLLM LLVM profil fotoğrafı
VLLM LLVM14 gün önce

Now the question is will it make tp4 even better?

isimsiz profil fotoğrafı
isimsiz14 gün önce

Which herness do you use ?

zhixingheyi profil fotoğrafı
zhixingheyi14 gün önce

Wow 🤩 this is big improvement !!

Francesco Di Costanzo profil fotoğrafı
Francesco Di Costanzo14 gün önce

Wow looks pretty smooth. Now thinking when to move from DS4F to this one 🤔

CocaKova profil fotoğrafı
CocaKova14 gün önce

Downloading the updated version now! I saw a comment on vision. I see this one has it, but he the first iteration said text-language only. First an foremost!! Backing up my models to nextcloud (especially the cyber ones lmao)

心斋 profil fotoğrafı
心斋14 gün önce

How large is the kv cache and ctx window please?

Angel profil fotoğrafı
Angel14 gün önce

keep pushing, soon on 1x single DGX Spark

Mark profil fotoğrafı
Mark14 gün önce

The way that’s my usual prompt for all my tests 😅

Kutluk profil fotoğrafı
Kutluk14 gün önce

i assume it's gonna be not ideal to run under single dgx? since it's much larger than v4

The AI Therapist profil fotoğrafı
The AI Therapist14 gün önce

the real story isn't inference speed. it's the 70% gross margin that makes h100 buyers look like they're buying luxury handbags instead of hardware. buy early or pay rent forever.

Venelin Videnov profil fotoğrafı
Venelin Videnov14 gün önce

Awesome!!!

Atishay Jain profil fotoğrafı
Atishay Jain14 gün önce

interesting, what's the effect of 2x dgx sparks on latency vs throughput for deepseek v4.1 flash? did you explore any optimizations for the dgx-a100's hbm2 memory?

kb.park profil fotoğrafı
kb.park14 gün önce

👍

Julian Raxworthy profil fotoğrafı
Julian Raxworthy14 gün önce

Hey @grok how much is that rig in Aussie dollars?

hhh ppp profil fotoğrafı
hhh ppp14 gün önce

how is the quality?

GitGem.org profil fotoğrafı
GitGem.org14 gün önce

OMG it's truly insane tbh...

tenshin profil fotoğrafı
tenshin14 gün önce

A question are you able to run your prompts + Claude Code But using your local api for Ds v4.1 Flash? Be interesting if you get better outcomes :) I notice in CC DS V4.1 Flash works very well

Tanuj Prakash profil fotoğrafı
Tanuj Prakash14 gün önce

2x setup matching 3-4x old hardware is a nice trick until you realize the next model release makes that 2x setup the new baseline for everyone else.

themiamer🇺🇸 profil fotoğrafı
themiamer🇺🇸14 gün önce

What is this harness you are using?

Chris W. Burke profil fotoğrafı
Chris W. Burke14 gün önce

what about TP3?

111@@@555@@@888 profil fotoğrafı
111@@@555@@@88814 gün önce

1M context?

Subhash Dasyam profil fotoğrafı
Subhash Dasyam13 gün önce

how is the accuracy ? what quant ?

draslan.eth profil fotoğrafı
draslan.eth14 gün önce

Faster than GPT 4 Turbo.

Cryptoartmaster profil fotoğrafı
Cryptoartmaster13 gün önce

the smooth experience on 2x DGX shows DeepSeek v4.1 is finally production-ready for scale.

yi profil fotoğrafı
yi14 gün önce

2x hardware is not 2x productivity. The useful test is sustained tool-call throughput as context grows, not a smooth one-shot chat. Publish that trace and the cost claim becomes meaningful.

Benzer Videolar