Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Check out how DeepSeek v4.1 Flash performs on 2x DGX Sparks after the update. Smooth experience!

39,395 Aufrufe • vor 14 Tagen •via X (Twitter)

36 Kommentare

Profilbild von The Ultimate Accelerationist
The Ultimate Accelerationistvor 14 Tagen

At this stage, I still haven’t found a reason to replace your GLM-5.3-flash-exl3 with deepseek v4.1 flash. I don’t have any frontend requirements on my side.

Profilbild von Mia
Miavor 14 Tagen

Some prefer DS. Some GLM.

Profilbild von pd19
pd19vor 14 Tagen

我是真喜欢你这个工具,老妹

Profilbild von John Carl - e/acc
John Carl - e/accvor 14 Tagen

🤯

Profilbild von basedcapital
basedcapitalvor 14 Tagen

273 gb/s is bad for dense decode since every token reads all weights. moe reads only active experts, so a big moe split across two sparks is where this box stops looking silly.

Profilbild von Mike Gannotti
Mike Gannottivor 14 Tagen

Dang!!

Profilbild von Pepo Jiménez
Pepo Jiménezvor 14 Tagen

Which framework do you use for that test? Thanks!

Profilbild von Dan D'Ascenzo
Dan D'Ascenzovor 14 Tagen

it's indeed almost as fast as TP4!

Profilbild von Javier ⚛ priv/acc
Javier ⚛ priv/accvor 14 Tagen

this is beautiful 🤩, thanks for making it work on 2 sparks!

Profilbild von recursive_wave
recursive_wavevor 14 Tagen

Here's to models never treating that prompt as a refusal

Profilbild von Kamakura
Kamakuravor 14 Tagen

bruh i need to upgrade my internet to download these models 😭 just downloaded qwen 3.8 lol

Profilbild von VLLM LLVM
VLLM LLVMvor 14 Tagen

Now the question is will it make tp4 even better?

Profilbild von isimsiz
isimsizvor 14 Tagen

Which herness do you use ?

Profilbild von zhixingheyi
zhixingheyivor 14 Tagen

Wow 🤩 this is big improvement !!

Profilbild von Francesco Di Costanzo
Francesco Di Costanzovor 14 Tagen

Wow looks pretty smooth. Now thinking when to move from DS4F to this one 🤔

Profilbild von CocaKova
CocaKovavor 13 Tagen

Downloading the updated version now! I saw a comment on vision. I see this one has it, but he the first iteration said text-language only. First an foremost!! Backing up my models to nextcloud (especially the cyber ones lmao)

Profilbild von 心斋
心斋vor 14 Tagen

How large is the kv cache and ctx window please?

Profilbild von Angel
Angelvor 14 Tagen

keep pushing, soon on 1x single DGX Spark

Profilbild von Mark
Markvor 14 Tagen

The way that’s my usual prompt for all my tests 😅

Profilbild von Kutluk
Kutlukvor 14 Tagen

i assume it's gonna be not ideal to run under single dgx? since it's much larger than v4

Profilbild von The AI Therapist
The AI Therapistvor 14 Tagen

the real story isn't inference speed. it's the 70% gross margin that makes h100 buyers look like they're buying luxury handbags instead of hardware. buy early or pay rent forever.

Profilbild von Venelin Videnov
Venelin Videnovvor 14 Tagen

Awesome!!!

Profilbild von Atishay Jain
Atishay Jainvor 14 Tagen

interesting, what's the effect of 2x dgx sparks on latency vs throughput for deepseek v4.1 flash? did you explore any optimizations for the dgx-a100's hbm2 memory?

Profilbild von kb.park
kb.parkvor 14 Tagen

👍

Profilbild von Julian Raxworthy
Julian Raxworthyvor 14 Tagen

Hey @grok how much is that rig in Aussie dollars?

Profilbild von hhh ppp
hhh pppvor 14 Tagen

how is the quality?

Profilbild von GitGem.org
GitGem.orgvor 14 Tagen

OMG it's truly insane tbh...

Profilbild von tenshin
tenshinvor 14 Tagen

A question are you able to run your prompts + Claude Code But using your local api for Ds v4.1 Flash? Be interesting if you get better outcomes :) I notice in CC DS V4.1 Flash works very well

Profilbild von Tanuj Prakash
Tanuj Prakashvor 13 Tagen

2x setup matching 3-4x old hardware is a nice trick until you realize the next model release makes that 2x setup the new baseline for everyone else.

Profilbild von themiamer🇺🇸
themiamer🇺🇸vor 14 Tagen

What is this harness you are using?

Profilbild von Chris W. Burke
Chris W. Burkevor 14 Tagen

what about TP3?

Profilbild von 111@@@555@@@888
111@@@555@@@888vor 14 Tagen

1M context?

Profilbild von Subhash Dasyam
Subhash Dasyamvor 13 Tagen

how is the accuracy ? what quant ?

Profilbild von draslan.eth
draslan.ethvor 14 Tagen

Faster than GPT 4 Turbo.

Profilbild von Cryptoartmaster
Cryptoartmastervor 13 Tagen

the smooth experience on 2x DGX shows DeepSeek v4.1 is finally production-ready for scale.

Profilbild von yi
yivor 14 Tagen

2x hardware is not 2x productivity. The useful test is sustained tool-call throughput as context grows, not a smooth one-shot chat. Publish that trace and the cost claim becomes meaningful.

Ähnliche Videos