Загрузка видео...
Не удалось загрузить видео
Check out how DeepSeek v4.1 Flash performs on 2x DGX Sparks after the update. Smooth experience!
39,395 просмотров • 14 дней назад •via X (Twitter)
Комментарии: 36

At this stage, I still haven’t found a reason to replace your GLM-5.3-flash-exl3 with deepseek v4.1 flash. I don’t have any frontend requirements on my side.

Some prefer DS. Some GLM.

我是真喜欢你这个工具,老妹

🤯

273 gb/s is bad for dense decode since every token reads all weights. moe reads only active experts, so a big moe split across two sparks is where this box stops looking silly.

Dang!!

Which framework do you use for that test? Thanks!

it's indeed almost as fast as TP4!

this is beautiful 🤩, thanks for making it work on 2 sparks!

Here's to models never treating that prompt as a refusal

bruh i need to upgrade my internet to download these models 😭 just downloaded qwen 3.8 lol

Now the question is will it make tp4 even better?

Which herness do you use ?

Wow 🤩 this is big improvement !!

Wow looks pretty smooth. Now thinking when to move from DS4F to this one 🤔

Downloading the updated version now! I saw a comment on vision. I see this one has it, but he the first iteration said text-language only. First an foremost!! Backing up my models to nextcloud (especially the cyber ones lmao)

How large is the kv cache and ctx window please?

keep pushing, soon on 1x single DGX Spark

The way that’s my usual prompt for all my tests 😅

i assume it's gonna be not ideal to run under single dgx? since it's much larger than v4

the real story isn't inference speed. it's the 70% gross margin that makes h100 buyers look like they're buying luxury handbags instead of hardware. buy early or pay rent forever.

Awesome!!!

interesting, what's the effect of 2x dgx sparks on latency vs throughput for deepseek v4.1 flash? did you explore any optimizations for the dgx-a100's hbm2 memory?

👍

Hey @grok how much is that rig in Aussie dollars?

how is the quality?

OMG it's truly insane tbh...

A question are you able to run your prompts + Claude Code But using your local api for Ds v4.1 Flash? Be interesting if you get better outcomes :) I notice in CC DS V4.1 Flash works very well

2x setup matching 3-4x old hardware is a nice trick until you realize the next model release makes that 2x setup the new baseline for everyone else.

What is this harness you are using?

what about TP3?

1M context?

how is the accuracy ? what quant ?

Faster than GPT 4 Turbo.

the smooth experience on 2x DGX shows DeepSeek v4.1 is finally production-ready for scale.

2x hardware is not 2x productivity. The useful test is sustained tool-call throughput as context grows, not a smooth one-shot chat. Publish that trace and the cost claim becomes meaningful.
