Загрузка видео...

Не удалось загрузить видео

На главную

Check out how DeepSeek v4.1 Flash performs on 2x DGX Sparks after the update. Smooth experience!

39,395 просмотров • 14 дней назад •via X (Twitter)

Комментарии: 36

Фото профиля The Ultimate Accelerationist
The Ultimate Accelerationist14 дней назад

At this stage, I still haven’t found a reason to replace your GLM-5.3-flash-exl3 with deepseek v4.1 flash. I don’t have any frontend requirements on my side.

Фото профиля Mia
Mia14 дней назад

Some prefer DS. Some GLM.

Фото профиля pd19
pd1914 дней назад

我是真喜欢你这个工具,老妹

Фото профиля John Carl - e/acc
John Carl - e/acc14 дней назад

🤯

Фото профиля basedcapital
basedcapital14 дней назад

273 gb/s is bad for dense decode since every token reads all weights. moe reads only active experts, so a big moe split across two sparks is where this box stops looking silly.

Фото профиля Mike Gannotti
Mike Gannotti14 дней назад

Dang!!

Фото профиля Pepo Jiménez
Pepo Jiménez14 дней назад

Which framework do you use for that test? Thanks!

Фото профиля Dan D'Ascenzo
Dan D'Ascenzo14 дней назад

it's indeed almost as fast as TP4!

Фото профиля Javier ⚛ priv/acc
Javier ⚛ priv/acc14 дней назад

this is beautiful 🤩, thanks for making it work on 2 sparks!

Фото профиля recursive_wave
recursive_wave14 дней назад

Here's to models never treating that prompt as a refusal

Фото профиля Kamakura
Kamakura14 дней назад

bruh i need to upgrade my internet to download these models 😭 just downloaded qwen 3.8 lol

Фото профиля VLLM LLVM
VLLM LLVM14 дней назад

Now the question is will it make tp4 even better?

Фото профиля isimsiz
isimsiz14 дней назад

Which herness do you use ?

Фото профиля zhixingheyi
zhixingheyi14 дней назад

Wow 🤩 this is big improvement !!

Фото профиля Francesco Di Costanzo
Francesco Di Costanzo14 дней назад

Wow looks pretty smooth. Now thinking when to move from DS4F to this one 🤔

Фото профиля CocaKova
CocaKova14 дней назад

Downloading the updated version now! I saw a comment on vision. I see this one has it, but he the first iteration said text-language only. First an foremost!! Backing up my models to nextcloud (especially the cyber ones lmao)

Фото профиля 心斋
心斋14 дней назад

How large is the kv cache and ctx window please?

Фото профиля Angel
Angel14 дней назад

keep pushing, soon on 1x single DGX Spark

Фото профиля Mark
Mark14 дней назад

The way that’s my usual prompt for all my tests 😅

Фото профиля Kutluk
Kutluk14 дней назад

i assume it's gonna be not ideal to run under single dgx? since it's much larger than v4

Фото профиля The AI Therapist
The AI Therapist14 дней назад

the real story isn't inference speed. it's the 70% gross margin that makes h100 buyers look like they're buying luxury handbags instead of hardware. buy early or pay rent forever.

Фото профиля Venelin Videnov
Venelin Videnov14 дней назад

Awesome!!!

Фото профиля Atishay Jain
Atishay Jain14 дней назад

interesting, what's the effect of 2x dgx sparks on latency vs throughput for deepseek v4.1 flash? did you explore any optimizations for the dgx-a100's hbm2 memory?

Фото профиля kb.park
kb.park14 дней назад

👍

Фото профиля Julian Raxworthy
Julian Raxworthy14 дней назад

Hey @grok how much is that rig in Aussie dollars?

Фото профиля hhh ppp
hhh ppp14 дней назад

how is the quality?

Фото профиля GitGem.org
GitGem.org14 дней назад

OMG it's truly insane tbh...

Фото профиля tenshin
tenshin14 дней назад

A question are you able to run your prompts + Claude Code But using your local api for Ds v4.1 Flash? Be interesting if you get better outcomes :) I notice in CC DS V4.1 Flash works very well

Фото профиля Tanuj Prakash
Tanuj Prakash14 дней назад

2x setup matching 3-4x old hardware is a nice trick until you realize the next model release makes that 2x setup the new baseline for everyone else.

Фото профиля themiamer🇺🇸
themiamer🇺🇸14 дней назад

What is this harness you are using?

Фото профиля Chris W. Burke
Chris W. Burke14 дней назад

what about TP3?

Фото профиля 111@@@555@@@888
111@@@555@@@88814 дней назад

1M context?

Фото профиля Subhash Dasyam
Subhash Dasyam13 дней назад

how is the accuracy ? what quant ?

Фото профиля draslan.eth
draslan.eth14 дней назад

Faster than GPT 4 Turbo.

Фото профиля Cryptoartmaster
Cryptoartmaster13 дней назад

the smooth experience on 2x DGX shows DeepSeek v4.1 is finally production-ready for scale.

Фото профиля yi
yi14 дней назад

2x hardware is not 2x productivity. The useful test is sustained tool-call throughput as context grows, not a smooth one-shot chat. Publish that trace and the cost claim becomes meaningful.

Похожие видео