Loading video...
Video Failed to Load
DeepSeek v4.1 Flash running on 2x DGX Sparks
57,434 views • 3 days ago •via X (Twitter)
46 Comments

looks ~25 tps, nice!

Amazing! API access out just chat?

This is a demo of chatting with it to test prose performance real time. Don't expect the same performance and kv cache pool as v4

What interface is this?

it's my harness, unreleased yet

🙌🏿🥚🔥

Do you think DeepSeek is worth looking into compared to Qwen 3.8 Flash Next at the moment? I’m asking specifically in terms of running it on 2x DGX Spark.

Nice! getting close to release? Which quant is it?

Yes sooner than later My own harness unreleased yet

Nice!!

Great! Two sparks, not four. That’s the setup most people can actually buy.

Do you sleep?

sometimes

Looks amazing!! How does it compare in overall intelligence to glm flash?

GLM 5.3 Flash will perform better in terms speed and prefill, as the overall size of the quant is noticeably smaller. Intelligence I think they are comparable, still testing

That's great, thanks Mia

does it run well on two strix halo 128? i plan on getting a second one

This will only work on Nvidia

is there anything you would recommend please, is it worth getting a secondary amd strix halo? and which model do you recommend? AFAIK qwen 3.8 flash next is the best option for me right now until i have two strix halo machines

Game on!

... wtf Mia's made her own whole harness just to test these LLMs :D

will max studio M5 Ultra 256GB produce same or better result ?

2× Spark for DeepSeek Flash is a tidy local footprint. That's the case where prompts stay in the building. Cloud still covers everything else.

wow that is good speed, glm flash level. Please also do some benchmarks for quality so we know what to expect when compared to full model

Which harness are you using @MiaAI_lab ?

it's my harness, unreleased yet

I'm rooting for it to be a success then, can't wait to try it 😉

就知道你一定可以。

What kind of tokens per second are you getting out of that setup?

@MystiqueMide what do you think?

What kind of token speed are you getting

@voluntas 这么快?

2x DGX Sparks is a bandwidth play, not FLOPs. For v4.1 Flash, MoE routing + KV cache pressure will decide if it feels snappy or just technically on-device.

Ok make it usable mcp and take my credit card along with it, deal ?

how about the speed? Comparing GML 5.3 flash

a week ago the community said this one was too big for two of these machines. now it's running on two. the "too big" line keeps moving

Decent speed enough for anyone

@voluntas 还是我老妹儿的工具好看

Not everyone can afford unfortunately

Can you give some more details about recipe? What quant and how did you fit it?

👀 on a saturday too 🔥

Holy shit

What web UI is that?

do you like it?

Yes!

it's my harness, unreleased yet
Related Videos
How DeepSeek v4.1 Flash looks on 3x DGX Sparks with 4 concurrent streams
Mia
19,730 views • 4 days ago
