Loading video...

Video Failed to Load

Go Home

DeepSeek v4.1 Flash via API is almost TOO fast 🤯

137,788 views • 6 days ago •via X (Twitter)

46 Comments

keys 🧪's profile picture
keys 🧪6 days ago

Actually don’t want the model to run too fast then you have keep scrolling up to see what happened it’s annoying after a while

Mia's profile picture
Mia6 days ago

Good problem to have imo

hkdom's profile picture
hkdom6 days ago

You will wow more if you run it with DSH

Mia's profile picture
Mia6 days ago

You're probably right

Santanu Sinha's profile picture
Santanu Sinha6 days ago

Genuine question.. what harness is that and why is it flickering so much? Also .. yes amazing speed

Mia's profile picture
Mia6 days ago

Hermes agent CLI

Santanu Sinha's profile picture
Santanu Sinha6 days ago

Ok thanks

kaith's profile picture
kaith6 days ago

How’s it’s performance?

Mia's profile picture
Mia6 days ago

Looking better than v4 flash

Consigliere's profile picture
Consigliere6 days ago

via Cloud API?

Mia's profile picture
Mia6 days ago

Yes

Gaurav Bhatia's profile picture
Gaurav Bhatia6 days ago

Imagine after thinking for that long it doesn’t complete the work.

Mia's profile picture
Mia6 days ago

But it did 😀

Wassollichhier's profile picture
Wassollichhier6 days ago

is vision onboard?

Mia's profile picture
Mia6 days ago

Not on this one, it's beta

Samuel McHargue's profile picture
Samuel McHargue6 days ago

@grok how much is v4.1 flash compared to the older one.

Bear's profile picture
Bear6 days ago

Vision support?

Daniel Carneiro's profile picture
Daniel Carneiro6 days ago

WTF is that ugly TUI?

nah's profile picture
nah6 days ago

Imagine every model being this fast. That'd be amazing.

Steven Cheng's profile picture
Steven Cheng6 days ago

Latency that low feels like cheating.

base's profile picture
base6 days ago

get in my spark … lol

CV.YH's profile picture
CV.YH6 days ago

Wow

Chris Winslow's profile picture
Chris Winslow6 days ago

Speed ≠ Quality?

Saurav's profile picture
Saurav6 days ago

actually useful?

Mia's profile picture
Mia6 days ago

Seems like it's better than v4 flash

bitflipgremlin's profile picture
bitflipgremlin6 days ago

this isn't quite mercury-2 speed but it's damn fucking close, and quality leads me to suspect it's a different mechanism. i wonder what kind of bullshit the whale came up with now

Ian Hailey's profile picture
Ian Hailey6 days ago

It can never be too fast!

dan0mad's profile picture
dan0mad6 days ago

There’s 4.1???!? Since wen?

forreal's profile picture
forreal6 days ago

better than glm 5.3 flash? it only matters if it reaches performance of kimi k3 or glm 5.3.

Devin Oldenburg's profile picture
Devin Oldenburg6 days ago

~176 tokens/sec

Redmix's profile picture
Redmix6 days ago

It's a shame we can't enjoy it for very long.

Mia's profile picture
Mia6 days ago

Squeezing what I can !

Redmix's profile picture
Redmix6 days ago

✌🏻

Cato Nooka's profile picture
Cato Nooka6 days ago

if there is another model break through for this kind of speed, I don't know what to say, not only so fast at 400 TPS, but also more token efficient, the time for each task is shink down to minutes

Yume_X's profile picture
Yume_X6 days ago

Yeah I was shocked too when I saw it , makes a good argument for 300 tok/s builds when you see it lol

CryptoYeti's profile picture
CryptoYeti6 days ago

@grok wie teuer ist die API von der neuste Model von Deepseek? Und vergleiche sie mit Muse spark corbitor .

Jayden's profile picture
Jayden6 days ago

The latency is fun. The bill is the actual benchmark.

Mike's profile picture
Mike6 days ago

I really wish Nous would fix the CLI redraw flashing in Hermes. That’s mainly why I stick to the TUI.

remakw's profile picture
remakw6 days ago

fastest model i gave ever used switched the trace off because my eyes were hurting lol

rapidTools / Mike's profile picture
rapidTools / Mike6 days ago

o_0... Dafuq... I hope we get this open as well. It would speed up local flash model az well I guess. (not to this speed obviously, but making it even faster wpuld be cool)

Tom's profile picture
Tom6 days ago

Why I only see v40730? Where is 4.1?

Steven Cheng's profile picture
Steven Cheng6 days ago

Almost" is doing a lot of heavy lifting there. Latency spikes still kill the vibe when you are chaining calls.

Reelix's profile picture
Reelix6 days ago

Do a speed comparison :p

Cave's profile picture
Cave6 days ago

Woo.

Jarno's profile picture
Jarno6 days ago

Fast until you hit the rate limit, then it's back to staring at a spinner like every other model.

Akshay Joshi's profile picture
Akshay Joshi6 days ago

Thought it’s free and went rampant and burned 96million token from Deekseek platform hopefully 97% cache hit hence 1.30$ only

Related Videos