Video yükleniyor...
Video Yüklenemedi
The new DeepSeek V4.1 Flash model is mindblowing - back on top of the open-source model leaderboard and extremely cheap. It has a lot of very smart ways to be efficient and highly capable so I made a video of the forward pass to give you a view of... show more
58,020 görüntüleme • 1 gün önce •via X (Twitter)
17 Yorum

Beautiful pass. Still a run. Still hours.

It's almost impossible to keep up with the local AI world nowadays. It's moving at an insane speed.

I'm watching KV-cache pressure more than raw tokens per second. Better cache reuse is what decides whether parallel coding sessions feel cheap in practice.

open weights the same day as the drop is the part labs hate.

An MLX version would be nice.

The forward-pass view is useful, but the production constraint is memory traffic, not only FLOPs. I’d want the same breakdown with KV-cache size and tokens/sec at long context; that’s where cheap inference gets less cheap.

The scaffold table is the interesting part: DeepSWE v1.1 runs 65.5 to 74.2 across harnesses and the headline 74.2 is mini-SWE, while Terminal-Bench 2.1 runs 84.1 to 90.6 and the headline 90.6 is DeepSeek's own harness.

this should be the default model release artifact. benchmarks tell you where it won. an animated forward pass tells builders where the cost went.

Do we know if this has anything approaching an internal world model? I believe Astra is moving this way.

The interesting part isn’t just that DeepSeek is back on top. It’s how they’re getting this much capability at this price.

前向传播可视化比单看榜单直观多了。

DeepSeek keeps pushing the boundaries of efficiency. ���� Cheap + highly capable open-source models are a huge win for developers. 🤖🚀

We've seen this pattern before, but closer to home. Thus, whether east or west, the play--whatever it is--I strongly suspect is not in our favor. That is wisdom. If you can accept it, stop using cloud models.

So fast and efficient!

Let's connect Thomas😊

Open weights plus this forward-pass view make the efficiency claims much easier to inspect

开源模型更新,现在连 forward pass 都做成短视频了。可读性本身成了发布资产:别人看懂你怎么省算力,比再刷一分榜更容易被采用。

