Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

The new DeepSeek V4.1 Flash model is mindblowing - back on top of the open-source model leaderboard and extremely cheap. It has a lot of very smart ways to be efficient and highly capable so I made a video of the forward pass to give you a view of...

58,020 görüntüleme • 1 gün önce •via X (Twitter)

17 Yorum

AionCloud profil fotoğrafı
AionCloud1 gün önce

Beautiful pass. Still a run. Still hours.

Neil profil fotoğrafı
Neil1 gün önce

It's almost impossible to keep up with the local AI world nowadays. It's moving at an insane speed.

Mehdi Zare profil fotoğrafı
Mehdi Zare1 gün önce

I'm watching KV-cache pressure more than raw tokens per second. Better cache reuse is what decides whether parallel coding sessions feel cheap in practice.

Shesaidmewakeup profil fotoğrafı
Shesaidmewakeup1 gün önce

open weights the same day as the drop is the part labs hate.

Information Machine profil fotoğrafı
Information Machine1 gün önce

An MLX version would be nice.

Nick profil fotoğrafı
Nick1 gün önce

The forward-pass view is useful, but the production constraint is memory traffic, not only FLOPs. I’d want the same breakdown with KV-cache size and tokens/sec at long context; that’s where cheap inference gets less cheap.

Marius Laurusevicius profil fotoğrafı
Marius Laurusevicius1 gün önce

The scaffold table is the interesting part: DeepSWE v1.1 runs 65.5 to 74.2 across harnesses and the headline 74.2 is mini-SWE, while Terminal-Bench 2.1 runs 84.1 to 90.6 and the headline 90.6 is DeepSeek's own harness.

John Rood profil fotoğrafı
John Rood1 gün önce

this should be the default model release artifact. benchmarks tell you where it won. an animated forward pass tells builders where the cost went.

Gentleman Tech Bro profil fotoğrafı
Gentleman Tech Bro1 gün önce

Do we know if this has anything approaching an internal world model? I believe Astra is moving this way.

Naveen Saradhi profil fotoğrafı
Naveen Saradhi1 gün önce

The interesting part isn’t just that DeepSeek is back on top. It’s how they’re getting this much capability at this price.

安叫兽|Bird🕊️ 🔶 BNB profil fotoğrafı
安叫兽|Bird🕊️ 🔶 BNB1 gün önce

前向传播可视化比单看榜单直观多了。

MOHAMMED ISMAIL profil fotoğrafı
MOHAMMED ISMAIL1 gün önce

DeepSeek keeps pushing the boundaries of efficiency. ���� Cheap + highly capable open-source models are a huge win for developers. 🤖🚀

David Moore profil fotoğrafı
David Moore1 gün önce

We've seen this pattern before, but closer to home. Thus, whether east or west, the play--whatever it is--I strongly suspect is not in our favor. That is wisdom. If you can accept it, stop using cloud models.

Tanjiro profil fotoğrafı
Tanjiro1 gün önce

So fast and efficient!

Srinivas Avula✨ profil fotoğrafı
Srinivas Avula✨1 gün önce

Let's connect Thomas😊

Xman profil fotoğrafı
Xman1 gün önce

Open weights plus this forward-pass view make the efficiency claims much easier to inspect

VastPlan profil fotoğrafı
VastPlan1 gün önce

开源模型更新,现在连 forward pass 都做成短视频了。可读性本身成了发布资产:别人看懂你怎么省算力,比再刷一分榜更容易被采用。

Benzer Videolar

Dario Amodei just dismantled the biggest myth in the AI industry. Open source AI isn’t free. It never was. Amodei: “It’s not free. You have to run it on inference and someone has to make it fast on inference.” For decades, open source meant something real. It meant a teenager in a basement could download the same tools as a Fortune 500 company. Could read the code. Could modify it. Could build something that competed with the giants. That was genuine democratization. That actually happened. AI is different. Fundamentally. Physically. In ways the ideology hasn’t caught up to yet. Downloading the weights is the easy part. The part that actually costs something is turning the weights into a running system. Into responses. Into intelligence operating in real time at scale. That requires compute. Power. Infrastructure. The kind measured in billions of dollars and years of construction. Amodei: “These are big models. They’re hard to do inference on. Ultimately you have to host it on the cloud. The people who host it on the cloud do inference.” The open source debate was never about who owns the model. It was always about who owns the cloud. And Amodei goes further. When a competitor drops a new open model, he doesn’t ask whether it’s open or closed. He doesn’t care about the licensing. He doesn’t engage the ideology. Amodei: “I don’t think it mattered that DeepSeek is open source. I think I ask, is it a good model? Is it better than us at the things that matter? That’s the only thing that I care about.” That’s the ruthless clarity of someone actually trying to win. While the media debates licensing frameworks, Amodei is asking one question. Is it better. Everything else is a distraction. Amodei: “I don’t think open source works the same way in AI that it has worked in other areas. Here we can’t see inside the model.” This isn’t Linux. You can’t read it. You can’t fork it. You can’t understand it the way generations of developers understood the tools they inherited. You can download it. And then you need a data center to run it. The teenager in the basement who was supposed to be empowered by this revolution needs a billion dollars of infrastructure before the empowerment starts. The era of the basement coder rewriting civilization on a laptop is over. The future belongs to whoever commands the compute, owns the power grid, and can actually turn the intelligence on. Open weights without infrastructure isn’t democratization. It’s a promise the physics of the universe won’t let us keep.

Dustin

687,842 görüntüleme • 6 ay önce