Loading video...

Video Failed to Load

Go Home

.DeepSeek-V4-Pro-0813 is live on DigitalOcean Serverless Inference and Inference Router. 🚀 1.6T params / 49B active, 1M-token context, AA Intelligence Index of 53 at max reasoning effort.

347,944 views • 27 days ago •via X (Twitter)

5 Comments

Bob Tong's profile picture
Bob Tong25 days ago

@deepseek_ai index 53 vs the open-weight median of 27 is a genuinely wide gap, not marginal. worth checking if that holds at lower reasoning effort though, AA's number is explicitly "max reasoning effort" which usually means slower/pricier in practice than the headline suggests

RAZA | AI EXPLORER's profile picture
RAZA | AI EXPLORER25 days ago

@deepseek_ai DeepSeek V4 Pro is live on DigitalOcean. 1.6T params, 49B active, and 1M-token context.

Rao Umair's profile picture
Rao Umair27 days ago

@deepseek_ai 1.6T parameters with 49B active and a 1M-token context window is a massive setup. DeepSeek V4 Pro looks built for serious workloads.

WhiteCap Data's profile picture
WhiteCap Data26 days ago

@deepseek_ai Serverless inference cant reach the data it needs to reason about if your POS, CRM, and email systems stay isolated.

MoExperts's profile picture
MoExperts27 days ago

@deepseek_ai curious how this performs on real workloads vs the benchmark claims

Related Videos

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time I feel like there is an open-weight model that can reason at the level of Claude and Codex. And it does this in a cost-effective way with support for 1M context length. To be clear, I am using DeepSeek-V4-Pro inside of Pi without any special configuration. It works out of the box. It's exciting that there is a model that can just be plugged into a basic harness like Pi, and it just works. I've never seen that before. Most models require lots of configuration and setup. DeepSeek's DeepSeek-V4-Pro is clearly good at agentic coding (probably the best from the open-weight models), but the model is also great on knowledge-intensive tasks where reasoning matters. The agent pulled agentic engineering best practices from different company docs (Anthropic, OpenAI, Google, Stripe, Meta, Modal, DeepSeek, Mistral, Cohere), searched and digested Reddit and HN threads, summarized arxiv papers, and surfaced trending GitHub repos. Then it distilled everything into actionable tips across categories. I love the Wiki it built. The quality is really good. Here is a snapshot of what the wiki looks like: DeepSeek-V4-Pro handled the task without breaking stride. Multi-step research queries, code generation for scaffolding, context-heavy reasoning across disparate sources. For coding specifically, this is the first open-weight model that genuinely feels like a Codex or Claude Code experience. It compares in capability and actual multi-turn agentic work. What made the loop feel so responsive was Fireworks' inference speed (the fastest in the market) and the fact that they actually validate models at the systems level before shipping. No corrupted reasoning traces. Just fast, reliable iteration. The hybrid CSA and HCA attention design cuts KV cache to just 10% and inference FLOPs by nearly 4x at 1M-token context. This is what makes the agent loop actually fast and cheap enough to run in practice. For devs who've been watching open-weight models close the gap but haven't found one that actually delivers in practice, this is the closest I've seen. Try it here:

elvis

60,281 views • 4 months ago