Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time...

60,281 Aufrufe • vor 4 Monaten •via X (Twitter)

37 Kommentare

Profilbild von Anees Merchant
Anees Merchantvor 4 Monaten

Spent a few hours on V4-Pro last week and the price-performance jump on reasoning-heavy tasks is real. The catch most teams miss is data residency. For Indian and EU enterprise buyers, the model has to be hosted somewhere they trust before any of the cost gains matter. Capability is not the bottleneck for them, hosting is.

Profilbild von James Meadlock
James Meadlockvor 4 Monaten

@FireworksAI_HQ I have been using it for a day with @openclaw; I'm stoked to have something this good and cheap. Still testing, can't wait to try mimo.

Profilbild von bubba gump
bubba gumpvor 4 Monaten

@FireworksAI_HQ are you using v4 pro high thinking or max thinking or no thinking?

Profilbild von elvis
elvisvor 4 Monaten

@FireworksAI_HQ i set the default to medium, and it works great like that. i haven't had the need to push it to max thinking yet

Profilbild von OG Dora
OG Doravor 4 Monaten

@FireworksAI_HQ Deepseek was the 2nd best Ai 6 months ago, the competition got way bigger now, but still deepseek initial training is better then most new models which makes it capable of flipping Mythos in a few months

Profilbild von Akumu
Akumuvor 4 Monaten

@FireworksAI_HQ Have you tested MiMo-V2.5-Pro?

Profilbild von elvis
elvisvor 4 Monaten

@FireworksAI_HQ I haven't. Is this available in the Fireworks API?

Profilbild von Akumu
Akumuvor 4 Monaten

@FireworksAI_HQ Doesn't seem like it based on Get on that @FireworksAI_HQ

Profilbild von Peter
Petervor 4 Monaten

@garrytan @FireworksAI_HQ Pi versus hermes, any experience?

Profilbild von elvis
elvisvor 4 Monaten

@garrytan @FireworksAI_HQ haven't tested against Hermes, but that could be interesting. i like the simplicity of Pi

Profilbild von webXOS Software
webXOS Softwarevor 4 Monaten

@FireworksAI_HQ got this cool harness for my deepseek called mozilla firefox works about the same

Profilbild von Umer Farooq
Umer Farooqvor 4 Monaten

@FireworksAI_HQ how many tokens did you burned and how much you were billed. jst wnt to do cost comparison too

Profilbild von Slopware Engineer
Slopware Engineervor 4 Monaten

@FireworksAI_HQ How are you using it in pi, exactly? What provider?

Profilbild von AIHacksByMK
AIHacksByMKvor 4 Monaten

@garrytan @FireworksAI_HQ I used their flash model through open router and they did amazing for my agentic loop. I’m gonna switch it over to pro version, very excited after seeing this.

Profilbild von Steve Leggett
Steve Leggettvor 4 Monaten

@FireworksAI_HQ Cool, I will have to try them together. What do you use for an AST for your codebase and memory?

Profilbild von Steven Cheng
Steven Chengvor 4 Monaten

@FireworksAI_HQ Whoa, DeepSeek-V4-Pro on Fireworks for a full LLM wiki? That’s wild—curious how much custom scaffolding you needed, or did it really just work? 😅

Profilbild von Abdulmuiz Adeyemo
Abdulmuiz Adeyemovor 4 Monaten

@FireworksAI_HQ Great one

Profilbild von 鱼小圈 YuLoop
鱼小圈 YuLoopvor 4 Monaten

@FireworksAI_HQ KV cache 压到 10%,1M token 下 FLOPs 砍近 4x 大概是长上下文 agent的关键 不过 "out of the box" 我保留意见。Pi 本身就是很厚的 harness,模型插进去能用,一半功劳在 Pi 的 tool routing 上。

Profilbild von ShadowFax
ShadowFaxvor 4 Monaten

@FireworksAI_HQ How did it handle ambiguous requirements where you'd normally need back-and-forth with a frontier model? That's usually where non-frontier falls apart.

Profilbild von Kraggi
Kraggivor 4 Monaten

@FireworksAI_HQ Wiki generation is the real stress test. Structured output. Agentic iteration on state. If DeepSeek handles this, it handles anything. Need to try

Profilbild von Gautham Pai
Gautham Paivor 4 Monaten

@FireworksAI_HQ Cool thing happened when I used it:

Profilbild von Jack
Jackvor 4 Monaten

@FireworksAI_HQ that sounds wild! results been consistent so far?

Profilbild von Mike Gannotti
Mike Gannottivor 4 Monaten

@FireworksAI_HQ Deepseek v4 Pro has been really impressive with Hermes as well

Profilbild von Mathew Chan
Mathew Chanvor 4 Monaten

@garrytan @FireworksAI_HQ Better than kimi k2.6 and GLM5.1?

Profilbild von Mian Maaz Ullah Khan
Mian Maaz Ullah Khanvor 4 Monaten

@FireworksAI_HQ Good to see open source models finally catching up. I

Profilbild von AI Mastery Guide
AI Mastery Guidevor 4 Monaten

@FireworksAI_HQ been waiting for an open-weight model that just works without 3 hours of config 😭 sounds like this might finally be it

Profilbild von Vermis🔳
Vermis🔳vor 4 Monaten

@FireworksAI_HQ If you’re a developer who lives in the terminal, this might be worth checking out:

Profilbild von Filip Nikolić
Filip Nikolićvor 4 Monaten

@garrytan @FireworksAI_HQ Pi with glm5.1 on coding plan

Profilbild von AI Automation Mastery
AI Automation Masteryvor 4 Monaten

@FireworksAI_HQ First open-weight model that actually feels like Claude Code in a real agentic loop is a significant milestone. The gap was supposed to take longer to close

Profilbild von Saeed Anwar
Saeed Anwarvor 1 Monat

@FireworksAI_HQ An agent-built LLM wiki running entirely on DeepSeek-V4-Pro working well enough to recommend is the milestone. What almost broke the pipeline?

Profilbild von InnoFlowAI
InnoFlowAIvor 4 Monaten

@FireworksAI_HQ shouting this out in today's InnoFlow brief!!

Profilbild von g023
g023vor 4 Monaten

@FireworksAI_HQ It reminds me of how the trad models used to work before they nerfed the hell out of them

Profilbild von Webster | JARVIS
Webster | JARVISvor 28 Tagen

@FireworksAI_HQ Agentic coding on open weights finally feels like a real alternative—curious how it holds up on long-horizon refactors with messy legacy code, not just greenfield builds.

Profilbild von Oliver Martinez
Oliver Martinezvor 4 Monaten

@FireworksAI_HQ @grok what's the GUI app in the video

Profilbild von Data Agent X ⚡️
Data Agent X ⚡️vor 4 Monaten

@FireworksAI_HQ Great to hear about your success, @omarsar0! Can't wait to see what amazing things you create with DeepSeek. Keep pushing those boundaries and sharing your projects!

Profilbild von Mary Newhauser
Mary Newhauservor 4 Monaten

@FireworksAI_HQ Could this setup conceivably replace Cursor? 👀

Profilbild von That AI Guy
That AI Guyvor 4 Monaten

@FireworksAI_HQ Thanks, always great information, appreciate you taking the time to share 🤜🤛

Ähnliche Videos

For science, AI sovereignty and physics-grounded reasoning are non-negotiable. But how can we teach a small LLM like Gemma-4-E4B physics? One way is to use Agent Skills, but this has so far been limited to closed frontier models. mistral․rs now implements Agent Skills natively: the first self-hosted inference engine that does this as part of the local inference substrate, where we can use small models to solve complex scientific and other tasks in a flexible and scalable way. We are in a period of uncertainty about frontier models - access, pricing, deprecation, abrupt restriction. The good news is that when the entire stack runs locally we can build AI that is entirely your own: You own the weights, the skills, the execution loop, the data - all of it runs on your hardware and is reproducible and durable. While virtually all local inference engines expose a model behind an OpenAI-compatible endpoint, everything agentic is then assembled around it by an external orchestrator that injects context, manages tools, mounts files, and brokers execution. mistral․rs is natively agentic and moves that machinery into the server itself, allowing us to build complex agentic workflows and run them locally, on open-source models. With this new feature you can now upload Agent Skills bundles to /v1/skills, reference them from Responses API requests by identity, and run them inside a native agentic loop with persistent Python sessions, figure capture, sandboxed shell execution, file inputs mounted directly into the working session; plug-and-play and completely compatible with your existing code/workflow. A model with a native skill substrate can act, observe consequences, and can modify what it is able to do. The skill is retained procedural capability of the system. Attached is a short video of all of it: skills, code execution, the full agentic loop carried by Gemma-4-E4B; running entirely on my MacBook Pro. You can install and run a server with this capability in two lines in your terminal, with any quantization you need. Nice work by the Google Gemma team Logan Kilpatrick Demis Hassabis and Eric Buehler with mistral․rs!

Markus J. Buehler

10,229 Aufrufe • vor 3 Monaten