Загрузка видео...

Не удалось загрузить видео

На главную

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time...

60,281 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 37

Фото профиля Anees Merchant
Anees Merchant4 месяцев назад

Spent a few hours on V4-Pro last week and the price-performance jump on reasoning-heavy tasks is real. The catch most teams miss is data residency. For Indian and EU enterprise buyers, the model has to be hosted somewhere they trust before any of the cost gains matter. Capability is not the bottleneck for them, hosting is.

Фото профиля James Meadlock
James Meadlock4 месяцев назад

@FireworksAI_HQ I have been using it for a day with @openclaw; I'm stoked to have something this good and cheap. Still testing, can't wait to try mimo.

Фото профиля bubba gump
bubba gump4 месяцев назад

@FireworksAI_HQ are you using v4 pro high thinking or max thinking or no thinking?

Фото профиля elvis
elvis4 месяцев назад

@FireworksAI_HQ i set the default to medium, and it works great like that. i haven't had the need to push it to max thinking yet

Фото профиля OG Dora
OG Dora4 месяцев назад

@FireworksAI_HQ Deepseek was the 2nd best Ai 6 months ago, the competition got way bigger now, but still deepseek initial training is better then most new models which makes it capable of flipping Mythos in a few months

Фото профиля Akumu
Akumu4 месяцев назад

@FireworksAI_HQ Have you tested MiMo-V2.5-Pro?

Фото профиля elvis
elvis4 месяцев назад

@FireworksAI_HQ I haven't. Is this available in the Fireworks API?

Фото профиля Akumu
Akumu4 месяцев назад

@FireworksAI_HQ Doesn't seem like it based on Get on that @FireworksAI_HQ

Фото профиля Peter
Peter4 месяцев назад

@garrytan @FireworksAI_HQ Pi versus hermes, any experience?

Фото профиля elvis
elvis4 месяцев назад

@garrytan @FireworksAI_HQ haven't tested against Hermes, but that could be interesting. i like the simplicity of Pi

Фото профиля webXOS Software
webXOS Software4 месяцев назад

@FireworksAI_HQ got this cool harness for my deepseek called mozilla firefox works about the same

Фото профиля Umer Farooq
Umer Farooq4 месяцев назад

@FireworksAI_HQ how many tokens did you burned and how much you were billed. jst wnt to do cost comparison too

Фото профиля Slopware Engineer
Slopware Engineer4 месяцев назад

@FireworksAI_HQ How are you using it in pi, exactly? What provider?

Фото профиля AIHacksByMK
AIHacksByMK4 месяцев назад

@garrytan @FireworksAI_HQ I used their flash model through open router and they did amazing for my agentic loop. I’m gonna switch it over to pro version, very excited after seeing this.

Фото профиля Steve Leggett
Steve Leggett4 месяцев назад

@FireworksAI_HQ Cool, I will have to try them together. What do you use for an AST for your codebase and memory?

Фото профиля Steven Cheng
Steven Cheng4 месяцев назад

@FireworksAI_HQ Whoa, DeepSeek-V4-Pro on Fireworks for a full LLM wiki? That’s wild—curious how much custom scaffolding you needed, or did it really just work? 😅

Фото профиля Abdulmuiz Adeyemo
Abdulmuiz Adeyemo4 месяцев назад

@FireworksAI_HQ Great one

Фото профиля 鱼小圈 YuLoop
鱼小圈 YuLoop4 месяцев назад

@FireworksAI_HQ KV cache 压到 10%,1M token 下 FLOPs 砍近 4x 大概是长上下文 agent的关键 不过 "out of the box" 我保留意见。Pi 本身就是很厚的 harness,模型插进去能用,一半功劳在 Pi 的 tool routing 上。

Фото профиля ShadowFax
ShadowFax4 месяцев назад

@FireworksAI_HQ How did it handle ambiguous requirements where you'd normally need back-and-forth with a frontier model? That's usually where non-frontier falls apart.

Фото профиля Kraggi
Kraggi4 месяцев назад

@FireworksAI_HQ Wiki generation is the real stress test. Structured output. Agentic iteration on state. If DeepSeek handles this, it handles anything. Need to try

Фото профиля Gautham Pai
Gautham Pai4 месяцев назад

@FireworksAI_HQ Cool thing happened when I used it:

Фото профиля Jack
Jack4 месяцев назад

@FireworksAI_HQ that sounds wild! results been consistent so far?

Фото профиля Mike Gannotti
Mike Gannotti4 месяцев назад

@FireworksAI_HQ Deepseek v4 Pro has been really impressive with Hermes as well

Фото профиля Mathew Chan
Mathew Chan4 месяцев назад

@garrytan @FireworksAI_HQ Better than kimi k2.6 and GLM5.1?

Фото профиля Mian Maaz Ullah Khan
Mian Maaz Ullah Khan4 месяцев назад

@FireworksAI_HQ Good to see open source models finally catching up. I

Фото профиля AI Mastery Guide
AI Mastery Guide4 месяцев назад

@FireworksAI_HQ been waiting for an open-weight model that just works without 3 hours of config 😭 sounds like this might finally be it

Фото профиля Vermis🔳
Vermis🔳4 месяцев назад

@FireworksAI_HQ If you’re a developer who lives in the terminal, this might be worth checking out:

Фото профиля Filip Nikolić
Filip Nikolić4 месяцев назад

@garrytan @FireworksAI_HQ Pi with glm5.1 on coding plan

Фото профиля AI Automation Mastery
AI Automation Mastery4 месяцев назад

@FireworksAI_HQ First open-weight model that actually feels like Claude Code in a real agentic loop is a significant milestone. The gap was supposed to take longer to close

Фото профиля Saeed Anwar
Saeed Anwar1 месяц назад

@FireworksAI_HQ An agent-built LLM wiki running entirely on DeepSeek-V4-Pro working well enough to recommend is the milestone. What almost broke the pipeline?

Фото профиля InnoFlowAI
InnoFlowAI4 месяцев назад

@FireworksAI_HQ shouting this out in today's InnoFlow brief!!

Фото профиля g023
g0234 месяцев назад

@FireworksAI_HQ It reminds me of how the trad models used to work before they nerfed the hell out of them

Фото профиля Webster | JARVIS
Webster | JARVIS28 дней назад

@FireworksAI_HQ Agentic coding on open weights finally feels like a real alternative—curious how it holds up on long-horizon refactors with messy legacy code, not just greenfield builds.

Фото профиля Oliver Martinez
Oliver Martinez4 месяцев назад

@FireworksAI_HQ @grok what's the GUI app in the video

Фото профиля Data Agent X ⚡️
Data Agent X ⚡️4 месяцев назад

@FireworksAI_HQ Great to hear about your success, @omarsar0! Can't wait to see what amazing things you create with DeepSeek. Keep pushing those boundaries and sharing your projects!

Фото профиля Mary Newhauser
Mary Newhauser4 месяцев назад

@FireworksAI_HQ Could this setup conceivably replace Cursor? 👀

Фото профиля That AI Guy
That AI Guy4 месяцев назад

@FireworksAI_HQ Thanks, always great information, appreciate you taking the time to share 🤜🤛

Похожие видео

For science, AI sovereignty and physics-grounded reasoning are non-negotiable. But how can we teach a small LLM like Gemma-4-E4B physics? One way is to use Agent Skills, but this has so far been limited to closed frontier models. mistral․rs now implements Agent Skills natively: the first self-hosted inference engine that does this as part of the local inference substrate, where we can use small models to solve complex scientific and other tasks in a flexible and scalable way. We are in a period of uncertainty about frontier models - access, pricing, deprecation, abrupt restriction. The good news is that when the entire stack runs locally we can build AI that is entirely your own: You own the weights, the skills, the execution loop, the data - all of it runs on your hardware and is reproducible and durable. While virtually all local inference engines expose a model behind an OpenAI-compatible endpoint, everything agentic is then assembled around it by an external orchestrator that injects context, manages tools, mounts files, and brokers execution. mistral․rs is natively agentic and moves that machinery into the server itself, allowing us to build complex agentic workflows and run them locally, on open-source models. With this new feature you can now upload Agent Skills bundles to /v1/skills, reference them from Responses API requests by identity, and run them inside a native agentic loop with persistent Python sessions, figure capture, sandboxed shell execution, file inputs mounted directly into the working session; plug-and-play and completely compatible with your existing code/workflow. A model with a native skill substrate can act, observe consequences, and can modify what it is able to do. The skill is retained procedural capability of the system. Attached is a short video of all of it: skills, code execution, the full agentic loop carried by Gemma-4-E4B; running entirely on my MacBook Pro. You can install and run a server with this capability in two lines in your terminal, with any quantization you need. Nice work by the Google Gemma team Logan Kilpatrick Demis Hassabis and Eric Buehler with mistral․rs!

Markus J. Buehler

10,229 просмотров • 3 месяцев назад

Progress in open models is keeping Big AI labs up at night, and I'm here for it! We have a brand new open-weight multimodal model optimized for long-horizon tasks. This model is really good at something: it can work on tasks that keep evolving over time. • 280B total parameters, but only 16B active • 512K context window • Understands text, images, and audio • Strong reasoning, coding, and tool use But the best of all: the model learns and adapts to new information! Imagine you start running an agent today to solve a problem, and while it's working, you get new information that changes the initial conditions, or you change your mind. The agents you run today don't have issues with short tasks and goals that don't change, but reality is messy, and that makes it hard for long-horizon agents to succeed. The new dots3-note Preview model introduces TEMPO. TEMPO is a new reinforcement learning technique that lets the model periodically pause and critique its own progress. Basically, from time to time, the agent asks itself: "Am I getting closer to the goal, or am I wasting my time?" The same model switches between actor and critic. The actor works on the problem. The critic looks at the current state, reasons about how much progress it has made, and determines what should happen next. TEMPO gives the model feedback along the way. This is huge for any agent that can work on long-horizon tasks without wasting its time.

Santiago

80,792 просмотров • 1 месяц назад