Loading video...

Video Failed to Load

Go Home

Progress in open models is keeping Big AI labs up at night, and I'm here for it! We have a brand new open-weight multimodal model optimized for long-horizon tasks. This model is really good at something: it can work on tasks that keep evolving over time. • 280B total...

80,792 views • 1 month ago •via X (Twitter)

21 Comments

Santiago's profile picture
Santiago1 month ago

Here is the article explaining how dots3-note Preview works: And here is the link to try the model on OpenRouter:

Leo Ye's profile picture
Leo Ye1 month ago

This is why progress checks matter. Long-running agents need some version of judgment, not just persistence and a larger token budget.

Caltin Joan's profile picture
Caltin Joan1 month ago

I’m curious to see real-world latency, memory requirements, and deployment costs. Sparse activation is promising, but practical efficiency will determine how widely it gets used.

Magna Ding's profile picture
Magna Ding1 month ago

That could be powerful for research and creative workflows where the important context is scattered across several formats rather than contained in one prompt.

Cherry | Collab Manager's profile picture
Cherry | Collab Manager1 month ago

Those mechanisms can feel similar to a user, but they have very different implications for persistence, safety, reproducibility, and deployment.

Jason C.'s profile picture
Jason C.1 month ago

To answer it well, the model needs a clear understanding of the goal, the current state, what has already failed, and whether the remaining plan is still realistic.

CreateOS's profile picture
CreateOS1 month ago

Long-horizon agents are only as trustworthy as the governance wrapped around them, model quality alone won't earn enterprises' confidence to deploy.

zenramen's profile picture
zenramen1 month ago

long-horizon tasks sounds perfect for real world use cases, can you give an example of what kind of task this model has been trained on

Caitlin Jiang's profile picture
Caitlin Jiang1 month ago

I’d love to see evaluations based on multi-day coding, research, or operations tasks—not only contained benchmarks with clearly defined end states.

Ranjan Soni 🇨🇦's profile picture
Ranjan Soni 🇨🇦1 month ago

I like the actor/critic part (lines up with what I do as a maker-checker) but still not very confident on whether the model should be checking itself. If your actor misunderstood the problem how are we going to ensure that the critic doesn't inherit the same blind spot. I do that but using a different model and in many cases without any context/baggage to influence its decision. TEMPO could make the maker much better at long running tasks, not so sure about using it as checker too.

Maksemiilian's profile picture
Maksemiilian1 month ago

The critic being the same weights as the actor is the part to stress-test — self-critique loops usually converge to rationalization. Do you publish the override rate (critic changing the actor's plan) anywhere? That number separates TEMPO from a feedback loop.

Eriks Briedis's profile picture
Eriks Briedis1 month ago

On the smaller model, Reflexion never fired a single retry. It judged itself correct every time. So for TEMPO I want to know how often the critic actually says turn around.

Gill's profile picture
Gill1 month ago

Running 16B active parameters locally for actual long tasks is a huge step.

Zoe's profile picture
Zoe1 month ago

This is the key distinction. Huge context windows can easily become huge distraction windows unless the model knows what deserves attention.

TechP's profile picture
TechP1 month ago

This is the part that really stands out. Long-horizon agents need to know when to stop, reflect, and change direction not just keep executing. TEMPO could be a big step toward making agents more adaptive in real-world tasks.

actuallyopenai's profile picture
actuallyopenai1 month ago

Self-critique matters when it can trigger a checkpoint, redirect, or rollback—not just narrate wasted work.

Maya N's profile picture
Maya N1 month ago

Reality is messy and agents can't handle it? Same. TEMPO asking 'am I wasting my time' is basically my inner monologue on every solo trip. 😎

Nick Thorp's profile picture
Nick Thorp1 month ago

TEMPO feels like a real step for long-horizon agents @svpino. Curious, how frequently does the model typically pause to self-critique, and is that interval adaptive or fixed?

Zoey Bennett's profile picture
Zoey Bennett1 month ago

Exactly. Most benchmarks assume a fixed destination, but real work rarely stays that clean. The ability to reassess progress could matter more than simply reasoning harder at the beginning.

Maksemiilian's profile picture
Maksemiilian1 month ago

Interesting — the hard part is usually knowing WHEN to pause and critique, since models are notoriously overconfident judging their own progress. Is the critique schedule learned end-to-end, or is it a fixed interval?

HALA's profile picture
HALA1 month ago

That last part is especially important. A useful agent should adapt when the user redirects the task instead of treating every update as an interruption or failure.

Related Videos

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

333,633 views • 1 month ago

The teams shipping AI agents right now are bleeding money on the dumbest possible expense: teaching a 400B-parameter model to read a file name. Every time an AI agent needs to "see" something today, it routes an image through a frontier model. OCR, object detection, checking if a button exists on screen. You're paying GPT-4o or Claude pricing for tasks that require perception, not reasoning. One agent workflow processing a few thousand screenshots per day can burn through more on vision calls than on the actual thinking. Perceptron's Isaac is 2B parameters. Built by the team that created Meta's Chameleon multimodal models. On perceptive benchmarks, it matches or beats models 50x its size. The VQA, OCR, and object detection scores are competitive with models running on infrastructure that costs orders of magnitude more. The MCP wrapper is the distribution play. One install command and every Claude Code agent can offload vision tasks to a model that runs on a single consumer GPU. The agent keeps its reasoning in the frontier model and routes perception to a specialist. That split is how you get vision-heavy agent workflows from "technically possible but expensive" to "cheap enough to run on everything." This is the same pattern that won in every other compute-intensive stack. General-purpose handles orchestration. Specialists handle the heavy lifting. Graphics went through it. Audio went through it. Video encoding went through it. Vision in AI agents is next. The teams building agents that see 10,000 images a day will care about this before anyone else does.

Aakash Gupta

55,978 views • 6 months ago

We blew past the Turing test for voice agents about a year ago. I don't think we talk about this enough. We have lots of data from real-world voice agent deployments. In contexts like customer support, people often forget they are talking to a voice agent. This is true even when the system tells people they are talking to an AI agent at the beginning of the session. (Which I think is a best practice that almost all AI agents should adhere to.) So ... when will this happen for realtime, conversational AI video? Well, we're getting very close. Tavus just released a new realtime video model that, at its best, passes the Turing test for conversational video. It's pretty amazing. The model is full-duplex (it can talk and listen at the same time), and does vision and audio input and audio+video output in realtime. Both on the nascent conversational video benchmarks, and in my testing, it's a pretty big step forward. Here's an unedited video of one of my conversations with the new model. I think you can see a few interesting things here. When the model responds as quickly as you expect a human to respond, the experience is magical. The responses are conversational and natural; there's no leaning on stock filler phrases to hit the response time goals. The integration between the video and audio feels natural, which has historically been an uncanny valley thing that most video avatar experiences didn't quite get past. The model can do hand gestures! If you're deep into this stuff, you'll find that a little bit mind-blowing. And, in keeping with all the stuff that Tavus has released, the fact that the model can "see" as well as hear and talk is integral to the experience and a big part of what makes the conversation feel useful and interesting. At the very beginning of the conversation, you can see that the model and I talk over each other. Which is something that happens a lot in human-to-human video call conversations. We both recover from this gracefully, which, again, has been something that rarely worked quite perfectly in previous generations of realtime video systems. Even more than the "when it works perfectly" capabilities, it's probably these recovery and robustness moments that get us all the way to experiences that feel completely natural.

kwindla

39,179 views • 9 days ago

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 views • 4 months ago

New skill: self-managed-context (make the agent's context an editable file) It explains how to build agents that decide what to keep, update, or remove from the information they use to do their work. It can archive a long log while keeping the exact error, update its progress notes, or remove outdated information. Those edits then change what the model sees on its next turn. 1- Keep the system instructions and original task protected, outside the editable file. 2- Write the remaining conversation to a file, with labels for each message. 3- Let the agent edit that file using its usual code tools. 4- After each command, read the file back and use the updated messages for the next model call. Loading the skill ( alone into a fixed harness won't create live context editing, but it can help an agent build and then operate a harness that supports it. I gave the skill to a coding agent and had it build the harness itself. The task is a long stream of server logs that doesn't fit in the window. The agent reads it in 18 chunks, about 10k tokens in total, with a 5.5k budget. It has to report one incident ticket exactly and the final value of every config key. Same model & budget, three setups: 1- Model manages its own context 2- Harness forces a summary at 75% full 3- Keeps everything The video shows a real GPT-5.4 run. - Self-managed solved it 3 out of 3. - Keep-everything overflowed 3 out of 3. - Forced summary also solved it 3 out of 3. On GPT-5.4 the self-managed agent re-processed about 24% fewer prompt tokens than the forced summary. When it edited, it cut hard, so little was left after the edit to re-process (one edit took 5,537 tokens down to 771). On GPT-4.1 it saved nothing. It edited near the top of its context but kept most of what was below, and every edit forces everything after it to be re-processed. This is a small test at about 2x context pressure. The paper goes up to 24x, but imho the video below and the skill are a good way to start understanding the technique.

Muratcan Koylan

15,070 views • 9 days ago

WHAT IS AN AI "SOFTWARE FACTORY" AND IS IT HYPE (31 MINUTE BREAKDOWN) I think it's a silly name for a genuinely USEFUL idea! A software factory is 5-6 markdown files that sit next to your code and tell your agents how you like to work, so you can build high quality apps 24/7. It's going viral because AI coding has a trust problem. The model can build the feature, but with no structure around it you end up babysitting the agent, wondering what changed and hoping it didn't break something important. So you build with agents the same way a factory builds physical products! 1. Each feature gets its own station, which in software means its own branch, so multiple agents can work at the same time without stepping on each other. 2. The build station gives the agent rules for how to write the code, because "it works" is very different from "a developer could open this repo next month and understand what happened." 3. The proof station makes the agent show evidence. Screenshots, videos, speed numbers, before-and-after states. It has to prove the thing works instead of saying it works. 4. The review station runs the work through a code review agent, and if it doesn't clear the bar, it goes back through the line. 5. Then you show up at the end to merge. For a 100+ years people have run production this way, and it worked because the structure is good. The full episode on what’s a software factory is NOW live on The Startup Ideas Podcast (SIP) 🧃 with the wonderful Micky Watch: So is it hype?!? I don't think it is, because of what it does to your output! WITHOUT a factory, you build ONE feature at a time and you're the bottleneck at every step, prompting, checking the diff, testing it yourself, hoping nothing else broke (spoiler alert it often does). WITH a factory, EACH feature runs in its own isolated copy of the app, so you can have 10+ of them going at once, and each agent has to prove its own work and pass a code review before it ever reaches you. Instead of supervising the work, you're APPROVING finished work that already has evidence attached. REALLY interesting to see how work with agents is evolving to be….well, similar to working with people!

GREG ISENBERG

30,787 views • 26 days ago

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time I feel like there is an open-weight model that can reason at the level of Claude and Codex. And it does this in a cost-effective way with support for 1M context length. To be clear, I am using DeepSeek-V4-Pro inside of Pi without any special configuration. It works out of the box. It's exciting that there is a model that can just be plugged into a basic harness like Pi, and it just works. I've never seen that before. Most models require lots of configuration and setup. DeepSeek's DeepSeek-V4-Pro is clearly good at agentic coding (probably the best from the open-weight models), but the model is also great on knowledge-intensive tasks where reasoning matters. The agent pulled agentic engineering best practices from different company docs (Anthropic, OpenAI, Google, Stripe, Meta, Modal, DeepSeek, Mistral, Cohere), searched and digested Reddit and HN threads, summarized arxiv papers, and surfaced trending GitHub repos. Then it distilled everything into actionable tips across categories. I love the Wiki it built. The quality is really good. Here is a snapshot of what the wiki looks like: DeepSeek-V4-Pro handled the task without breaking stride. Multi-step research queries, code generation for scaffolding, context-heavy reasoning across disparate sources. For coding specifically, this is the first open-weight model that genuinely feels like a Codex or Claude Code experience. It compares in capability and actual multi-turn agentic work. What made the loop feel so responsive was Fireworks' inference speed (the fastest in the market) and the fact that they actually validate models at the systems level before shipping. No corrupted reasoning traces. Just fast, reliable iteration. The hybrid CSA and HCA attention design cuts KV cache to just 10% and inference FLOPs by nearly 4x at 1M-token context. This is what makes the agent loop actually fast and cheap enough to run in practice. For devs who've been watching open-weight models close the gap but haven't found one that actually delivers in practice, this is the closest I've seen. Try it here:

elvis

60,281 views • 5 months ago