Loading video...
Video Failed to Load
Progress in open models is keeping Big AI labs up at night, and I'm here for it! We have a brand new open-weight multimodal model optimized for long-horizon tasks. This model is really good at something: it can work on tasks that keep evolving over time. • 280B total... show more
80,792 views • 1 month ago •via X (Twitter)
21 Comments

Here is the article explaining how dots3-note Preview works: And here is the link to try the model on OpenRouter:

This is why progress checks matter. Long-running agents need some version of judgment, not just persistence and a larger token budget.

I’m curious to see real-world latency, memory requirements, and deployment costs. Sparse activation is promising, but practical efficiency will determine how widely it gets used.

That could be powerful for research and creative workflows where the important context is scattered across several formats rather than contained in one prompt.

Those mechanisms can feel similar to a user, but they have very different implications for persistence, safety, reproducibility, and deployment.

To answer it well, the model needs a clear understanding of the goal, the current state, what has already failed, and whether the remaining plan is still realistic.

Long-horizon agents are only as trustworthy as the governance wrapped around them, model quality alone won't earn enterprises' confidence to deploy.

long-horizon tasks sounds perfect for real world use cases, can you give an example of what kind of task this model has been trained on

I’d love to see evaluations based on multi-day coding, research, or operations tasks—not only contained benchmarks with clearly defined end states.

I like the actor/critic part (lines up with what I do as a maker-checker) but still not very confident on whether the model should be checking itself. If your actor misunderstood the problem how are we going to ensure that the critic doesn't inherit the same blind spot. I do that but using a different model and in many cases without any context/baggage to influence its decision. TEMPO could make the maker much better at long running tasks, not so sure about using it as checker too.

The critic being the same weights as the actor is the part to stress-test — self-critique loops usually converge to rationalization. Do you publish the override rate (critic changing the actor's plan) anywhere? That number separates TEMPO from a feedback loop.

On the smaller model, Reflexion never fired a single retry. It judged itself correct every time. So for TEMPO I want to know how often the critic actually says turn around.

Running 16B active parameters locally for actual long tasks is a huge step.

This is the key distinction. Huge context windows can easily become huge distraction windows unless the model knows what deserves attention.

This is the part that really stands out. Long-horizon agents need to know when to stop, reflect, and change direction not just keep executing. TEMPO could be a big step toward making agents more adaptive in real-world tasks.

Self-critique matters when it can trigger a checkpoint, redirect, or rollback—not just narrate wasted work.

Reality is messy and agents can't handle it? Same. TEMPO asking 'am I wasting my time' is basically my inner monologue on every solo trip. 😎

TEMPO feels like a real step for long-horizon agents @svpino. Curious, how frequently does the model typically pause to self-critique, and is that interval adaptive or fixed?

Exactly. Most benchmarks assume a fixed destination, but real work rarely stays that clean. The ability to reassess progress could matter more than simply reasoning harder at the beginning.

Interesting — the hard part is usually knowing WHEN to pause and critique, since models are notoriously overconfident judging their own progress. Is the critique schedule learned end-to-end, or is it a fixed interval?

That last part is especially important. A useful agent should adapt when the user redirects the task instead of treating every update as an interruption or failure.
