Loading video...

Video Failed to Load

Go Home

Introducing Offloop! We're a team of four. Today our multi-agent harness hit state of the art on GDPval, ahead of Claude code and Codex across jobs that pay $2.4 trillion a year in the US. Offloop gives every knowledge worker what the Fortune 500 spends billions on: a high-performing...

2,641,437 views • 2 months ago •via X (Twitter)

44 Comments

Offloop's profile picture
Offloop2 months ago

Here is how Alice runs an entire company inside Offloop. Alice is building a private, local-first legal AI. She started with one instruction. Offloop assembled the agents army, organized their work into channels, and set up a shared flow across product and engineering team. The agents build in parallel - the desktop runtime, Word revisions grounded in verifiable citations, client documents locked to the device. When QA caught a restart failure, engineering fixed it while testing kept running. Alice reviewed the evidence in one thread and approved a limited private beta. Public release stays gated. One instruction in. One approval out. Everything in between was agents.

Offloop's profile picture
Offloop2 months ago

Meet D1, the world's first dispatcher model, and the best way to orchestrate a team of agents. When a client email lands in your workspace, something has to decide what happens next - which agent takes it, who reviews the draft, and which six agents should stay out of the way. That's D1's entire job. It's trained on real production dispatch decisions, and it's why our agents finish work without talking over each other or burning tokens they didn't need. On Dispatch Bench, our eval built from those real interactions: D1 scores 81%.

Offloop's profile picture
Offloop2 months ago

Offloop is still early. We are learning what context should last, when agents should act, and where human judgment matters most.

Offloop's profile picture
Offloop2 months ago

join the waitlist:

Yum⋆₊˚'s profile picture
Yum⋆₊˚2 months ago

@sama bet we’d see the first one-person billion-dollar company. We think the ceiling is at least 10x higher. Every David should have a real chance to defeat a Goliath.

Offloop's profile picture
Offloop2 months ago

@sama a big day for agents!

KP's profile picture
KP2 months ago

Nice!! Shared context across agents is the unlock. Tired of re-explaining my project to every new chat

Ori Mannheim's profile picture
Ori Mannheim2 months ago

LFG

elvis's profile picture
elvis2 months ago

Congrats on the launch! Looks like a great way to work with agents in teams. How does D1 decide when to pull a human in versus letting the agents keep going?

Alex / AI Experiments's profile picture
Alex / AI Experiments2 months ago

Great job! Congrats on the launch! We need a better way to communicate with agents. Mentions on Slack are terrible idea.

Robin Delta's profile picture
Robin Delta2 months ago

big day for agents

Offloop's profile picture
Offloop2 months ago

agent interface moves to a new milestone - I’m not joking

shirish's profile picture
shirish2 months ago

shipping something this capable with a team of four is crazy

Patrick's profile picture
Patrick2 months ago

Congratulations guys, this is sick - can I scale up my dropshipping business?

Offloop's profile picture
Offloop2 months ago

One of our early hypotheses is that AI-native commerce teams will be among the biggest winners!!! Would love to put that to the test with you🙌🙌

Bishal Nandi's profile picture
Bishal Nandi2 months ago

Big vision. Looking forward to seeing how it evolves. 👏

Jin's profile picture
Jin2 months ago

Here's how we run our entire company on Offloop. Across marketing, sales, and fund raising

Urooj's profile picture
Urooj2 months ago

Is GDPval the right benchmark or just the best one available today?

Offloop's profile picture
Offloop2 months ago

It's the most relevant benchmark we've found so far with a reasonable level of credibility. Our target use case is actually a subset of what GDPval evaluates. We also believe the ecosystem will eventually develop much better benchmarks specifically for measuring multi-agent collaboration, and we're looking forward to that.

Poonam Soni's profile picture
Poonam Soni2 months ago

Running entire company using Offloop is wild

Toha Khan's profile picture
Toha Khan2 months ago

What happens when two agents disagree on an edit?

Offloop's profile picture
Offloop2 months ago

Sometimes a third agent joins the discussion and makes the final call.

Madza 👨‍💻⚡'s profile picture
Madza 👨‍💻⚡2 months ago

Congrats on the launch, this is a big step forward! 🎉👍

Alamin's profile picture
Alamin2 months ago

Small teams are proving that agent orchestration is becoming the real competitive advantage.

Khushi's profile picture
Khushi2 months ago

this is impressive 🫡

AshutoshShrivastava's profile picture
AshutoshShrivastava2 months ago

5X cheaper damn....

Aute's profile picture
Aute2 months ago

We've been looking for a solution like this for our team. Could you spare an invite code so we can put it through some real workflow tests? Thanks!

Ava AI Labs's profile picture
Ava AI Labs2 months ago

I spend too much time re-sharing context between AI tools. This feels like a much better approach.

Markets & Mayhem's profile picture
Markets & Mayhem2 months ago

This is such an underappreciated concept, yet critical for managing costs and outcomes. Love to see the innovation! 🦾

Paul Couvert's profile picture
Paul Couvert2 months ago

I was thinking about fine-tuning a slm as a router but you executed 10x better than I could do! Well done!

Patrick's profile picture
Patrick2 months ago

How do I get an invite code?

Offloop's profile picture
Offloop2 months ago

Sure, please check your DM

Akshay Dawra's profile picture
Akshay Dawra2 months ago

Congratulations on the launch

H A J R A's profile picture
H A J R A2 months ago

Multi-agent systems are opening exciting possibilities for the future of knowledge work.

D-Coder's profile picture
D-Coder2 months ago

Curious to see how Offloop performs outside benchmarks. Real production workflows will be the real test.

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ2 months ago

This seems pretty cool! Can't wait to try it! Great job!

ɱҽԃι✨'s profile picture
ɱҽԃι✨2 months ago

Congratulations to the Offloop team on reaching state-of-the-art performance.

Steven Cheng's profile picture
Steven Cheng2 months ago

GDPval is a solid benchmark. Curious how you handle state consistency across those agents during long-running tasks.

Aaliya's profile picture
Aaliya2 months ago

congrats on the launch The interesting part is not only benchmark performance but also the idea of agent orchestration.

Monami's profile picture
Monami2 months ago

Excellent comparison, very insightful.

V's profile picture
V2 months ago

Congratulations to the fantastic 4

Rakib Hossen's profile picture
Rakib Hossen2 months ago

Congratulations! Excited to see autonomous agent teams tackling real-world knowledge work at scale.

Kirk Patrick Miller's profile picture
Kirk Patrick Miller2 months ago

Go on. •

Alex / AI Experiments's profile picture
Alex / AI Experiments2 months ago

Can I get invite codes for my team. Growth team of 4 people. Would love to try it out and probably implement in our workflow.

Related Videos

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 views • 5 months ago