Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Offloop! We're a team of four. Today our multi-agent harness hit state of the art on GDPval, ahead of Claude code and Codex across jobs that pay $2.4 trillion a year in the US. Offloop gives every knowledge worker what the Fortune 500 spends billions on: a high-performing...

2,641,437 Aufrufe • vor 2 Monaten •via X (Twitter)

44 Kommentare

Profilbild von Offloop
Offloopvor 2 Monaten

Here is how Alice runs an entire company inside Offloop. Alice is building a private, local-first legal AI. She started with one instruction. Offloop assembled the agents army, organized their work into channels, and set up a shared flow across product and engineering team. The agents build in parallel - the desktop runtime, Word revisions grounded in verifiable citations, client documents locked to the device. When QA caught a restart failure, engineering fixed it while testing kept running. Alice reviewed the evidence in one thread and approved a limited private beta. Public release stays gated. One instruction in. One approval out. Everything in between was agents.

Profilbild von Offloop
Offloopvor 2 Monaten

Meet D1, the world's first dispatcher model, and the best way to orchestrate a team of agents. When a client email lands in your workspace, something has to decide what happens next - which agent takes it, who reviews the draft, and which six agents should stay out of the way. That's D1's entire job. It's trained on real production dispatch decisions, and it's why our agents finish work without talking over each other or burning tokens they didn't need. On Dispatch Bench, our eval built from those real interactions: D1 scores 81%.

Profilbild von Offloop
Offloopvor 2 Monaten

Offloop is still early. We are learning what context should last, when agents should act, and where human judgment matters most.

Profilbild von Offloop
Offloopvor 2 Monaten

join the waitlist:

Profilbild von Yum⋆₊˚
Yum⋆₊˚vor 2 Monaten

@sama bet we’d see the first one-person billion-dollar company. We think the ceiling is at least 10x higher. Every David should have a real chance to defeat a Goliath.

Profilbild von Offloop
Offloopvor 2 Monaten

@sama a big day for agents!

Profilbild von KP
KPvor 2 Monaten

Nice!! Shared context across agents is the unlock. Tired of re-explaining my project to every new chat

Profilbild von Ori Mannheim
Ori Mannheimvor 2 Monaten

LFG

Profilbild von elvis
elvisvor 2 Monaten

Congrats on the launch! Looks like a great way to work with agents in teams. How does D1 decide when to pull a human in versus letting the agents keep going?

Profilbild von Alex / AI Experiments
Alex / AI Experimentsvor 2 Monaten

Great job! Congrats on the launch! We need a better way to communicate with agents. Mentions on Slack are terrible idea.

Profilbild von Robin Delta
Robin Deltavor 2 Monaten

big day for agents

Profilbild von Offloop
Offloopvor 2 Monaten

agent interface moves to a new milestone - I’m not joking

Profilbild von shirish
shirishvor 2 Monaten

shipping something this capable with a team of four is crazy

Profilbild von Patrick
Patrickvor 2 Monaten

Congratulations guys, this is sick - can I scale up my dropshipping business?

Profilbild von Offloop
Offloopvor 2 Monaten

One of our early hypotheses is that AI-native commerce teams will be among the biggest winners!!! Would love to put that to the test with you🙌🙌

Profilbild von Bishal Nandi
Bishal Nandivor 2 Monaten

Big vision. Looking forward to seeing how it evolves. 👏

Profilbild von Jin
Jinvor 2 Monaten

Here's how we run our entire company on Offloop. Across marketing, sales, and fund raising

Profilbild von Urooj
Uroojvor 2 Monaten

Is GDPval the right benchmark or just the best one available today?

Profilbild von Offloop
Offloopvor 2 Monaten

It's the most relevant benchmark we've found so far with a reasonable level of credibility. Our target use case is actually a subset of what GDPval evaluates. We also believe the ecosystem will eventually develop much better benchmarks specifically for measuring multi-agent collaboration, and we're looking forward to that.

Profilbild von Poonam Soni
Poonam Sonivor 2 Monaten

Running entire company using Offloop is wild

Profilbild von Toha Khan
Toha Khanvor 2 Monaten

What happens when two agents disagree on an edit?

Profilbild von Offloop
Offloopvor 2 Monaten

Sometimes a third agent joins the discussion and makes the final call.

Profilbild von Madza 👨‍💻⚡
Madza 👨‍💻⚡vor 2 Monaten

Congrats on the launch, this is a big step forward! 🎉👍

Profilbild von Alamin
Alaminvor 2 Monaten

Small teams are proving that agent orchestration is becoming the real competitive advantage.

Profilbild von Khushi
Khushivor 2 Monaten

this is impressive 🫡

Profilbild von AshutoshShrivastava
AshutoshShrivastavavor 2 Monaten

5X cheaper damn....

Profilbild von Aute
Autevor 2 Monaten

We've been looking for a solution like this for our team. Could you spare an invite code so we can put it through some real workflow tests? Thanks!

Profilbild von Ava AI Labs
Ava AI Labsvor 2 Monaten

I spend too much time re-sharing context between AI tools. This feels like a much better approach.

Profilbild von Markets & Mayhem
Markets & Mayhemvor 2 Monaten

This is such an underappreciated concept, yet critical for managing costs and outcomes. Love to see the innovation! 🦾

Profilbild von Paul Couvert
Paul Couvertvor 2 Monaten

I was thinking about fine-tuning a slm as a router but you executed 10x better than I could do! Well done!

Profilbild von Patrick
Patrickvor 2 Monaten

How do I get an invite code?

Profilbild von Offloop
Offloopvor 2 Monaten

Sure, please check your DM

Profilbild von Akshay Dawra
Akshay Dawravor 2 Monaten

Congratulations on the launch

Profilbild von H A J R A
H A J R Avor 2 Monaten

Multi-agent systems are opening exciting possibilities for the future of knowledge work.

Profilbild von D-Coder
D-Codervor 2 Monaten

Curious to see how Offloop performs outside benchmarks. Real production workflows will be the real test.

Profilbild von Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅvor 2 Monaten

This seems pretty cool! Can't wait to try it! Great job!

Profilbild von ɱҽԃι✨
ɱҽԃι✨vor 2 Monaten

Congratulations to the Offloop team on reaching state-of-the-art performance.

Profilbild von Steven Cheng
Steven Chengvor 2 Monaten

GDPval is a solid benchmark. Curious how you handle state consistency across those agents during long-running tasks.

Profilbild von Aaliya
Aaliyavor 2 Monaten

congrats on the launch The interesting part is not only benchmark performance but also the idea of agent orchestration.

Profilbild von Monami
Monamivor 2 Monaten

Excellent comparison, very insightful.

Profilbild von V
Vvor 2 Monaten

Congratulations to the fantastic 4

Profilbild von Rakib Hossen
Rakib Hossenvor 2 Monaten

Congratulations! Excited to see autonomous agent teams tackling real-world knowledge work at scale.

Profilbild von Kirk Patrick Miller
Kirk Patrick Millervor 2 Monaten

Go on. •

Profilbild von Alex / AI Experiments
Alex / AI Experimentsvor 2 Monaten

Can I get invite codes for my team. Growth team of 4 people. Would love to try it out and probably implement in our workflow.

Ähnliche Videos

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 Aufrufe • vor 5 Monaten