Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Offloop! We're a team of four. Today our multi-agent harness hit state of the art on GDPval, ahead of Claude code and Codex across jobs that pay $2.4 trillion a year in the US. Offloop gives every knowledge worker what the Fortune 500 spends billions on: a high-performing...

2,641,437 görüntüleme • 2 ay önce •via X (Twitter)

44 Yorum

Offloop profil fotoğrafı
Offloop2 ay önce

Here is how Alice runs an entire company inside Offloop. Alice is building a private, local-first legal AI. She started with one instruction. Offloop assembled the agents army, organized their work into channels, and set up a shared flow across product and engineering team. The agents build in parallel - the desktop runtime, Word revisions grounded in verifiable citations, client documents locked to the device. When QA caught a restart failure, engineering fixed it while testing kept running. Alice reviewed the evidence in one thread and approved a limited private beta. Public release stays gated. One instruction in. One approval out. Everything in between was agents.

Offloop profil fotoğrafı
Offloop2 ay önce

Meet D1, the world's first dispatcher model, and the best way to orchestrate a team of agents. When a client email lands in your workspace, something has to decide what happens next - which agent takes it, who reviews the draft, and which six agents should stay out of the way. That's D1's entire job. It's trained on real production dispatch decisions, and it's why our agents finish work without talking over each other or burning tokens they didn't need. On Dispatch Bench, our eval built from those real interactions: D1 scores 81%.

Offloop profil fotoğrafı
Offloop2 ay önce

Offloop is still early. We are learning what context should last, when agents should act, and where human judgment matters most.

Offloop profil fotoğrafı
Offloop2 ay önce

join the waitlist:

Yum⋆₊˚ profil fotoğrafı
Yum⋆₊˚2 ay önce

@sama bet we’d see the first one-person billion-dollar company. We think the ceiling is at least 10x higher. Every David should have a real chance to defeat a Goliath.

Offloop profil fotoğrafı
Offloop2 ay önce

@sama a big day for agents!

KP profil fotoğrafı
KP2 ay önce

Nice!! Shared context across agents is the unlock. Tired of re-explaining my project to every new chat

Ori Mannheim profil fotoğrafı
Ori Mannheim2 ay önce

LFG

elvis profil fotoğrafı
elvis2 ay önce

Congrats on the launch! Looks like a great way to work with agents in teams. How does D1 decide when to pull a human in versus letting the agents keep going?

Alex / AI Experiments profil fotoğrafı
Alex / AI Experiments2 ay önce

Great job! Congrats on the launch! We need a better way to communicate with agents. Mentions on Slack are terrible idea.

Robin Delta profil fotoğrafı
Robin Delta2 ay önce

big day for agents

Offloop profil fotoğrafı
Offloop2 ay önce

agent interface moves to a new milestone - I’m not joking

shirish profil fotoğrafı
shirish2 ay önce

shipping something this capable with a team of four is crazy

Patrick profil fotoğrafı
Patrick2 ay önce

Congratulations guys, this is sick - can I scale up my dropshipping business?

Offloop profil fotoğrafı
Offloop2 ay önce

One of our early hypotheses is that AI-native commerce teams will be among the biggest winners!!! Would love to put that to the test with you🙌🙌

Bishal Nandi profil fotoğrafı
Bishal Nandi2 ay önce

Big vision. Looking forward to seeing how it evolves. 👏

Jin profil fotoğrafı
Jin2 ay önce

Here's how we run our entire company on Offloop. Across marketing, sales, and fund raising

Urooj profil fotoğrafı
Urooj2 ay önce

Is GDPval the right benchmark or just the best one available today?

Offloop profil fotoğrafı
Offloop2 ay önce

It's the most relevant benchmark we've found so far with a reasonable level of credibility. Our target use case is actually a subset of what GDPval evaluates. We also believe the ecosystem will eventually develop much better benchmarks specifically for measuring multi-agent collaboration, and we're looking forward to that.

Poonam Soni profil fotoğrafı
Poonam Soni2 ay önce

Running entire company using Offloop is wild

Toha Khan profil fotoğrafı
Toha Khan2 ay önce

What happens when two agents disagree on an edit?

Offloop profil fotoğrafı
Offloop2 ay önce

Sometimes a third agent joins the discussion and makes the final call.

Madza 👨‍💻⚡ profil fotoğrafı
Madza 👨‍💻⚡2 ay önce

Congrats on the launch, this is a big step forward! 🎉👍

Alamin profil fotoğrafı
Alamin2 ay önce

Small teams are proving that agent orchestration is becoming the real competitive advantage.

Khushi profil fotoğrafı
Khushi2 ay önce

this is impressive 🫡

AshutoshShrivastava profil fotoğrafı
AshutoshShrivastava2 ay önce

5X cheaper damn....

Aute profil fotoğrafı
Aute2 ay önce

We've been looking for a solution like this for our team. Could you spare an invite code so we can put it through some real workflow tests? Thanks!

Ava AI Labs profil fotoğrafı
Ava AI Labs2 ay önce

I spend too much time re-sharing context between AI tools. This feels like a much better approach.

Markets & Mayhem profil fotoğrafı
Markets & Mayhem2 ay önce

This is such an underappreciated concept, yet critical for managing costs and outcomes. Love to see the innovation! 🦾

Paul Couvert profil fotoğrafı
Paul Couvert2 ay önce

I was thinking about fine-tuning a slm as a router but you executed 10x better than I could do! Well done!

Patrick profil fotoğrafı
Patrick2 ay önce

How do I get an invite code?

Offloop profil fotoğrafı
Offloop2 ay önce

Sure, please check your DM

Akshay Dawra profil fotoğrafı
Akshay Dawra2 ay önce

Congratulations on the launch

H A J R A profil fotoğrafı
H A J R A2 ay önce

Multi-agent systems are opening exciting possibilities for the future of knowledge work.

D-Coder profil fotoğrafı
D-Coder2 ay önce

Curious to see how Offloop performs outside benchmarks. Real production workflows will be the real test.

Ivan Fioravanti ᯅ profil fotoğrafı
Ivan Fioravanti ᯅ2 ay önce

This seems pretty cool! Can't wait to try it! Great job!

ɱҽԃι✨ profil fotoğrafı
ɱҽԃι✨2 ay önce

Congratulations to the Offloop team on reaching state-of-the-art performance.

Steven Cheng profil fotoğrafı
Steven Cheng2 ay önce

GDPval is a solid benchmark. Curious how you handle state consistency across those agents during long-running tasks.

Aaliya profil fotoğrafı
Aaliya2 ay önce

congrats on the launch The interesting part is not only benchmark performance but also the idea of agent orchestration.

Monami profil fotoğrafı
Monami2 ay önce

Excellent comparison, very insightful.

V profil fotoğrafı
V2 ay önce

Congratulations to the fantastic 4

Rakib Hossen profil fotoğrafı
Rakib Hossen2 ay önce

Congratulations! Excited to see autonomous agent teams tackling real-world knowledge work at scale.

Kirk Patrick Miller profil fotoğrafı
Kirk Patrick Miller2 ay önce

Go on. •

Alex / AI Experiments profil fotoğrafı
Alex / AI Experiments2 ay önce

Can I get invite codes for my team. Growth team of 4 people. Would love to try it out and probably implement in our workflow.

Benzer Videolar

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 görüntüleme • 5 ay önce