Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Offloop! We're a team of four. Today our multi-agent harness hit state of the art on GDPval, ahead of Claude code and Codex across jobs that pay $2.4 trillion a year in the US. Offloop gives every knowledge worker what the Fortune 500 spends billions on: a high-performing...

2,641,437 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 44

Фото профиля Offloop
Offloop2 месяцев назад

Here is how Alice runs an entire company inside Offloop. Alice is building a private, local-first legal AI. She started with one instruction. Offloop assembled the agents army, organized their work into channels, and set up a shared flow across product and engineering team. The agents build in parallel - the desktop runtime, Word revisions grounded in verifiable citations, client documents locked to the device. When QA caught a restart failure, engineering fixed it while testing kept running. Alice reviewed the evidence in one thread and approved a limited private beta. Public release stays gated. One instruction in. One approval out. Everything in between was agents.

Фото профиля Offloop
Offloop2 месяцев назад

Meet D1, the world's first dispatcher model, and the best way to orchestrate a team of agents. When a client email lands in your workspace, something has to decide what happens next - which agent takes it, who reviews the draft, and which six agents should stay out of the way. That's D1's entire job. It's trained on real production dispatch decisions, and it's why our agents finish work without talking over each other or burning tokens they didn't need. On Dispatch Bench, our eval built from those real interactions: D1 scores 81%.

Фото профиля Offloop
Offloop2 месяцев назад

Offloop is still early. We are learning what context should last, when agents should act, and where human judgment matters most.

Фото профиля Offloop
Offloop2 месяцев назад

join the waitlist:

Фото профиля Yum⋆₊˚
Yum⋆₊˚2 месяцев назад

@sama bet we’d see the first one-person billion-dollar company. We think the ceiling is at least 10x higher. Every David should have a real chance to defeat a Goliath.

Фото профиля Offloop
Offloop2 месяцев назад

@sama a big day for agents!

Фото профиля KP
KP2 месяцев назад

Nice!! Shared context across agents is the unlock. Tired of re-explaining my project to every new chat

Фото профиля Ori Mannheim
Ori Mannheim2 месяцев назад

LFG

Фото профиля elvis
elvis2 месяцев назад

Congrats on the launch! Looks like a great way to work with agents in teams. How does D1 decide when to pull a human in versus letting the agents keep going?

Фото профиля Alex / AI Experiments
Alex / AI Experiments2 месяцев назад

Great job! Congrats on the launch! We need a better way to communicate with agents. Mentions on Slack are terrible idea.

Фото профиля Robin Delta
Robin Delta2 месяцев назад

big day for agents

Фото профиля Offloop
Offloop2 месяцев назад

agent interface moves to a new milestone - I’m not joking

Фото профиля shirish
shirish2 месяцев назад

shipping something this capable with a team of four is crazy

Фото профиля Patrick
Patrick2 месяцев назад

Congratulations guys, this is sick - can I scale up my dropshipping business?

Фото профиля Offloop
Offloop2 месяцев назад

One of our early hypotheses is that AI-native commerce teams will be among the biggest winners!!! Would love to put that to the test with you🙌🙌

Фото профиля Bishal Nandi
Bishal Nandi2 месяцев назад

Big vision. Looking forward to seeing how it evolves. 👏

Фото профиля Jin
Jin2 месяцев назад

Here's how we run our entire company on Offloop. Across marketing, sales, and fund raising

Фото профиля Urooj
Urooj2 месяцев назад

Is GDPval the right benchmark or just the best one available today?

Фото профиля Offloop
Offloop2 месяцев назад

It's the most relevant benchmark we've found so far with a reasonable level of credibility. Our target use case is actually a subset of what GDPval evaluates. We also believe the ecosystem will eventually develop much better benchmarks specifically for measuring multi-agent collaboration, and we're looking forward to that.

Фото профиля Poonam Soni
Poonam Soni2 месяцев назад

Running entire company using Offloop is wild

Фото профиля Toha Khan
Toha Khan2 месяцев назад

What happens when two agents disagree on an edit?

Фото профиля Offloop
Offloop2 месяцев назад

Sometimes a third agent joins the discussion and makes the final call.

Фото профиля Madza 👨‍💻⚡
Madza 👨‍💻⚡2 месяцев назад

Congrats on the launch, this is a big step forward! 🎉👍

Фото профиля Alamin
Alamin2 месяцев назад

Small teams are proving that agent orchestration is becoming the real competitive advantage.

Фото профиля Khushi
Khushi2 месяцев назад

this is impressive 🫡

Фото профиля AshutoshShrivastava
AshutoshShrivastava2 месяцев назад

5X cheaper damn....

Фото профиля Aute
Aute2 месяцев назад

We've been looking for a solution like this for our team. Could you spare an invite code so we can put it through some real workflow tests? Thanks!

Фото профиля Ava AI Labs
Ava AI Labs2 месяцев назад

I spend too much time re-sharing context between AI tools. This feels like a much better approach.

Фото профиля Markets & Mayhem
Markets & Mayhem2 месяцев назад

This is such an underappreciated concept, yet critical for managing costs and outcomes. Love to see the innovation! 🦾

Фото профиля Paul Couvert
Paul Couvert2 месяцев назад

I was thinking about fine-tuning a slm as a router but you executed 10x better than I could do! Well done!

Фото профиля Patrick
Patrick2 месяцев назад

How do I get an invite code?

Фото профиля Offloop
Offloop2 месяцев назад

Sure, please check your DM

Фото профиля Akshay Dawra
Akshay Dawra2 месяцев назад

Congratulations on the launch

Фото профиля H A J R A
H A J R A2 месяцев назад

Multi-agent systems are opening exciting possibilities for the future of knowledge work.

Фото профиля D-Coder
D-Coder2 месяцев назад

Curious to see how Offloop performs outside benchmarks. Real production workflows will be the real test.

Фото профиля Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅ2 месяцев назад

This seems pretty cool! Can't wait to try it! Great job!

Фото профиля ɱҽԃι✨
ɱҽԃι✨2 месяцев назад

Congratulations to the Offloop team on reaching state-of-the-art performance.

Фото профиля Steven Cheng
Steven Cheng2 месяцев назад

GDPval is a solid benchmark. Curious how you handle state consistency across those agents during long-running tasks.

Фото профиля Aaliya
Aaliya2 месяцев назад

congrats on the launch The interesting part is not only benchmark performance but also the idea of agent orchestration.

Фото профиля Monami
Monami2 месяцев назад

Excellent comparison, very insightful.

Фото профиля V
V2 месяцев назад

Congratulations to the fantastic 4

Фото профиля Rakib Hossen
Rakib Hossen2 месяцев назад

Congratulations! Excited to see autonomous agent teams tackling real-world knowledge work at scale.

Фото профиля Kirk Patrick Miller
Kirk Patrick Miller2 месяцев назад

Go on. •

Фото профиля Alex / AI Experiments
Alex / AI Experiments2 месяцев назад

Can I get invite codes for my team. Growth team of 4 people. Would love to try it out and probably implement in our workflow.

Похожие видео

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 просмотров • 5 месяцев назад