正在加载视频...
视频加载失败
We’re Scaled Cognition, developing the first ever models trained specifically for agentic applications: 1. Our first system, APT-1, is now #1 on agentic benchmarks. 2. It was developed by a US team for a total cost of less than $11M. 3. Khosla Ventures led our seed round ($21M closed... show more
94,532 次观看 • 1 年前 •via X (Twitter)
9 条评论

We’re also announcing our Agent Builder platform, which allows you to create, test, and deploy an enterprise-grade AI agent using APT-1 in under an hour. Our GenAPI technology lets you test agent behaviors without needing to integrate with real APIs during development. Learn more about Agent Builder and GenAPI:

APT-1 currently outperforms all other models on the Tau-Bench and ComplexFuncBench agentic leaderboards, which test the ability to invoke sequences of complex APIs and comply with business policies.

We’ve accomplished these gains through (1) optimizing the model for actions rather than tokens, (2) a new kind of synthetic agentic training data, and (3) a novel RL approach using agent-to-agent self play. (1) Standard models are focused on token sequences, but business logic applies to actions (like policy rules governing API calls). Our models are trained to optimize licensed, well-formed actions, producing training regimes better suited to agentic tasks. (2) Training an agentic system requires data that contains both conversations and the actions that go along with them. Neither the web nor enterprise datastores have this kind of grounded data, so we generated it using a new, fully synthetic data pipeline. (3) RL through self-play has been a powerful tool in AI for domains like Chess or Go where win/loss outcomes are clear. We developed solutions for getting self-play to work in the agentic case, using simulated agent-to-agent interactions. Our approach applies to any agentic application and teaches the system to take actions correctly, subject to policies and instructions. Learn more and register for early access:

Very cool! Did you also compute the pass^k curves? Curious to see how reliable it is over multiple runs

Not sure if I understand what's happening here. Is the last action-token the model outputted in this screenshot “generate response”?

Congrats! Open source?

Impressive breakthrough in agentic AI! 🚀 APT-1 leading benchmarks with a fully synthetic RL-based pipeline is a game-changer. Excited to see its impact! 🔥 For PhD research support in AI/ML, check out @PhDPRIMA! 🎓✨ #PhDPRIMA

Yummy

which AI video recording did he use?
