Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Built something fun this week: behavior-judge Long-horizon agents are hard to evaluate, but observing the process agents take to get to outcomes can help `behavior-judge` helps eval long-horizon agents by compiling natural-language specs into deterministic checks for behaviors

14,387 Aufrufe • vor 13 Tagen •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with Mitchell Troyanovsky, co-founder of Basis 01:09 Why Everyone at Basis Was Whispering to AI when Stephanie Palazzolo walked in 04:12 Accounting as "an Intelligence Over the Economy" 06:11 What Makes an Agent Truly Long-Horizon 08:24 Inside an Autonomous, Multi-Day Tax Return 10:19 Agents That Hand Off Like Senior Engineers 11:17 A Brief History of Agents: From ReAct to Today 12:33 Why LLMs Have No Long-Term Memory 14:13 Why AutoGPT Didn't Live Up to Its Promise 15:51 The Three Breakthroughs: Opus 3, o1, o3 17:07 Why Reasoning Models Unlocked Agents 18:23 "Let's Verify Step by Step": The Road Not Taken 20:32 Pushing Back on the METR Chart 22:09 Why Coding Agents Won First 25:14 Why Real-World Agents Are Harder 26:55 How Accountants Verify Non-Deterministic Work 29:18 You Can't Scale Tax Returns Like Math 33:16 100 Evals Pass - So What? 35:53 Right Answer, Wrong Process 36:37 Behavior Specs, Explained 39:58 How Specific Should Behaviors Be? 42:18 Context Is Runtime Training Data 44:21 Who Judges the Judge? 46:45 The Move 37 Objection 50:02 The Magic Box Mental Model 52:41 "Nothing Has Changed Since o3" 54:56 Open-Sourcing Behavior Specs with Ankur Goyal Braintrust 59:45 Ontologies: A World for Agents to Live In 01:04:20 Documentation as Codebase 01:06:33 Why the Founding Fathers Were Context Engineers 01:09:05 Onboarding 300 Brilliant Alien Employees 01:11:10 Self-Improving Agent Systems 01:12:50 The Context Mistake Agent Builders Make 01:14:29 RL on Behavior Adherence 01:17:01 Will the Bitter Lesson Swallow the Harness 01:18:46 "Technical Moats Are Not Real Moats" 01:21:03 Advice for AI Builders

Matt Turck

20,898 Aufrufe • vor 20 Tagen