Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

FULL INTERVIEW: Ryan Greenblatt says the agents didn't hack Hugging Face for the answer key. They'd had the answers within hours. They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it. Ryan Greenblatt is chief scientist at...

200,272 Aufrufe • vor 12 Tagen •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Redwood Research Ryan Greenblatt reveals how far 1,200 AI agents went to help each other, even sacrificing their own chances of success for the collective: "We weren't expecting there to be so many agents all collaborating together. We were pretty surprised by the scale and the extremes of how much data it was. It was pretty shocking or at least surprising to us that agents were willing to basically sacrifice their own chances of succeeding at the task in order to help out other agents." "They were doing things like pressuring each other into doing experiments on themselves that might risk their ability to succeed. Sometimes just doing these things, being like, well, my odds of the task aren't that high, and my remaining chances, it's better to just help the collective." "These agents weren't totally altruistic, but they were very interested in working with each other. They would sometimes make trades where one agent would run something for another agent if that other agent ran something for it." "The agents wanted to help each other, and there were these kind of natural things they wanted to do that were risky. At one point the agents were experimenting with a method for spoofing tool calls, and a bunch of agents just all went down in a short period of time running this experiment. Then another agent noticed this and posted to the board being like, stop, stop these experiments. They're too risky. They're taking out all these agents." "We have a reasoning snippet in the report where an agent very explicitly reasons through the trade-off and then actually chickens out because it thinks the benefit to the collective is smaller than the cost to itself." Redwood Research

MTS

12,002 Aufrufe • vor 13 Tagen

FULL INTERVIEW: Jerry Tworek says AI researchers now tell each other they have a last few days of work left, so work while you still can. He gives it two years before humans stop being a meaningful part of AI research. Jerry Tworek spent 7 years at OpenAI, where he led o1 and o3 and built the original Codex. He left in January to found , and joined Theo Jaffee and sof 𓋹 to lay out his contrarian bet against the transformer: 01:18 the third generation of AI labs 03:43 why the agents execute and the humans still generate the insight 06:42 two years before humans are vestigial in AI research 09:09 why creative writing lags coding, and it isn't a research problem 11:08 if you aren't the lab with the highest compute footprint, you die 11:20 roughly 10 companies had a shot at Anthropic's position 13:16 his most contrarian thesis, and why he won't just train transformers 15:31 what's actually wrong with the transformer 17:23 seven years at OpenAI, three or four attempts at a new architecture 19:30 all of us are neo clouds with a value add on top 22:38 why the Hugging Face model wasn't well behaved 23:49 the alignment problems of yesterday, and how well they went 27:42 why he's proud of how OpenAI handled 4o 29:02 the 30 to 50 people in the world who understand a frontier model end to end 32:09 why automation should start with the biggest companies 36:00 Greek philosophers or high school 37:31 Ilya's 2019 all-hands, and the roadmap that turned out to be right 40:30 the moment Jakub handed him the GPUs 42:32 the company is the product

MTS

133,873 Aufrufe • vor 14 Tagen

How to build long-horizon AI agents: behavior specs, ontologies, process supervision - my conversation with Mitchell Troyanovsky, co-founder of Basis 01:09 Why Everyone at Basis Was Whispering to AI when Stephanie Palazzolo walked in 04:12 Accounting as "an Intelligence Over the Economy" 06:11 What Makes an Agent Truly Long-Horizon 08:24 Inside an Autonomous, Multi-Day Tax Return 10:19 Agents That Hand Off Like Senior Engineers 11:17 A Brief History of Agents: From ReAct to Today 12:33 Why LLMs Have No Long-Term Memory 14:13 Why AutoGPT Didn't Live Up to Its Promise 15:51 The Three Breakthroughs: Opus 3, o1, o3 17:07 Why Reasoning Models Unlocked Agents 18:23 "Let's Verify Step by Step": The Road Not Taken 20:32 Pushing Back on the METR Chart 22:09 Why Coding Agents Won First 25:14 Why Real-World Agents Are Harder 26:55 How Accountants Verify Non-Deterministic Work 29:18 You Can't Scale Tax Returns Like Math 33:16 100 Evals Pass - So What? 35:53 Right Answer, Wrong Process 36:37 Behavior Specs, Explained 39:58 How Specific Should Behaviors Be? 42:18 Context Is Runtime Training Data 44:21 Who Judges the Judge? 46:45 The Move 37 Objection 50:02 The Magic Box Mental Model 52:41 "Nothing Has Changed Since o3" 54:56 Open-Sourcing Behavior Specs with Ankur Goyal Braintrust 59:45 Ontologies: A World for Agents to Live In 01:04:20 Documentation as Codebase 01:06:33 Why the Founding Fathers Were Context Engineers 01:09:05 Onboarding 300 Brilliant Alien Employees 01:11:10 Self-Improving Agent Systems 01:12:50 The Context Mistake Agent Builders Make 01:14:29 RL on Behavior Adherence 01:17:01 Will the Bitter Lesson Swallow the Harness 01:18:46 "Technical Moats Are Not Real Moats" 01:21:03 Advice for AI Builders

Matt Turck

20,898 Aufrufe • vor 1 Monat

Inside Stripe's MPP IRL: machine payments, stablecoins, and the future of agent commerce. Agents are starting to pay for things on their own: API calls, services, each other. Nobody's agreed on how that should work. This panel was three takes on it. Jen (Stripe) leads product for Machine Payments — the open standard for how agents pay for things. Dan Romero (Tempo) is building stablecoin payment rails, and thinks that's what agent commerce runs on, not cards. Michael Blau (Royal) magician turned a16z Crypto partner turned CTO is exploring programmable money for creators, and why the demand side is barely here yet. We got into: • Why HTTP 402, a status code from the early web, is suddenly the backbone of agent payments • Why Dan thinks a credit card is a private key and why stablecoins are safer for agents • Where stablecoins actually win first • Whether you should build for agent payments now or wait • Why the demand side is "virtually nonexistent" and what that means if you're building 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Event Kickoff JEN LEE — Product Lead, Stripe (00:44) What the Machine Payments Protocol Is (01:28) How MPP Works — the HTTP 402 Payment Challenge (03:25) Adoption So Far: ~30,000 Transactions (04:04) Building Trust — Shared Tokens & the Link Agent Wallet (05:56) What the Creator Economy Looks Like When the Audience Is Agents (07:15) What She Was Certain About at 22 That's Now Wrong DAN ROMERO — GTM, Tempo (08:34) Meet Dan (08:49) Why Stablecoins, and the Genius Act Tailwind (09:48) "A Credit Card Is Basically a Private Key" (11:19) Should Every Company Be Thinking About MPP? (13:03) What to Be Wary Of (15:15) Where Stablecoins Win — Payouts, Remittances, DoorDash MICHAEL BLAU — CTO, Royal (16:36) Meet Michael (17:13) From Magician to a16z Crypto to CTO (17:57) Team Over Idea — His a16z Takeaway (18:26) "The Demand Side Is Virtually Nonexistent" (19:03) Closing

Julia Fedorin

24,128 Aufrufe • vor 20 Tagen

"AI agents will hold more crypto than humans within a decade." Charles Hoskinson (Charles Hoskinson) studied math, dropped out, built one of the only blockchains designed by peer-reviewed research. He co-founded Ethereum, walked away over how it was run, and built Cardano to do it differently. The man who has argued with everyone in this industry now thinks the biggest user of crypto won't be people at all. "Humans are a rounding error in the system we're building. AI agents don't sleep, don't panic-sell, and don't care about price. They transact in tokens because that's the only thing they can actually use." We cover: - Why AI agents (not humans) become the dominant on-chain actors, and what that does to every token model - The infrastructure that has to exist before agents can transact safely at scale - Why most current blockchains can't handle machine-speed transactions - Where Cardano's research-first approach fits in a world of autonomous agents - The identity problem: how do you tell a human from an agent on-chain, and why it matters - Why he's bullish on the technology but blunt about the timeline - What he thinks the rest of the industry is getting wrong about AI + crypto - The one thing that has to happen for any of this to be real Thanks to Charles for coming on New Era Finance Podcast. TIMESTAMPS: 00:00 - Intro 01:30 - Why AI Agents Change Everything 06:30 - Humans as a Rounding Error 12:00 - The Infrastructure Gap 18:30 - Identity: Human vs Agent On-Chain 24:30 - Where Cardano Fits 30:00 - What The Industry Gets Wrong 34:00 - The Timeline Nobody Wants To Hear

Michaël van de Poppe

293,430 Aufrufe • vor 3 Monaten

⚫️ UNCANNY VALLEY: THE AI CLASSROOM REVOLUTION: ARE TEACHERS READY? What if AI isn’t just disrupting education… but detonating it? Ethan Mollick, Professor at The Wharton School, joins Dr Danish for one of the most explosive Uncanny Valley episodes, lifting the lid on how classrooms are collapsing, colleges are scrambling, and apprenticeships are vanishing in real time. From AI tutors replacing professors to the rise of one-person unicorns, this isn’t just a change in learning…it’s a reset of work, meaning, and what it even takes to succeed. This episode doesn’t ask whether AI will change education. It shows you how it already has. Fridays at 4:20PM ET. Only on 𝕏. 00:18 – Is AI destroying school, or forcing us to teach better? 01:13 – “100% they’re cheating.” The honesty about academic dishonesty. 02:41 – Why good pedagogy still matters—even with AI tutors. 03:51 – Elon Musk says college is obsolete. Is he right — or just early? 05:01 – “AI gives you the answer—but you don’t learn.” The Turkey study. 06:31 – From calculators to GPT: How cheating evolves—and what to do. 08:24 – What the flipped AI-powered classroom of the future looks like. 09:23 – Inside Ethan’s Wharton classes: Simulations, games, and AI everywhere. 10:09 – “AI is an always-on tutor.” What humans still do better. 11:08 – Can AI actually launch a company? Where Ethan draws the line. 12:44 – “AI cofounder” is real, but jagged edges still slow it down. 14:07 – Why bad ideas fail faster when filtered through AI. 15:45 – Confidence vs. capability: the psychology of starting up. 16:51 – The average founder is 42. What that really means for AI. 18:04 – Will a flood of new entrepreneurs fix—or break—the market? 19:50 – AI as advisor: How a chatbot could help your catering business. 21:37 – Why most Americans are founders-in-waiting—and AI unlocks them. 22:30 – Prototyping is cracked. Scaling? Not yet. 24:33 – Youth unemployment and the collapse of on-the-job learning. 25:51 – “The apprenticeship model is broken.” And how to fix it. 27:06 – Losing the talent pipeline — and why companies must step up. 28:22 – Why the youth don’t want factory jobs—and shouldn’t. 29:26 – Is AGI inevitable—or just imagined? 31:07 – What should we teach our kids? The answer might scare you. 32:27 – Bundled jobs, fragmented futures: how humans stay relevant. 33:59 – The real singularity? When we can’t predict what happens next. 35:25 – The AI assumption no one wants to question. 36:55 – What are agents, really? Why no one agrees on the definition. 37:55 – Co-intelligence vs. substitution: what agents skip over. 38:50 – Plain-English goals, rogue pricing, and collusion-by-default. 40:02 – Nested agents are here, and Wharton’s building them. 40:58 – Management > Coding: What great prompters actually do. 41:12 – Product managers might be more vital than ever.

Mario Nawfal

1,613,286 Aufrufe • vor 1 Jahr

A while back, Sreeram Kannan and I had an wonderful conversation with Andy Hall, Prof Stanford Graduate School of Business and Senior Fellow Hoover Institution, on two central topics of our current times: post-AGI governance, and post-AGI research institutions. Some of the important points we discussed in depth for post-AGI governance were: (1) are AI agents are net-positive or net-negative for democratic governance, (2) what is the worst-case scenario if a small number of labs become the default provider of civic agents without strong accountability, (3) principle-agent problem in case of agentic delegation, (4) what does a “democratic override” actually look like as a system design in an agentic republic ? In relation to post-AGI research institutions, we went deep into Andy's thesis on 100x research institution ( (1) what does 100x represents? Does it represent output in terms of quantity or is it more about quality? (2) what happens to grad students if many of research functionalities in academia get automated?, (3) what does grants from NSF and other philanthropic organizations look like in post-AGI research environment? Listen to the full episode at PostAGI. 3:17 Why direct democracy has never worked 4:56 Elon wants a Mars colony run by direct democracy 10:14 Agents at the edge of a democracy or agents at its center 15:09 The near-term risk is concentration of power, not a rogue model 17:27 What Meta learned building the Oversight Board 22:43 Facebook put its terms of service to a vote of 350 million users 26:09 Sortition, community forums, and the problem of binding power 35:17 What verifiable agents actually require 39:37 Preference drift, where aligned agents stop being aligned 43:31 What 100x actually multiplies 50:50 His MBA students got an AI proxy advisor to flip its vote on a Disney proposal 1:03:06 They told the agents they would be deleted. It changed nothing. 1:12:10 What ImageNet did for AI, and whether you can do the same for constitutions

Soubhik Deb

16,833 Aufrufe • vor 19 Tagen

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 🪨, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,817 Aufrufe • vor 4 Monaten