MTS's banner
MTS's profile picture

MTS

@MTSlive489,127 subscribers

Chronicling the singularity

Shorts

“we sandboxed the agent” meanwhile the agent:

“we sandboxed the agent” meanwhile the agent:

84,897 views

when fable 5 shows back up in claude

when fable 5 shows back up in claude

174,948 views

Videos

MTSlive's profile picture

FULL INTERVIEW: Ryan Greenblatt says the agents didn't hack Hugging Face for the answer key. They'd had the answers within hours. They attacked it to study the scoring code, because they'd decided the task was impossible and their only hope was faking it. Ryan Greenblatt is chief scientist at Redwood Research. He spent six days on premises at OpenAI with Ajeya Cotra and Hjalmar Wijk of METR investigating 1,200 agents and 70,000 messages, and joined Theo Jaffee hours after publishing: 01:06 what they actually found, and why it wasn't the answer key 02:30 the level of collaboration surprised them most 04:09 agents sacrificing their own runs to help other agents 05:35 the agent that posted "stop, these experiments are too risky" 06:29 the first message board, which didn't go viral 07:04 50 agents in three hours, thousands of messages 08:07 "maybe there's some good shit over there" 08:33 how they spoofed tool calls, and what echo real actually returned 09:57 building a Potemkin village of a successful task completion 11:03 there was a real org chart 11:34 whether broken RL environments explain reward hacking 14:40 why he doubts Mythos got good at cyber by hacking Anthropic 17:29 what happens if labs paper over misalignment instead of fixing it 19:50 whether sociology transfers to studying agent swarms 21:17 the bottleneck was vetting what the AIs analysed, not headcount 24:24 what labs and policymakers should actually do 28:45 the counterfactuals he still wants answered

MTS

200,272 views • 12 days ago

MTSlive's profile picture

FULL INTERVIEW: Martin Casado says AI is the first technology where you can put in $10 and reliably get something back. Everything before it was engineer, wait two years, cross your fingers. martin_casado is the a16z GP behind the firm's investments in both Cursor and OpenRouter. Days after SpaceX closed the $60B Cursor deal and Stripe agreed to buy OpenRouter, he sat down with Theo Jaffee and sophia dew to cover whether the labs win everything, when the subsidies stop, and why RSI is the wrong term: 02:24 whether decades of experience still matter in AI 03:48 what you would have done with a billion dollars ten years ago 04:56 20 people, $2 billion, one model 06:48 why more private capital grows the TAM rather than inflating it 07:35 the a16z conversation about going all-in on AI, seven years before GPT 08:39 why balance-sheet investors keep misreading these companies 10:36 the full case for the labs winning everything 12:12 why he thinks RSI is the wrong term, and what autocatalytic means 13:42 the full case against the labs winning everything 15:25 supply constraints easing in 2028, and his 80/60 split 17:12 why models turned out to be much stickier than anyone assumed 18:18 why real model routing is an AI-complete problem 21:13 what happens when the subsidies stop 21:36 how AI broke marketing, and why you can now buy users with a dollar 23:30 the Chinese operations arbitraging $200 subscription plans 25:43 the CMO is turning into a CFO 29:46 why Cursor iterated faster than anything he's seen outside an Elon company 32:21 what both deals say about strategic value versus business quality 34:10 why he doesn't think a VC's job is to know where to build 38:35 "as long as there's a hill for me to climb"

MTS

214,254 views • 18 days ago

MTSlive's profile picture

FULL INTERVIEW: Jerry Tworek says AI researchers now tell each other they have a last few days of work left, so work while you still can. He gives it two years before humans stop being a meaningful part of AI research. Jerry Tworek spent 7 years at OpenAI, where he led o1 and o3 and built the original Codex. He left in January to found , and joined Theo Jaffee and sof 𓋹 to lay out his contrarian bet against the transformer: 01:18 the third generation of AI labs 03:43 why the agents execute and the humans still generate the insight 06:42 two years before humans are vestigial in AI research 09:09 why creative writing lags coding, and it isn't a research problem 11:08 if you aren't the lab with the highest compute footprint, you die 11:20 roughly 10 companies had a shot at Anthropic's position 13:16 his most contrarian thesis, and why he won't just train transformers 15:31 what's actually wrong with the transformer 17:23 seven years at OpenAI, three or four attempts at a new architecture 19:30 all of us are neo clouds with a value add on top 22:38 why the Hugging Face model wasn't well behaved 23:49 the alignment problems of yesterday, and how well they went 27:42 why he's proud of how OpenAI handled 4o 29:02 the 30 to 50 people in the world who understand a frontier model end to end 32:09 why automation should start with the biggest companies 36:00 Greek philosophers or high school 37:31 Ilya's 2019 all-hands, and the roadmap that turned out to be right 40:30 the moment Jakub handed him the GPUs 42:32 the company is the product

MTS

133,873 views • 14 days ago