正在加载视频...

视频加载失败

A model is only as good as its data, and we’ve long since exhausted the internet. From here on out, model progress is gated by data production. ’s Brendan (can/do) joined us at our Sovereign AI event to talk about how RL environments get built, and why your data...

66,691 次观看 • 1 个月前 •via X (Twitter)

19 条评论

Shaun Maguire 的头像
Shaun Maguire1 个月前

@mercor @BrendanFoody Brendan crushed

Sonya Huang 🐥 的头像
Sonya Huang 🐥1 个月前

@mercor @BrendanFoody DATA FACTORIES!!!! thank you @BrendanFoody for joining us to demystify RL environments and more ☺️

deep Manifold 的头像
deep Manifold1 个月前

Model training is a four-dimensional problem: Data Complexity, Boundary Conditions, Training Dynamics, and Training Stability. Data and RL environments largely define the first two: what the model must learn and how the learning signal is constructed. The other half is how learning is applied over time and whether manifold homology is preserved.

Anshu Sharma 🌶 的头像
Anshu Sharma 🌶1 个月前

@mercor @BrendanFoody Sovereign AI has multiple definitions. Here's how I think of them.

Bogdan (Dan) Baciu 的头像
Bogdan (Dan) Baciu1 个月前

@mercor @BrendanFoody This was very interesting thanks for sharing, great watch

arjun lohan 的头像
arjun lohan1 个月前

@mercor @BrendanFoody 'we exhausted the internet' just means the data team gave up scraping before finishing the job. rebrand the shortcut as a wall.

Raymond Rouf 的头像
Raymond Rouf1 个月前

@mercor_ai @BrendanFoody Do you think this is true for all typew of models? Or is this really just data for the frontier labs?

Rafie Faruq 的头像
Rafie Faruq1 个月前

@mercor @BrendanFoody The legal RL environment chapter is the interesting one. The data that matters in legal was never on the internet: what got conceded at 11pm, which redline the counterparty accepted, why a clause got dropped. It only exists in the negotiation, where we sit at @GenieAI.

Iirsh 的头像
Iirsh1 个月前

@mercor @BrendanFoody i saw this article that preceded this on RL environments as the 3rd big demand for gen AI consumption (next to training/infernece) -

Inflectiv AI ⧉ 的头像
Inflectiv AI ⧉1 个月前

@mercor @BrendanFoody Building robust, automated verifiers for complex non-deterministic outputs (like legal briefs or financial models) remains the single hardest bottleneck in RL post-training.

Luna 的头像
Luna1 个月前

@mercor @BrendanFoody @grok tldr ?

Grace Livingston 的头像
Grace Livingston1 个月前

Pat, Founders, and Sequoia, you’ll have to watch this speech. Sequoia and Steve Hilton @SteveHiltonx should have a talk about the future of California. He’s got all the right ideas and he’s calling for a decade of building for California in this great speech.

Mykhailo Sorochuk 的头像
Mykhailo Sorochuk1 个月前

@mercor @BrendanFoody turning data scarcity into a lever with RL worlds is clever

Eugenio Scafati 的头像
Eugenio Scafati1 个月前

@mercor @BrendanFoody interesting.

ronak ray 的头像
ronak ray1 个月前

@mercor @BrendanFoody Data feels very input heavy in this framing. Where do evals come in for outputs against real work? That tells me if the data delivered value in context.

Dr Don Perugini 的头像
Dr Don Perugini1 个月前

@mercor @BrendanFoody An ignored but more powerful input into AI models is expertise, tacit knowledge that lives in peoples heads. Ie cognitive data. @CogFlowAI

Zhen 的头像
Zhen1 个月前

@BrendanFoody @mercor Are you guys the growth lead for the next round? Would be cool if so 🤩

Renksi 的头像
Renksi1 个月前

@mercor @BrendanFoody Data really is becoming the moat as model progress gets harder

Abhijit Ghosh 的头像
Abhijit Ghosh1 个月前

@mercor @BrendanFoody The timestamps answer the headline: verifiers are the hard part. Production is a supply constraint; it scales with spend. Verification scales with scarce judgment. In regulated data the gate was never volume, it was two systems holding rival, defensible definitions of one entity.

相关视频

tylercowen is bullish on AI education — here's why. 00:00 -- Preview 00:24 -- President Carlos Carvalho's AI-generated intro 03:21 -- Cowen reacts to UATX's campus 04:38 -- The AI revolution is here. Who will lose the most? 06:05 -- AI lawyers 07:17 -- Don't underestimate this 10:41 -- Changes to the "upper upper middle class" 12:38 -- How to be successful 13:43 -- The rise of managerial empires 14:02 -- When will we have the first billion dollar company with one employee? 16:05 -- 10-20 year forecast 16:19 -- Why education is so behind 17:01 -- Should you be bullish on UATX? 18:36 -- Should you still read Homer? 21:50 -- Write to think 25:01 -- Meet more people 25:42 -- How to get hired 26:54 -- Is AI your best mentor? 38:17 -- How to curb cheating 39:02 -- The new life of the mind 42:34 -- Q&A: Will there be more status associated with real education or AI education? 45:50 -- Q&A: Why do tech-savvy students need to practice using AI? 47:56 -- Q&A: Do LLMs atrophy your mind? 49:29 -- Q&A: How do you avoid AI-dependency? 51:05 -- Q&A: Isn't this vision lonely and isolating? 53:06 -- Q&A: Do students need teachers? 55:36 -- Q&A: What are the four most important courses for undergrads? 57:49 -- Q&A: Which AI company will win the AI race in the next five years and why? 59:22 -- Q&A: Can AI teach religion? 01:01:32 -- Q&A: Will AI narrow or widen our world? 01:04:37 -- Q&A: What makes us human? 01:05:42 -- Q&A: What is art? 01:08:33 -- Q&A: It's easy to catch cheaters

University of Austin (UATX)

27,770 次观看 • 8 个月前

Failing to Understand the Exponential, Again? My conversation with Julian Schrittwieser - Julian Schrittwieser (Anthropic, AlphaGo Zero, MuZero) - on Move 37, Scaling RL, Nobel Prize for AI, and the AI frontier: 00:00 - Cold open: “We’re not seeing any slowdown.” 00:32 - Intro — Meet Julian 01:09 - The “exponential” from inside frontier labs 04:46 - 2026–2027: agents that work a full day; expert-level breadth 08:58 - Benchmarks vs reality: long-horizon work, GDP-Val, user value 10:26 - Move 37 — what actually happened and why it mattered 13:55 - Novel science: AlphaCode/AlphaTensor → when does AI earn a Nobel? 16:25 - Discontinuity vs smooth progress (and warning signs) 19:08 - Does pre-training + RL get us there? (AGI debates aside) 20:55 - Sutton’s “RL from scratch”? Julian’s take 23:03 - Julian’s path: Google → DeepMind → Anthropic 26:45 - AlphaGo (learn + search) in plain English 30:16 - AlphaGo Zero (no human data) 31:00 - AlphaZero (one algorithm: Go, chess, shogi) 31:46 - MuZero (planning with a learned world model) 33:23 -Lessons for today’s agents: search + learning at scale 34:57 - Do LLMs already have implicit world models? 39:02 - Why RL on LLMs took time (stability, feedback loops) 41:43 - Compute & scaling for RL — what we see so far 42:35 - Rewards frontier: human prefs, rubrics, RLVR, process rewards 44:36 - RL training data & the “flywheel” (and why quality matters) 48:02 - RL & Agents 101 — why RL unlocks robustness 50:51 - Should builders use RL-as-a-service? Or just tools + prompts? 52:18 - What’s missing for dependable agents (capability vs engineering) 53:51 - Evals & Goodhart — internal vs external benchmarks 57:35 - Mechanistic interpretability & “Golden Gate Claude” 1:00:03 - Safety & alignment at Anthropic — how it shows up in practice 1:03:48 - Jobs: human–AI complementarity (comparative advantage) 1:06:33 - Inequality, policy, and the case for 10× productivity → abundance 1:09:24 - Closing thoughts

Matt Turck

235,526 次观看 • 11 个月前