正在加载视频...

视频加载失败

capability != learning new benchtalks with Parth Asawa on continual learning, where we discuss teaching models to learn from experience, measuring learning ability, the bet on parametric models, and more 01:06 What is continual learning? 04:10 Why capability and learning are different 06:13 Why build a benchmark? 08:07 Continual...

41,444 次观看 • 3 个月前 •via X (Twitter)

3 条评论

vincent sunn chen 的头像
vincent sunn chen3 个月前

Kudos to the full Continual Learning Bench team for the work! + @chris_m_glaze @Gorlanski Benji Xu @RamyaRamakri @_asimbiswal @fredsala @matei_zaharia @profjoeyg

vincent sunn chen 的头像
vincent sunn chen3 个月前

Youtube here:

James Alcorn 的头像
James Alcorn3 个月前

@pgasawa @pgasawa a lad destined for greatness

相关视频

Why AI Can Now Make Discoveries - my conversation with Dan Roberts, Lead of the Foundations of Reinforcement Learning team at OpenAI 00:00 Intro: AI's wild week in mathematics 01:21 What OpenAI's Foundations of RL team does 03:08 Dan's journey: from black holes and quantum gravity to frontier AI 07:04 Are AI systems becoming useful for real science 08:21 The AI math moment: Erdős, OpenAI, DeepMind, and Anthropic 08:52 Why the OpenAI result was an act of exploration 10:25 OpenAI vs. DeepMind: informal reasoning vs. formal proof 12:13 RL 101: learning by doing, not just watching 15:10 Why reinforcement learning works 15:58 How RL breaks: sparse feedback and long-horizon tasks 17:03 RLHF: how human feedback shaped early language models 18:48 Move 37, self-play, and the search for novel strategies 22:16 Explore vs. exploit in scientific discovery 24:49 Why RL may now be "the cake," not the cherry on top 25:46 Why RL started working with large language models 27:29 Is RL "sucking supervision through a straw"? 28:47 Why language may be the grounding layer for intelligence 31:46 A contrarian take on the Bitter Lesson 32:41 What test-time compute actually is 34:50 How RL gives models the ability to think 35:40 Verifiable rewards, math, coding, and the messy real world 38:00 What physics can teach us about AI 42:08 Is there a thermodynamics of AI? 43:08 From Erdős problems to Einstein-level AI 45:16 Is AI already doing original science? 45:51 How far are we from AI automating AI research 47:41 Why Dan is excited about the future of science

Matt Turck

69,801 次观看 • 3 个月前

Today we release my favorite episode of Training Data yet: the great Rich Sutton. Richard Sutton wrote the textbook, wrote The Bitter Lesson (and many other on-point essays like "Self-Verification, The Key to AI"), and trained a mafia of talented students who went on to change the AI landscape forever including David Silver, inventor of built AlphaGo. Khurram Javed was Rich's PhD student at Alberta and wrote The Big World Hypothesis. They just left academia to start Oak Lab Their core argument: (1) The Bitter Lesson: the world is massively more complex than any model of it, so anything trained on human-curated data has a ceiling (2) Continual Learning: intelligence is continual by definition, and today's models stop learning the moment they ship. The conversation covers: — what The Bitter Lesson actually says, and what people get wrong — why synthetic data is "just a big mistake," and the Big World Hypothesis behind it — how LLMs are both a positive and a negative example of his own essay — why no animal learns by supervised learning, and what squirrels can do that we can't — the cure for catastrophic forgetting: per-weight step sizes and continual backprop — why the biggest labs can't take a path where performance gets worse before it gets better — a trillion parameters on 20 watts, and the Moore's Law math that makes it plausible — why the endpoint isn't one mind but one design, running as many minds It was both a fun generative idea- and debate-filled conversation, and a surprisingly human one too. Rich, thank you for beating cancer and changing the trajectory of AI. 💙 00:00 Introduction 02:10 An AI winter, a cancer diagnosis, and the move to Alberta 07:07 Writing "The Bitter Lesson," and what people get wrong 09:53 Are LLMs a positive or a negative example of it? 11:03 Synthetic data is "just a big mistake," and the Big World Hypothesis 18:01 AlphaGo, human priors, and why prior knowledge and learning should be friends 22:37 "Their weights never change": do LLM assistants actually learn? 26:09 Babies, squirrels, and why no animal learns by supervised learning 32:02 Rockets, imagination, and where paradigm shifts come from 36:42 The Alberta Plan and its 12 steps 38:53 Catastrophic forgetting and the cure 43:43 Oak's biggest ambition: a self-maintaining mind 47:56 Why the big labs are stuck in a local minimum 49:13 If everything goes right: LLMs, many minds, and hiring The man who pioneered reinforcement learning thinks the rest of the field is weird, and lays it all out in today's episode. Together w/ Alfred Lin Sequoia Capital

Sonya Huang 🐥

119,665 次观看 • 1 个月前