正在加载视频...

视频加载失败

Interpretability research has made only minor contributions to AI safety so far. What can we do to change that? (Clip from a longer talk; YouTube link in the thread):

29,228 次观看 • 10 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

AI models currently have a 50% chance of doing something that takes a human expert one hour. This doubles every 7 months. In 2 years? They could automate full workdays. In 4 years? A full month. I discuss the most important graph in AI today with Beth Barnes, the CEO of METR, which uncovered this rule of AI progress. Her bottom line: "It really doesn't seem like 2 years would be surprising for recursively self-improving AI." Beth also explains: where company safety testing fails, why there are no true closed-weight models, AI undermines leading powers, why she's come around on open weighting, and why models might be about to start playing dumb much more often. Enjoy! Available on the 80,000 Hours Podcast in all apps. Links below. 1:51 Can we see AI scheming in the chain of thought? 12:50 Alignment faking 17:33 We have to test models before they're even used inside AI companies 31:56 Each 7 months models can do tasks twice as long 51:31 METR's research finds AIs are solid at AI research already 58:18 AI may turn out to be strong at novel and creative research 1:07:55 Recursively self-improving AI might even be here in two years 1:14:29 Could evaluations backfire? 1:39:55 Do we need external auditors doing AI safety tests? 1:54:09 Why not work at AI companies 2:08:40 The new more dire situation has forced changes to METR's strategy 2:21:49 Overrated: Interpretability research 2:32:55 Overrated: Major AI companies' contributions to safety research 2:39:15 Could we ban using AI to enhance AI, or is that just naive? 2:45:31 Open-weighting models is often good 2:50:22 What we can learn about AGI from the nuclear arms race 3:10:43 AI is more like bioweapons because it undermines the leading power 3:42:09 What research METR plans to do next

Rob Wiblin

93,669 次观看 • 1 年前