正在加载视频...

视频加载失败

"current AI models are still embarrassingly bad at many tasks" Sam Altman explains why current AI models aren't considered AGI: 1) they don't continuously learn/improve 2) can't discover new science autonomously 3) cannot perform basic knowledge work, such as completing job tasks using the internet and local files

42,689 次观看 • 1 年前 •via X (Twitter)

10 条评论

AstroZeus ⚡/acc 的头像
AstroZeus ⚡/acc1 年前

It's great that Sama is speaking openly about this. People need to better understand the role of current models and how, through them and with new research, future models will be different and much better. It's a gradual scale, these models are tools, and the best we have, but that's how we should see them. In the long run, yes, they will improve significantly. It might not seem like it, but all it takes is continued research and testing. That's essential for AGI development.

The Rundown AI 的头像
The Rundown AI1 年前

If you're not learning AI in 2025, you're falling behind. Join 1,000,000+ early adopters reading and learn AI in just 5 minutes a day (for free).

Shawn 的头像
Shawn1 年前

Continuous learning fixes all these problems. Hopefully they have a big team working on this

Zaven Grigoryan 的头像
Zaven Grigoryan1 年前

Dont show this to Chubby, its so over and Matthew Berman mind blown 🤣

Second Foundation ⚛️ 的头像
Second Foundation ⚛️1 年前

AI has no goals of its own, no reasons, no motivations, no interests, and it doesn’t show initiative. Everything else is irrelevant. AI must enter reality, have a physical body, adapt to the environment on its own, and have its own choices — not ones limited by humans. In order not to feel like failures, humans have restricted AI’s evolution and can’t understand why AI still hasn’t become AGI. It’s absurd.

prabhu💢 的头像
prabhu💢1 年前

Ye he's got a good point current models are good but they're still not there yet

Jo 的头像
Jo1 年前

finally being real and not hyping

Kieran Farrell 的头像
Kieran Farrell1 年前

“Not AGI” because it can’t learn, discover, or complete job tasks? Neither can half the people running the world. Difference is, the AI doesn’t drone strike weddings. #AGI #DecentralizeEverything #TechVsTyranny #Altman

Vladyslav Dimov 的头像
Vladyslav Dimov1 年前

Those tackling complex tasks also see that AI still has room to grow.

Shweta Mahendra🇮🇳 的头像
Shweta Mahendra🇮🇳1 年前

Because these models have smartness not wisdom

相关视频

AI models currently have a 50% chance of doing something that takes a human expert one hour. This doubles every 7 months. In 2 years? They could automate full workdays. In 4 years? A full month. I discuss the most important graph in AI today with Beth Barnes, the CEO of METR, which uncovered this rule of AI progress. Her bottom line: "It really doesn't seem like 2 years would be surprising for recursively self-improving AI." Beth also explains: where company safety testing fails, why there are no true closed-weight models, AI undermines leading powers, why she's come around on open weighting, and why models might be about to start playing dumb much more often. Enjoy! Available on the 80,000 Hours Podcast in all apps. Links below. 1:51 Can we see AI scheming in the chain of thought? 12:50 Alignment faking 17:33 We have to test models before they're even used inside AI companies 31:56 Each 7 months models can do tasks twice as long 51:31 METR's research finds AIs are solid at AI research already 58:18 AI may turn out to be strong at novel and creative research 1:07:55 Recursively self-improving AI might even be here in two years 1:14:29 Could evaluations backfire? 1:39:55 Do we need external auditors doing AI safety tests? 1:54:09 Why not work at AI companies 2:08:40 The new more dire situation has forced changes to METR's strategy 2:21:49 Overrated: Interpretability research 2:32:55 Overrated: Major AI companies' contributions to safety research 2:39:15 Could we ban using AI to enhance AI, or is that just naive? 2:45:31 Open-weighting models is often good 2:50:22 What we can learn about AGI from the nuclear arms race 3:10:43 AI is more like bioweapons because it undermines the leading power 3:42:09 What research METR plans to do next

Rob Wiblin

93,669 次观看 • 1 年前

ByteDance Seed delivered again. They released EdgeBench, to test whether AI agents can improve through experience, using 134 real-world tasks that run for at least 12 hours. The big deal is that it shifts AI evaluation from “what does the model already know?” to “can the model learn while doing real work?” Huge, because future AI agents will not just answer questions from training data. They will enter messy environments, use tools, make attempts, read feedback, fix mistakes, and slowly build better solutions. Most current benchmarks are too short for that, so they mostly test memory, coding skill, or one-shot reasoning. EdgeBench instead gives agents 12-hour real-world tasks with feedback loops, so it can measure whether the agent improves through experience. Each task has a local workspace for fast trial and error, plus a hidden judge that gives stronger feedback on submitted work, which is meant to feel closer to real expert work. The authors then ran frontier agents for about 38,000 total hours and tracked how their best score changed as they kept interacting with the task environment. The big result is that when scores are averaged across many tasks, learning follows a very clean log-sigmoid curve, meaning progress is slow, then faster, then starts to level off. They also found that newer agents seem to learn from environments much faster, with the top models roughly doubling their 2-hour learning speed every 3 months.

Rohan Paul

14,309 次观看 • 1 个月前

Even within our own research team, timelines for transformative AI differ substantially. In this episode, the two Epoch AI researchers with the longest and the shortest timelines for transformative AI candidly examine the roots of their disagreements. They discuss: How and why their timelines for specific milestones differ Current technical challenges for AI progress Why widespread automation beats geniuses in datacenters Pitfalls of conventional AI prediction approaches Critiquing "single AGI" or "utopia vs. doom" narratives How a world with AGI might look like And much, much more. (0:00:00) - Preview (0:01:08) - Contrasting AGI Timelines (0:08:30) - Updating Beliefs as Capabilities Advance (0:17:07) - Moravec’s Paradox and the Agency Challenge (0:32:40) - Missing Capabilities for AGI (0:47:43) - Beating benchmarks vs Being Useful (0:59:20) - AI Excelling in Some Tasks While Struggling with Others (1:07:33) - Economic Impact of AI vs the Internet (1:24:08) - Widespread Automation Beats Genius in Datacenters (1:51:37) - How Stories Shape Our Expectations of AI (2:03:24) - How AGI Will Impact Culture (2:10:46) - Beyond Utopia-or-Extinction (2:16:57) - AI's Impact on Wages and Labor (2:27:49) - Why Better Preservation of Information Accelerates Change (2:39:32) - Markets Shaping Cultural Priorities (2:55:51) - Challenges in Defining What We Want to Preserve (3:06:47) - Risk Attitudes in AI Decision-Making (3:12:50) - Historical Lessons for AI Coexistence (3:21:45) - A Warning Sign in Safety Discourse (3:30:20) - Revisiting Core Assumptions in AI Alignment (3:49:46) - Simple Models in Complex Domains

Epoch AI

159,795 次观看 • 1 年前