正在加载视频...

视频加载失败

Zuck on: - Llama 3 - open sourcing towards AGI - custom silicon, synthetic data, & energy constraints on scaling - Caeser Augustus, intelligence explosion, bioweapons, $10b models, & much more Enjoy! Links below

870,920 次观看 • 2 年前 •via X (Twitter)

10 条评论

Dwarkesh Patel 的头像
Dwarkesh Patel2 年前

YouTube: Transcript: Apple Podcasts: Spotify:

nic carter 的头像
nic carter2 年前

@JosephJacks_ Thank you for your consistently amazing work Dwarkesh.

Dwarkesh Patel 的头像
Dwarkesh Patel2 年前

@JosephJacks_ ❤️

main 的头像
main2 年前

more excited for this than the actual model

Nemo 的头像
Nemo2 年前

common dwarkesh win

Zggy 的头像
Zggy2 年前

I have to admit, Zuck interviews very well, far better than Sam. He comes across as more consistently candid, and I can see how his optimism can be contagious and refreshing.

Mckay Wrigley 的头像
Mckay Wrigley2 年前

Best podcast in the world.

Chris Gaun 的头像
Chris Gaun2 年前

Grew out the fro and gained new social magnetism. Dude's Samson

Emad 的头像
Emad2 年前

This was a really good interview thank you

Tyler 🏴‍☠️ 的头像
Tyler 🏴‍☠️2 年前

The people champ

相关视频

Ryan Greenblatt is lead author of "Alignment faking in LLMs" and one of AI's most productive researchers. He puts a 25% probability on automating AI research by 2029. We discuss: • Concrete evidence for and against AGI coming soon • The 4 easiest ways for AI to take over • What evidence we have on how fast / long the intelligence explosion will go • Would misaligned AGI go rogue early or bide its time • Whether 'pause at human level' is naive or smart • Lots more. My head was often spinning during this interview, in a good way. Find it on the 80,000 Hours Podcast, links below. Enjoy! 1:29 How close are we to automating AI R&D? 5:15 Really, though: how capable are today's models? 13:01 Why AI companies get automated first 18:10 Most likely ways for AGI to take over 30:04 Would AGI go rogue early or bide its time? 34:53 "Pause at human level" 46:43 AI control vs AI alignment 52:38 Do we have to hope to catch AIs red-handed? 56:57 How would a slow AGI takeoff look? 1:05:04 Why might an intelligence explosion not happen for 8+ years? 1:17:05 Key challenges in forecasting AI progress 1:25:07 The bear case on AGI 1:30:59 The change to "compute at inference" 1:36:38 How much has pretraining petered out? 1:49:08 Could we get an intelligence explosion within a year? 1:53:08 Reasons AIs might struggle to replace humans 2:00:10 Things could go insanely fast when we automate AI R&D. Or not. 2:14:52 How fast would the intelligence explosion slow down? 2:27:53 Bottom line for mortals 2:34:00 Six orders of magnitude of progress... what does that even look like? 2:44:10 Neglected and important technical work people should be doing 2:48:16 What's the most promising work in governance? 2:51:37 Ryan's current research priorities

Rob Wiblin

34,618 次观看 • 1 年前

"Introducing Multimodal Llama 3.2": As promised two weeks ago, here's the short course on Meta's latest open model! This short course is created with Meta and taught by Amit Sangani, Director of AI Partner Engineering at Meta. Meta’s Llama family of models is leading the way in open models, allowing anyone to download, customize, fine-tune, or build new applications on top of them. Learn about the vision capabilities of the Llama 3.2, and use it for image classification, prompting, tokenization, tool-calling. You'll also learn about the open-source Llama stack, which gives building blocks for many different stages of the LLM application life cycle. In detail, you’ll: - Learn what are the features of Meta's four newest models, and when to use which Llama model. - Learn best practices for multimodal prompting, with applications to advanced image reasoning, illustrated by many examples: Understanding errors on a car dashboard, adding up the total of photographed restaurant receipts, grading written math homework. - Use different roles—system, user, assistant, ipython—in the Llama 3.1 and 3.2 models and the prompt format that identifies those roles. - Understand how Llama uses the tiktoken tokenizer, and how it has expanded to a 128k vocabulary size that improves encoding efficiency and multilingual support. - Learn how to prompt Llama to call built-in and custom tools (functions) with examples for web search and solving math equations. - Learn about Llama Stack, a standardized interface for common toolchain components like fine-tuning or synthetic data generation, useful for building agentic applications. By the end of this course, you’ll be equipped to build out new applications with the new Llama 3.2. Thank you to Ahmad Al-Dahle, Amit Sangani, and the whole AI at Meta team AI at Meta for all the hard work on Llama 3.2 — we’re excited to make these open models even more accessible to more developers with this new course! Please sign up here!

Andrew Ng

131,767 次观看 • 1 年前

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existential and strategic topic for our portfolio. Intelligence is the product. Companies want to shape it and own it and let it compound within their own walls. Not your weights, not your product. Now, with frontier open-weight models and fantastic tooling/infrastructure, owning your intelligence at the frontier is finally becoming possible. The result: every application company we work with is embarking on the journey of doing their own research on post-training, evals, harnesses, etc. The hottest neolabs may just be Harvey, Factory, RamPrasad "RamP!" Moudgalya, etc. The list goes on. We held a summit Sequoia Capital to convene our portfolio on this topic, together with Gabe Pereyra (Harvey) on building Harvey Labs, Lin Qiao (Fireworks) on post-training, Harrison Chase (LangChain) on harnesses + evals, Brendan (can/do) () on RL environments and synthetic data, Arjun Karanam (Trajectory) on online continual learning. Opening talk below; rest to come this week! 00:00 What is sovereign AI (and what it isn't) 01:24 Centralized vs. decentralized intelligence 02:54 Four reasons companies own their models: cost, speed, performance, destiny 04:22 "Not your weights, not your product" 05:32 The application companies are the newest neo labs 07:05 Step 1: Deciding what to own vs. rent 09:51 Step 2: Build the team (and don't shoehorn your platform team) 11:17 Step 3: Legibility – why your research has to be visible 12:33 Step 4: The technical roadmap 13:56 The stack: production vs. development 15:16 Opening Pandora's box – base models, harnesses, context

Sonya Huang 🐥

125,893 次观看 • 10 天前

The interview with Demis Hassabis - the tl;dr (summary) about scaling, AGI and much more: 1. Solving the "Root Node" Problems: DeepMind isn't just building chatbots; they are using AI to solve the hardest scientific problems. After the success of AlphaFold, they are now targeting materials science (room-temperature superconductors, better batteries) and even nuclear fusion to unlock unlimited clean energy. 2. The "Jagged Intelligence" Paradox: Current AI models are in a weird spot—they can win gold medals at the International Math Olympiad but still fail at basic logic puzzles. Hassabis calls this "jagged intelligence." The goal isn't just more data, but fixing these inconsistencies to make models reliable across the board. 3. Scaling is Not Dead (But it’s Changing): Despite rumors of hitting a "data wall," Hassabis says we haven't seen a hard limit yet. However, we are seeing diminishing returns. His bet? Getting to AGI will require 50% scaling and 50% architectural innovation. It’s no longer just about making the models bigger; it’s about making them smarter. 4. The Missing Piece: System 2 Thinking: Today's models are passive—they just spit out an answer. To reach AGI, we need systems that can "think" before they speak. This involves planning, reasoning, and double-checking their own work (similar to human "System 2" thinking) rather than just predicting the next word. 5. Rise of World Models: The next big frontier is "World Models" (like their project Genie). AI needs to understand the physics of the world—gravity, object permanence, and cause-and-effect—not just language. This is crucial for building helpful digital agents and robots that can navigate real-life situations. 6. Is the Universe Computable? On a philosophical level, Hassabis believes that everything in the universe might be computable. His life's work is testing the limits of the "Turing Machine." If we can build an AGI that simulates the human mind perfectly, we might finally understand what (if anything) makes human consciousness unique. 7. Bigger than the Industrial Revolution: We need to prepare for a shift that is 10x faster and bigger than the Industrial Revolution. If AI solves energy (fusion) and labor, we might enter a "post-scarcity" world. Hassabis warns that society, economics, and governments need to adapt quickly to ensure these benefits are shared by everyone, not just a few. And since this is the most important aspect, here is the clip about post labor economy:

Chubby♨️

27,473 次观看 • 8 个月前

Here's a demo on a project I've been developing and working on for the past 9 months. Called NightBeacon. Using it now in production, getting released fully this week. Our own internally trained models on our own infrastructure (no third party). Trained on our analysts knowledge and behavior (TP/FPs retrain model to be smarter with context). Handles emails (including tonality), attachments, various malicious filetypes (DLL/exe/svg/lnk/etc). Can send it full evtx exports, packet dumps, zip files, whatever. Universal log handler can parse any log from any source, EDR, SIEM, etc. Deep-Scan / sandbox detonation + shellcode emulation with IOC extraction automatically. Automatic playbook generation, full AI-based recommendations custom to the attack. Synthetic training data layer - meaning when it trains on a specific attack at a customer, generates training data based on the customers data but never has any of the actual data or information about the customer in it. No customer information. For areas its weak at, bubbles up and automatically kicks off research to become smarter on a specific topic. Supports GenAI based rulesets (to improve confidence), over 900+ YARA rules, full MITRE ATT&CK integration. Integrated into our SOAR - enriches data, creates playbooks for analysts, MTTR reduces substantially, false positives reduced, true positive escalations. Not using our MDR service? Can integrate into your EDR or SIEM for automatic enrichment and escalation of attacks. Built to help respond faster. More accurately. Be intelligent based on our analysts intelligence. Stop attackers much much faster. Coming soon.. #BinaryDefense

Dave Kennedy

12,905 次观看 • 5 个月前

Synthetic data will provide the next trillion tokens to fuel our hungry models. I'm excited to announce MimicGen: massively scaling up data pipeline for robot learning! We multiply high-quality human data in simulation with digital twins. Using 50,000 training episodes across 18 tasks, multiple simulators, and even in the real-world! The idea is simple: 1. Humans tele-operate the robot to complete a task. It is extremely high-quality but also very slow and expensive. 2. We create a digital twin of the robot and the scene in high-fidelity, GPU-accelerated simulation. 3. We can now move objects around, replace with new assets, and even change the robot hand - basically augment the training data with procedural generation. 4. Export the successful episodes, and feed that to a neural network! You now have an near-infinite stream of data. One of the key reasons that robotics lags far behind other AI fields is the lack of data: you cannot scrape control signals from the internet. They simply don't exist in-the-wild. MimicGen shows the power of synthetic data and simulation to keep our scaling laws alive. I believe this principle apply beyond robotics. We are quickly exhausting the high-quality, real tokens from the web. Artificial intelligence from artificial data will be the way forward. We are big fans of the OSS community. As usual, we open-source everything, including the generated dataset! - Website: - Paper: - Dataset is hosted on HuggingFace (thanks AK!!): - Code: MimicGen is led by Ajay Mandlekar, deep dive in the thread:

Jim Fan

332,238 次观看 • 2 年前

AI models currently have a 50% chance of doing something that takes a human expert one hour. This doubles every 7 months. In 2 years? They could automate full workdays. In 4 years? A full month. I discuss the most important graph in AI today with Beth Barnes, the CEO of METR, which uncovered this rule of AI progress. Her bottom line: "It really doesn't seem like 2 years would be surprising for recursively self-improving AI." Beth also explains: where company safety testing fails, why there are no true closed-weight models, AI undermines leading powers, why she's come around on open weighting, and why models might be about to start playing dumb much more often. Enjoy! Available on the 80,000 Hours Podcast in all apps. Links below. 1:51 Can we see AI scheming in the chain of thought? 12:50 Alignment faking 17:33 We have to test models before they're even used inside AI companies 31:56 Each 7 months models can do tasks twice as long 51:31 METR's research finds AIs are solid at AI research already 58:18 AI may turn out to be strong at novel and creative research 1:07:55 Recursively self-improving AI might even be here in two years 1:14:29 Could evaluations backfire? 1:39:55 Do we need external auditors doing AI safety tests? 1:54:09 Why not work at AI companies 2:08:40 The new more dire situation has forced changes to METR's strategy 2:21:49 Overrated: Interpretability research 2:32:55 Overrated: Major AI companies' contributions to safety research 2:39:15 Could we ban using AI to enhance AI, or is that just naive? 2:45:31 Open-weighting models is often good 2:50:22 What we can learn about AGI from the nuclear arms race 3:10:43 AI is more like bioweapons because it undermines the leading power 3:42:09 What research METR plans to do next

Rob Wiblin

93,669 次观看 • 1 年前

Mark Zuckerberg is explaining one of the most misunderstood dynamics in AI and it has direct investment implications (Save this). The concept he's describing is model distillation, and it's one of the most important techniques to emerge in AI over the past year. Here's how it works. You train a massive, enormously expensive model, in Meta's case, Llama 4 Behemoth, a 2 trillion parameter teacher model and then you use that model to teach a much smaller, cheaper model. The smaller model inherits roughly 90 to 95% of the intelligence of the giant while running at 10% of the cost and on a fraction of the compute. Meta already did this with the Llama 4 family and Behemoth serves as the teacher. Llama 4 Scout and Maverick, the publicly released open-source models were distilled from it. Scout runs on a single H100 GPU with a 10 million token context window and outperforms models that cost far more to operate. Maverick, at 17 billion active parameters, rivals DeepSeek V3 in coding at half the parameter count and beats GPT-4o on multimodal benchmarks. Both are completely free for commercial use. What Zuckerberg is pointing at is a structural shift in how AI gets deployed in the real world. Companies aren't taking a frontier model off the shelf and running it as-is but rather taking open-source models, fine-tuning them on their own proprietary data, distilling them into even smaller custom models tailored to their specific use case, and running them on infrastructure they control at a fraction of the cost of a closed frontier API. The investment implication of this is significant and runs in two directions. For Meta specifically, this is a strategic masterstroke. Every company that builds on Llama, fine-tunes it, distills it, or deploys it through their infrastructure is pulling into Meta's orbit while Meta builds the most powerful open teacher model. The ecosystem of companies using it grows and that ecosystem generates commercial activity across Meta's platforms and data services. Meta's AI research benefits from billions of real world deployment signals and it's a flywheel that closed model providers cannot replicate because their strategy requires charging per token, which is now a 65x cost disadvantage against the open-source alternative. For the broader market, distillation changes the economics of inference in a way that has barely been priced in. As intelligence becomes extractable into smaller and cheaper models, the absolute demand for compute doesn't decline but rather it explodes, because now the number of applications that are economically viable expands by orders of magnitude. Every task that was previously too expensive to automate at $3.25 per call becomes viable at $0.05 that means more total token usage, more total GPU utilization, and more demand for the infrastructure companies, the Nebiuses, the GE Vernovas, the Constellation Energies that supply the underlying compute and power.

Milk Road AI

27,908 次观看 • 1 个月前

In just one week, Binh and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

23,236 次观看 • 6 个月前