正在加载视频...

视频加载失败

This guy literally broke down step-by-step how to create frontier lab quality evals: 2:42 - What an eval actually is 3:29 - Offline evals vs prod 5:23 - Why old evals stopped working 7:10 - How to write one 8:29 - Why 100% means you failed 9:03 - Models...

44,671 次观看 • 8 天前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

tylercowen is bullish on AI education — here's why. 00:00 -- Preview 00:24 -- President Carlos Carvalho's AI-generated intro 03:21 -- Cowen reacts to UATX's campus 04:38 -- The AI revolution is here. Who will lose the most? 06:05 -- AI lawyers 07:17 -- Don't underestimate this 10:41 -- Changes to the "upper upper middle class" 12:38 -- How to be successful 13:43 -- The rise of managerial empires 14:02 -- When will we have the first billion dollar company with one employee? 16:05 -- 10-20 year forecast 16:19 -- Why education is so behind 17:01 -- Should you be bullish on UATX? 18:36 -- Should you still read Homer? 21:50 -- Write to think 25:01 -- Meet more people 25:42 -- How to get hired 26:54 -- Is AI your best mentor? 38:17 -- How to curb cheating 39:02 -- The new life of the mind 42:34 -- Q&A: Will there be more status associated with real education or AI education? 45:50 -- Q&A: Why do tech-savvy students need to practice using AI? 47:56 -- Q&A: Do LLMs atrophy your mind? 49:29 -- Q&A: How do you avoid AI-dependency? 51:05 -- Q&A: Isn't this vision lonely and isolating? 53:06 -- Q&A: Do students need teachers? 55:36 -- Q&A: What are the four most important courses for undergrads? 57:49 -- Q&A: Which AI company will win the AI race in the next five years and why? 59:22 -- Q&A: Can AI teach religion? 01:01:32 -- Q&A: Will AI narrow or widen our world? 01:04:37 -- Q&A: What makes us human? 01:05:42 -- Q&A: What is art? 01:08:33 -- Q&A: It's easy to catch cheaters

University of Austin (UATX)

27,770 次观看 • 6 个月前

Why AI Can Now Make Discoveries - my conversation with Dan Roberts, Lead of the Foundations of Reinforcement Learning team at OpenAI 00:00 Intro: AI's wild week in mathematics 01:21 What OpenAI's Foundations of RL team does 03:08 Dan's journey: from black holes and quantum gravity to frontier AI 07:04 Are AI systems becoming useful for real science 08:21 The AI math moment: Erdős, OpenAI, DeepMind, and Anthropic 08:52 Why the OpenAI result was an act of exploration 10:25 OpenAI vs. DeepMind: informal reasoning vs. formal proof 12:13 RL 101: learning by doing, not just watching 15:10 Why reinforcement learning works 15:58 How RL breaks: sparse feedback and long-horizon tasks 17:03 RLHF: how human feedback shaped early language models 18:48 Move 37, self-play, and the search for novel strategies 22:16 Explore vs. exploit in scientific discovery 24:49 Why RL may now be "the cake," not the cherry on top 25:46 Why RL started working with large language models 27:29 Is RL "sucking supervision through a straw"? 28:47 Why language may be the grounding layer for intelligence 31:46 A contrarian take on the Bitter Lesson 32:41 What test-time compute actually is 34:50 How RL gives models the ability to think 35:40 Verifiable rewards, math, coding, and the messy real world 38:00 What physics can teach us about AI 42:08 Is there a thermodynamics of AI? 43:08 From Erdős problems to Einstein-level AI 45:16 Is AI already doing original science? 45:51 How far are we from AI automating AI research 47:41 Why Dan is excited about the future of science

Matt Turck

66,559 次观看 • 2 个月前

I'm often asked for the best public example of AI evals done right for a real, production product. I finally have an answer. Teresa Torres shares how she shipped an AI interview coach, and used evals to rapidly squash bugs and improve the product. Teresa shows how she: 1. did error analysis FIRST to find real issues (instead of using generic metrics) 😍 2. used Jupyter notebooks to analyze errors 3. built custom annotation tools + custom widgets in notebooks 4. built a LLM-judge and assertions to test for specific errors 5. iterated through this feedback loop until it worked. 6. kept things simple the whole time It's also probably the best commercial for Jupyter notebooks you can imagine. 🥰 Chapter summary below. Link to YT in next thread 00:00:00 - Intro 00:01:45 - The Product: Building an AI Interview Coach 00:06:34 - The Problem: How Do I Know if My AI Coach is Any Good? 00:10:15 - Using Airtable for Traces and Annotation 00:12:15 - Discovering Jupyter Notebooks and Designing the First Evals 00:15:15 - Example Evals: LLM-as-Judge vs. Code-Based Assertions 00:21:00 - Learning Python with ChatGPT to Analyze Eval Results 00:31:00 - VS Code, Custom Tools, and an Eval Investigation Notebook 00:39:45 - Building a Custom Annotation Tool with Claude 00:41:00 - From Personal Project to Production App 00:46:02 - How Should PMs and Engineers Collaborate on AI Products? 00:55:45 - Q&A: Capturing Feedback and Annotations from End Users 00:58:11 - Q&A: Is a Technical Background Necessary to Build AI? 01:02:28 - Q&A: What's Next for Teresa? 01:03:13 - Q&A: Unpacking the Micro-Decisions of Building an AI App

Hamel Husain

51,376 次观看 • 11 个月前

E159: Hyperliquid: Housing all of Finance jeff.hl came back on the When Shift Happens Podcast to talk about the Hyperliquid journey since the TGE and what the future holds for one of the most loved and prolific protocols in the space Hyperliquid Timestamps 0:00 Intro 2:01 Singapore 2:27 Reminiscing on the Token Launch 5:00 Was This Scale Of Wealth Expected? 6:28 Doing The Right Thing In Crypto 9:07 The Responsibility that comes with Billions of $ 11:10 Jupiter KAST 11:51 Bringing Hyperliquid to the masses 15:21 Pre TGE and Post TGE: Operational difference 20:13 Choices on what to build Internally vs Externally 22:05 How to build a reliable team 24:51 Did the Team celebrate the HYPE wealth Generation event? 26:45 How to test talents for High Integrity 28:31 How much does the Hyperliquid team sleep? 30:05 Employee Vesting Fears 31:41 Dealing with FUD 32:28 How Does Jeff Personally Handle FUD 35:02 Token "Buybacks" critics 37:20 Why Hyperliquid can't have Discretionary "Buybacks" 39:04 HyperEVM, explained Simply 40:00 Paradex Zodl 40:41 HyperEVM: Success so Far? 44:05 HIP-3, explained Simply 47:44 What makes Hyperliquid's approach different 48:19 Why Should People Care? 51:33 Bring All Finance On Chain 52:08 Why Is The Hyperliquid Approach Better? 53:47 Key Numbers showing that Hyperliquid Is Doing it right 59:01 What Has the Unit team demonstrated with spot trading on Hyperliquid in 2025 1:03:29 HIP-4: Outcome Markets 1:08:01 Trezor Sui 1:08:58 What does "Housing All Of Finance" mean? 1:10:51 Why Hyperliquid is not a crypto company 1:12:23 Why Does Hyperliquid have A Stablecoin USDH (Native Markets) 1:14:39 What Is Kinetiq & Why Does It Matter? 1:16:15 Why Is What HyperLend Is Building Important For HyperLiquid 1:23:39 Where did Fairness cost the most? 1:24:47 What should Hyperliquid be Remembered for? 1:25:24 Why should people stay in Crypto when there's an AI brain drain? 1:28:10 Closing Thoughts

MR SHIFT 🦁

577,829 次观看 • 5 个月前

Inside Nemotron and NVIDIA's AI lab: my conversation with Bryan Catanzaro (Bryan Catanzaro). NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? We go deep into the Nemotron models, what it takes to build a top AI lab, and the future of frontier AI. 01:33 - Is open source AI catching the frontier? 05:29 - Do closed labs blocking distillation slow open source down? 07:42 - Is the US falling behind China? 10:30 - Why companies actually choose open models 12:39 - A "crazy" 2008 bet: machine learning on GPUs 15:33 - Working with Andrew Ng and Dario Amodei at Baidu 17:41 - Coming back to NVIDIA: DLSS and the birth of Megatron 21:55 - The real reason NVIDIA builds its own models 24:28 - Is Moore's Law really dead? 33:37 - The Nemotron family: Nano, Super, Ultra 35:09 - Built for agents: why NVIDIA bets on speed 36:02 - How you train a 550B model in 4 bits 39:25 - Hybrid Mamba-Transformer, explained simply 42:31 - Mixture of experts, and why NVIDIA built NVL72 around it 47:26 - Why a 1-million-token context window matters 49:26 - Multi-token prediction: how the model predicts 5 tokens at once 52:47 - Multi-teacher distillation: teaching one model from many 58:01 - Where reinforcement learning goes next 01:00:16 - Inside NVIDIA's research org: "the mission is the boss" 01:04:03 - How NVIDIA decides who gets the GPUs 01:10:53 - Why NVIDIA still feels entrepreneurial after 33 years 01:12:58 - Why Bryan doesn't believe in the singularity 01:17:50 - The AI backlash 01:19:18 - The controversial case: open AI is safer than closed

Matt Turck

56,728 次观看 • 1 个月前

.Ben Shapiro at UATX with Joe Lonsdale and Niall Ferguson. 00:00 — UATX President Carlos Carvalho 03:16 — Niall Ferguson introduces Ben Shapiro 07:26 — Why UATX is important 08:09 — Why Americans hate each other 10:12 — Emotivism 11:57 — The death of politics 12:30 — Hannibal Lecter skin suits 12:53 — Conspiracy theories 14:11 — Why people don't go to church anymore 15:04 — Vaccines 16:02 — Social engineering & weak professors 17:08 — The last time Harvard meant veritas 17:49 — Epistemic humility 18:30 — Read the Federalist Papers 19:56 — War of all against all 21:03 — Capitalism & soul sickness 22:42 — Tribalism 23:38 — JS Mill and debate culture 26:09 — How to restore our institutions 26:40 — Why UATX matters 27:52 — Joe Lonsdale interviews Ben Shapiro 28:10 — What is a college degree worth? 29:46 — What Jews should learn from Christians and vice versa 32:21 — Candace Owens 34:24 — Audience Q&A: Abraham Lincoln & the Declaration of Independence 36:10 — Q&A: Constitutional boundaries 38:31 — Q&A: Tucker Carlson 42:09 — Q&A: American ingratitude 45:02 — Q&A: How to unite our country 48:24 — Q&A: How to repair our institutions 51:41 — Q&A: The three most important words in the English language 54:25 — Q&A: The future of populism 57:08 — Q&A: Lizard brains 59:43 — Q&A: How Israeli politics work 01:04:03 — Q&A: How to strengthen America 01:07:29 — Q&A: Practical advice for students 01:09:07 — Standing ovation for Ben Shapiro 01:09:36 — President Carlos Carvalho's speech 01:11:08 — How to raise lions Recorded: Sunday, April 26.

University of Austin (UATX)

36,540 次观看 • 3 个月前

DROPS E38: Vanta Trading - The Best Traders Won't Be Human Arrash is the founder and CEO of Vanta Trading, a decentralized prop trading platform built on Bittensor. He spent years as a quant trader building his own strategies before deciding the entire funded-account industry needed to be rebuilt from the ground up. We talk about: - Why most "funded accounts" trade on money that doesn't exist - Why your payout was never real - How prop firms intentionally change rules and spreads before you cash out - Why the best traders of the future won't be human And much more… Timestamps: 0:00 Introduction 1:48 Founder of Vanta 3:08 Explaining Vanta to an Uber Driver 3:26 Unfair vs Fair Funding 6:37 How do they make money? 8:33 How a Legit Prop Firm Makes Money 9:43 Founder's Journey Into Entrepreneurship 10:42 Discovering Crypto & Blockchain 13:03 From LinkedIn Engineer to Quant Trader 15:22 Trading Strategies 16:13 How Trading Is Changing? 18:39 Why TradFi Should Fear Hyperliquid 19:16 Building on BitTensor 21:43 Explaining BitTensor Simply 22:03 How BitTensor Creates Value 23:17 BitTensor's Structure 24:36 BitTensor as Crypto AI 25:35 Is the BitTensor Hype Justified? 27:09 Revenue & Profitability in BitTensor 29:17 Role of TAO 29:51 Advantages & Limitations of BitTensor ecosystem 31:56 What Is Vanta? 32:26 Why No One Fixed Prop Trading Before 33:12 Is the Entire Industry a Scam? 34:17 How Prop Firms Really Make Money 35:35 How Vanta Is Different 38:33 Copy Trading Explained 39:30 Vanta's Business Model 40:28 Dark Reality Behind Funded Accounts 43:48 Long-Term Vision for Vanta 44:50 What Happens If Too Many Traders Win? 46:25 Future Belongs to AI Traders 47:59 Vanta's Endgame 49:14 Conclusion

MR SHIFT 🦁

37,896 次观看 • 1 个月前

I don't think most PMs realize the PRD is becoming obsolete. For the last decade, the PM's core artifact was a qualitative spec. Clear requirements, user stories, acceptance criteria. The engineering team interpreted it, built something close, and the PM spent two weeks reconciling what shipped with what they wrote. The best AI companies replaced that entire loop with evals. A set of inputs your product needs to handle. A task that generates outputs. A scoring function that produces a number between 0 and 1. No ambiguity. No interpretation gap. Ankur Goyal built the eval platform behind Vercel, Replit, Ramp, Notion, and Airtable. An $800M company. He walked through building an eval from zero on this episode and the score went from 0 to 0.75 in under 20 minutes. That's a PM shipping a measurable quality bar before a single line of product code exists. Here's the part that changes the PM role permanently. When the product passes the eval and users still hate it, the eval is wrong. That's on the PM. Evals make PM judgment quantifiable in a way PRDs never did. You can't hide behind "the spec was ambiguous." There's a number now. Six months ago, PM interviews asked "how do you use AI in your workflow." The next wave of interviews is going to ask you to write an eval. The PMs who can encode user intent as a scoring function are building the one skill that survives every model change, every framework swap, every agent rewrite. Write the eval.

Aakash Gupta

78,257 次观看 • 4 个月前