Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

This guy literally broke down step-by-step how to create frontier lab quality evals: 2:42 - What an eval actually is 3:29 - Offline evals vs prod 5:23 - Why old evals stopped working 7:10 - How to write one 8:29 - Why 100% means you failed 9:03 - Models...

44,671 görüntüleme • 1 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

tylercowen is bullish on AI education — here's why. 00:00 -- Preview 00:24 -- President Carlos Carvalho's AI-generated intro 03:21 -- Cowen reacts to UATX's campus 04:38 -- The AI revolution is here. Who will lose the most? 06:05 -- AI lawyers 07:17 -- Don't underestimate this 10:41 -- Changes to the "upper upper middle class" 12:38 -- How to be successful 13:43 -- The rise of managerial empires 14:02 -- When will we have the first billion dollar company with one employee? 16:05 -- 10-20 year forecast 16:19 -- Why education is so behind 17:01 -- Should you be bullish on UATX? 18:36 -- Should you still read Homer? 21:50 -- Write to think 25:01 -- Meet more people 25:42 -- How to get hired 26:54 -- Is AI your best mentor? 38:17 -- How to curb cheating 39:02 -- The new life of the mind 42:34 -- Q&A: Will there be more status associated with real education or AI education? 45:50 -- Q&A: Why do tech-savvy students need to practice using AI? 47:56 -- Q&A: Do LLMs atrophy your mind? 49:29 -- Q&A: How do you avoid AI-dependency? 51:05 -- Q&A: Isn't this vision lonely and isolating? 53:06 -- Q&A: Do students need teachers? 55:36 -- Q&A: What are the four most important courses for undergrads? 57:49 -- Q&A: Which AI company will win the AI race in the next five years and why? 59:22 -- Q&A: Can AI teach religion? 01:01:32 -- Q&A: Will AI narrow or widen our world? 01:04:37 -- Q&A: What makes us human? 01:05:42 -- Q&A: What is art? 01:08:33 -- Q&A: It's easy to catch cheaters

University of Austin (UATX)

27,770 görüntüleme • 8 ay önce

An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. LangChain founder Harrison Chase joined us at our Sequoia Capital Own Your Intelligence to talk about the piece that often gets the least attention: the harness. He offers a clear heuristic for when to build your own. The more out of distribution you are from what the models were trained on, the more you'll want to customize. And good technical content on how to actually measure performance with evals and langsmith. 00:00 Introduction 00:58 The three parts of an agent: harness, model, context 02:12 What a harness actually does 03:25 Customizing the core loop with middleware 04:41 Sandboxes, file systems, sub-agents, summarization 05:47 Cognitive architectures — and when you still need them 07:03 Build your own harness or use off the shelf? 08:24 In-distribution vs. out-of-distribution: the file-editing example 09:39 Why evals define what "good" means in an organization 11:04 Harbor: what an eval task actually looks like 12:11 Comparing harnesses and models on accuracy, latency, and cost 13:20 Why observability is underrated — it's usually the context 14:34 The data flywheel: traces → curation → experiments 15:42 Getting feedback through UX design and online evaluators 16:51 Demo: LangSmith Engine 19:23 Q&A: Running Engine on Engine, and "codex-ification" 20:44 Q&A: Will harnesses converge or diverge?

Sonya Huang 🐥

77,519 görüntüleme • 1 ay önce

Why AI Can Now Make Discoveries - my conversation with Dan Roberts, Lead of the Foundations of Reinforcement Learning team at OpenAI 00:00 Intro: AI's wild week in mathematics 01:21 What OpenAI's Foundations of RL team does 03:08 Dan's journey: from black holes and quantum gravity to frontier AI 07:04 Are AI systems becoming useful for real science 08:21 The AI math moment: Erdős, OpenAI, DeepMind, and Anthropic 08:52 Why the OpenAI result was an act of exploration 10:25 OpenAI vs. DeepMind: informal reasoning vs. formal proof 12:13 RL 101: learning by doing, not just watching 15:10 Why reinforcement learning works 15:58 How RL breaks: sparse feedback and long-horizon tasks 17:03 RLHF: how human feedback shaped early language models 18:48 Move 37, self-play, and the search for novel strategies 22:16 Explore vs. exploit in scientific discovery 24:49 Why RL may now be "the cake," not the cherry on top 25:46 Why RL started working with large language models 27:29 Is RL "sucking supervision through a straw"? 28:47 Why language may be the grounding layer for intelligence 31:46 A contrarian take on the Bitter Lesson 32:41 What test-time compute actually is 34:50 How RL gives models the ability to think 35:40 Verifiable rewards, math, coding, and the messy real world 38:00 What physics can teach us about AI 42:08 Is there a thermodynamics of AI? 43:08 From Erdős problems to Einstein-level AI 45:16 Is AI already doing original science? 45:51 How far are we from AI automating AI research 47:41 Why Dan is excited about the future of science

Matt Turck

69,801 görüntüleme • 3 ay önce

I'm often asked for the best public example of AI evals done right for a real, production product. I finally have an answer. Teresa Torres shares how she shipped an AI interview coach, and used evals to rapidly squash bugs and improve the product. Teresa shows how she: 1. did error analysis FIRST to find real issues (instead of using generic metrics) 😍 2. used Jupyter notebooks to analyze errors 3. built custom annotation tools + custom widgets in notebooks 4. built a LLM-judge and assertions to test for specific errors 5. iterated through this feedback loop until it worked. 6. kept things simple the whole time It's also probably the best commercial for Jupyter notebooks you can imagine. 🥰 Chapter summary below. Link to YT in next thread 00:00:00 - Intro 00:01:45 - The Product: Building an AI Interview Coach 00:06:34 - The Problem: How Do I Know if My AI Coach is Any Good? 00:10:15 - Using Airtable for Traces and Annotation 00:12:15 - Discovering Jupyter Notebooks and Designing the First Evals 00:15:15 - Example Evals: LLM-as-Judge vs. Code-Based Assertions 00:21:00 - Learning Python with ChatGPT to Analyze Eval Results 00:31:00 - VS Code, Custom Tools, and an Eval Investigation Notebook 00:39:45 - Building a Custom Annotation Tool with Claude 00:41:00 - From Personal Project to Production App 00:46:02 - How Should PMs and Engineers Collaborate on AI Products? 00:55:45 - Q&A: Capturing Feedback and Annotations from End Users 00:58:11 - Q&A: Is a Technical Background Necessary to Build AI? 01:02:28 - Q&A: What's Next for Teresa? 01:03:13 - Q&A: Unpacking the Micro-Decisions of Building an AI App

Hamel Husain

51,376 görüntüleme • 1 yıl önce

E159: Hyperliquid: Housing all of Finance jeff.hl came back on the When Shift Happens Podcast to talk about the Hyperliquid journey since the TGE and what the future holds for one of the most loved and prolific protocols in the space Hyperliquid Timestamps 0:00 Intro 2:01 Singapore 2:27 Reminiscing on the Token Launch 5:00 Was This Scale Of Wealth Expected? 6:28 Doing The Right Thing In Crypto 9:07 The Responsibility that comes with Billions of $ 11:10 Jupiter KAST 11:51 Bringing Hyperliquid to the masses 15:21 Pre TGE and Post TGE: Operational difference 20:13 Choices on what to build Internally vs Externally 22:05 How to build a reliable team 24:51 Did the Team celebrate the HYPE wealth Generation event? 26:45 How to test talents for High Integrity 28:31 How much does the Hyperliquid team sleep? 30:05 Employee Vesting Fears 31:41 Dealing with FUD 32:28 How Does Jeff Personally Handle FUD 35:02 Token "Buybacks" critics 37:20 Why Hyperliquid can't have Discretionary "Buybacks" 39:04 HyperEVM, explained Simply 40:00 Paradex Zodl 40:41 HyperEVM: Success so Far? 44:05 HIP-3, explained Simply 47:44 What makes Hyperliquid's approach different 48:19 Why Should People Care? 51:33 Bring All Finance On Chain 52:08 Why Is The Hyperliquid Approach Better? 53:47 Key Numbers showing that Hyperliquid Is Doing it right 59:01 What Has the Unit team demonstrated with spot trading on Hyperliquid in 2025 1:03:29 HIP-4: Outcome Markets 1:08:01 Trezor Sui 1:08:58 What does "Housing All Of Finance" mean? 1:10:51 Why Hyperliquid is not a crypto company 1:12:23 Why Does Hyperliquid have A Stablecoin USDH (Native Markets) 1:14:39 What Is Kinetiq & Why Does It Matter? 1:16:15 Why Is What HyperLend Is Building Important For HyperLiquid 1:23:39 Where did Fairness cost the most? 1:24:47 What should Hyperliquid be Remembered for? 1:25:24 Why should people stay in Crypto when there's an AI brain drain? 1:28:10 Closing Thoughts

MR SHIFT 🦁

584,416 görüntüleme • 7 ay önce

Inside Nemotron and NVIDIA's AI lab: my conversation with Bryan Catanzaro (Bryan Catanzaro). NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? We go deep into the Nemotron models, what it takes to build a top AI lab, and the future of frontier AI. 01:33 - Is open source AI catching the frontier? 05:29 - Do closed labs blocking distillation slow open source down? 07:42 - Is the US falling behind China? 10:30 - Why companies actually choose open models 12:39 - A "crazy" 2008 bet: machine learning on GPUs 15:33 - Working with Andrew Ng and Dario Amodei at Baidu 17:41 - Coming back to NVIDIA: DLSS and the birth of Megatron 21:55 - The real reason NVIDIA builds its own models 24:28 - Is Moore's Law really dead? 33:37 - The Nemotron family: Nano, Super, Ultra 35:09 - Built for agents: why NVIDIA bets on speed 36:02 - How you train a 550B model in 4 bits 39:25 - Hybrid Mamba-Transformer, explained simply 42:31 - Mixture of experts, and why NVIDIA built NVL72 around it 47:26 - Why a 1-million-token context window matters 49:26 - Multi-token prediction: how the model predicts 5 tokens at once 52:47 - Multi-teacher distillation: teaching one model from many 58:01 - Where reinforcement learning goes next 01:00:16 - Inside NVIDIA's research org: "the mission is the boss" 01:04:03 - How NVIDIA decides who gets the GPUs 01:10:53 - Why NVIDIA still feels entrepreneurial after 33 years 01:12:58 - Why Bryan doesn't believe in the singularity 01:17:50 - The AI backlash 01:19:18 - The controversial case: open AI is safer than closed

Matt Turck

56,954 görüntüleme • 2 ay önce

.Ben Shapiro at UATX with Joe Lonsdale and Niall Ferguson. 00:00 — UATX President Carlos Carvalho 03:16 — Niall Ferguson introduces Ben Shapiro 07:26 — Why UATX is important 08:09 — Why Americans hate each other 10:12 — Emotivism 11:57 — The death of politics 12:30 — Hannibal Lecter skin suits 12:53 — Conspiracy theories 14:11 — Why people don't go to church anymore 15:04 — Vaccines 16:02 — Social engineering & weak professors 17:08 — The last time Harvard meant veritas 17:49 — Epistemic humility 18:30 — Read the Federalist Papers 19:56 — War of all against all 21:03 — Capitalism & soul sickness 22:42 — Tribalism 23:38 — JS Mill and debate culture 26:09 — How to restore our institutions 26:40 — Why UATX matters 27:52 — Joe Lonsdale interviews Ben Shapiro 28:10 — What is a college degree worth? 29:46 — What Jews should learn from Christians and vice versa 32:21 — Candace Owens 34:24 — Audience Q&A: Abraham Lincoln & the Declaration of Independence 36:10 — Q&A: Constitutional boundaries 38:31 — Q&A: Tucker Carlson 42:09 — Q&A: American ingratitude 45:02 — Q&A: How to unite our country 48:24 — Q&A: How to repair our institutions 51:41 — Q&A: The three most important words in the English language 54:25 — Q&A: The future of populism 57:08 — Q&A: Lizard brains 59:43 — Q&A: How Israeli politics work 01:04:03 — Q&A: How to strengthen America 01:07:29 — Q&A: Practical advice for students 01:09:07 — Standing ovation for Ben Shapiro 01:09:36 — President Carlos Carvalho's speech 01:11:08 — How to raise lions Recorded: Sunday, April 26.

University of Austin (UATX)

36,540 görüntüleme • 4 ay önce