正在加载视频...

视频加载失败

It is very easy to make mistakes when creating evals for your AI product. Shreya Shankar and I run through the most common mistakes in this talk (with memes 🌶️!) . Chapter summaries below: 00:51 Foundation model benchmarks are not the same as your application evals 03:00 Generic Evals...

46,137 次观看 • 1 年前 •via X (Twitter)

10 条评论

Hamel Husain 的头像
Hamel Husain1 年前

35% discount to our upcoming course: Youtube version of this video:

Xeophon 的头像
Xeophon1 年前

@sh_reya While it almost is a meme at this point, „look at your data“ is the most important step. Now that I am reading so many benchmark papers, I immediately look whether they done human validation - if they didn’t, there’s always a follow-up paper which does human filtering

Oliver Beavers 的头像
Oliver Beavers1 年前

@sh_reya Signed up and excited to learn!

ikka 的头像
ikka1 年前

Thanks for making this course. But how do you "monitor" the performance of an AI system in production? >Metrics that indicate the correctness of the systems cannot be calculated in prod, as we don't have GT > The signals one can derive from production, such as implicit & explicit feedback, are often weak and assumed to be true. Would love to hear your thoughts.

Hamel Husain 的头像
Hamel Husain1 年前

@sh_reya We discussed this in the video! In prod you also often need to do A/B testing with additional metrics that are more aligned with business outcomes That part is fairly similar to the same paradigm as traditional ML monitoring

Raunak Chowdhuri 的头像
Raunak Chowdhuri1 年前

@sh_reya @raindrop_ai

Johnson Thomas, MD, FACE 的头像
Johnson Thomas, MD, FACE1 年前

@sh_reya Very true. General metrics might be useless or misleading in particular cases.

MO 的头像
MO1 年前

@sh_reya Impressive, thanks a lot for sharing.

Jose Lopez 的头像
Jose Lopez1 年前

@sh_reya @bytetweets can this be useful?

Sydney 的头像
Sydney1 年前

@sh_reya So many pitfalls to avoid! Thanks for the insightful talk and chapter summaries 🤩

相关视频

I'm often asked for the best public example of AI evals done right for a real, production product. I finally have an answer. Teresa Torres shares how she shipped an AI interview coach, and used evals to rapidly squash bugs and improve the product. Teresa shows how she: 1. did error analysis FIRST to find real issues (instead of using generic metrics) 😍 2. used Jupyter notebooks to analyze errors 3. built custom annotation tools + custom widgets in notebooks 4. built a LLM-judge and assertions to test for specific errors 5. iterated through this feedback loop until it worked. 6. kept things simple the whole time It's also probably the best commercial for Jupyter notebooks you can imagine. 🥰 Chapter summary below. Link to YT in next thread 00:00:00 - Intro 00:01:45 - The Product: Building an AI Interview Coach 00:06:34 - The Problem: How Do I Know if My AI Coach is Any Good? 00:10:15 - Using Airtable for Traces and Annotation 00:12:15 - Discovering Jupyter Notebooks and Designing the First Evals 00:15:15 - Example Evals: LLM-as-Judge vs. Code-Based Assertions 00:21:00 - Learning Python with ChatGPT to Analyze Eval Results 00:31:00 - VS Code, Custom Tools, and an Eval Investigation Notebook 00:39:45 - Building a Custom Annotation Tool with Claude 00:41:00 - From Personal Project to Production App 00:46:02 - How Should PMs and Engineers Collaborate on AI Products? 00:55:45 - Q&A: Capturing Feedback and Annotations from End Users 00:58:11 - Q&A: Is a Technical Background Necessary to Build AI? 01:02:28 - Q&A: What's Next for Teresa? 01:03:13 - Q&A: Unpacking the Micro-Decisions of Building an AI App

Hamel Husain

51,376 次观看 • 1 年前

An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. LangChain founder Harrison Chase joined us at our Sequoia Capital Own Your Intelligence to talk about the piece that often gets the least attention: the harness. He offers a clear heuristic for when to build your own. The more out of distribution you are from what the models were trained on, the more you'll want to customize. And good technical content on how to actually measure performance with evals and langsmith. 00:00 Introduction 00:58 The three parts of an agent: harness, model, context 02:12 What a harness actually does 03:25 Customizing the core loop with middleware 04:41 Sandboxes, file systems, sub-agents, summarization 05:47 Cognitive architectures — and when you still need them 07:03 Build your own harness or use off the shelf? 08:24 In-distribution vs. out-of-distribution: the file-editing example 09:39 Why evals define what "good" means in an organization 11:04 Harbor: what an eval task actually looks like 12:11 Comparing harnesses and models on accuracy, latency, and cost 13:20 Why observability is underrated — it's usually the context 14:34 The data flywheel: traces → curation → experiments 15:42 Getting feedback through UX design and online evaluators 16:51 Demo: LangSmith Engine 19:23 Q&A: Running Engine on Engine, and "codex-ification" 20:44 Q&A: Will harnesses converge or diverge?

Sonya Huang 🐥

75,702 次观看 • 22 天前

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existential and strategic topic for our portfolio. Intelligence is the product. Companies want to shape it and own it and let it compound within their own walls. Not your weights, not your product. Now, with frontier open-weight models and fantastic tooling/infrastructure, owning your intelligence at the frontier is finally becoming possible. The result: every application company we work with is embarking on the journey of doing their own research on post-training, evals, harnesses, etc. The hottest neolabs may just be Harvey, Factory, RamPrasad "RamP!" Moudgalya, etc. The list goes on. We held a summit Sequoia Capital to convene our portfolio on this topic, together with Gabe Pereyra (Harvey) on building Harvey Labs, Lin Qiao (Fireworks) on post-training, Harrison Chase (LangChain) on harnesses + evals, Brendan (can/do) () on RL environments and synthetic data, Arjun Karanam (Trajectory) on online continual learning. Opening talk below; rest to come this week! 00:00 What is sovereign AI (and what it isn't) 01:24 Centralized vs. decentralized intelligence 02:54 Four reasons companies own their models: cost, speed, performance, destiny 04:22 "Not your weights, not your product" 05:32 The application companies are the newest neo labs 07:05 Step 1: Deciding what to own vs. rent 09:51 Step 2: Build the team (and don't shoehorn your platform team) 11:17 Step 3: Legibility – why your research has to be visible 12:33 Step 4: The technical roadmap 13:56 The stack: production vs. development 15:16 Opening Pandora's box – base models, harnesses, context

Sonya Huang 🐥

127,724 次观看 • 24 天前

My conversation with OpenAI co-founder Greg Brockman This is the most detailed first-person account of the 72 hours after Sam Altman was fired. We also go deep on what comes next: the global race to AGI, why ChatGPT stopped showing reasoning, how much of OpenAI's own code is now written by AI ("it's hard to know what percent is not"), and the untold story of how OpenAI actually started in 2015. 00:00:00 Introduction 00:00:49 Meeting Sam Altman and Starting OpenAI 00:02:40 Building the Founding Team 00:04:25 DeepMind's Lead Over OpenAI 00:04:54 Changing OpenAI to a For-Profit Model 00:06:05 Breakthrough Moments at OpenAI 00:08:22 What Dota 2 Meant for OpenAI 00:10:04 Reasoning Versus Prediction 00:11:59 Tensions Grow at OpenAI 00:15:44 Sam Altman's Firing 00:17:49 Greg Quits OpenAI 00:19:56 Sam Explores Deal with Microsoft's Satya 00:20:28 Petition for Altman's Return 00:23:43 Ilya Sutskever Leaves OpenAI 00:24:59 Lessons Learned after Sam Ousting 00:28:22 The Thing Ilya Said that Greg Can't Forget 00:32:22 Is AI Going Parabolic? 00:33:24 How Much of OpenAI's Code is Written by AI? 00:36:21 Do AI Chatbots Tell Us What We Want to Hear? 00:38:06 The Global AI Race to Reach AGI 00:38:40 What Happens if US Doesn't Reach AGI First? 00:39:49 Are Countries Stealing AI Advancements? 00:40:38 Why ChatGPT No Longer Shows Reasoning 00:41:47 The Finite Constraints of Compute 00:43:38 On Investing Early in Data Centers 00:46:31 The Future of Data Center Specialization 00:47:52 How to Decide Whose Queries to Serve 00:49:08 OpenAI on Consumer vs Enterprise Models 00:53:05 Data Centers in Space? 01:00:56 What Should AI Regulation Look Like? 01:04:33 The Future of AI-Powered Entrepreneurship 01:04:44 AI and Job Loss 01:07:15 The Skills Young People Should Invest In 01:11:30 What Does Success Look Like For You? Full episode on X below. Also find it on: • YouTube: • Spotify: • Apple:

Shane Parrish

450,952 次观看 • 4 个月前

one thing that has saved my projects more time than I can count is evals boy was I excited when florian, quite literally an expert in benchmarks, agreed to hop into a ~2h interview to do a walkthrough of what the eval landscape looks like in 2026 (and also answer my personal business questions on the subject) given that now running frontier model through benchmarks is a vector for hacking other systems in order to avoid doing work (looking at you sol), I think it's more important than ever to educate folks on the evals situation. had a lot of fun throughout this session and I hope that you learn a thing or two! enjoy! 🌹 table of content: 0:00:00: are AI Benchmark broken? 0:05:45: Florian Brand background 0:09:00: what motivates florian to work on evaluation? 0:13:33: what is the mirrorcode benchmark about? 0:18:20: cheating in agent benchmark is insaneeeee 0:24:08: LLM benchmarks in era of agents 0:26:30: what’s up with the pelican man 0:28:27: evals are about capabilities 0:31:46: components of running evals 0:35:30: the volume of things to audit is huge!!! 0:40:20: expert answers are wrong hahahaha 0:46:00: api providers aren’t the same 0:48:00: benchmark narrow capabilities (synthetically) 0:50:56: link between eval and environment 0:53:45: small validated benchmark or massive bench? 0:56:11: what is your flow to review a benchmark? 0:58:30: tracking work capabilities with evaluation 1:00:20: slide deck in industry is all vibecoded 1:03:30: harness impact in the evaluation 1:07:39: hardware/sandboxes impact evaluation too! 1:11:00: “is it going to get worse?” 1:12:40: all components influence the final score 1:13:50: training models on different harnesses? 1:17:20: is the model just the weights or it’s all of it? 1:19:30: how to craft benchmark that prevent to cheating and undereliciting models in 2026 1:23:19: ways agents cheat and steal 1:26:00: correct elicitation of capabilities is important 1:36:00: building evaluation on prime intellect 1:45:10: how do you design interactivity benchmarks? 1:48:40: do you think evals are well set to reflect real world performance? 1:52:50: what will the benchmarking landscape will look like in 1 year

Yacine Mahdid

12,923 次观看 • 1 个月前