
Pat Grady
@gradypb • 30,689 subscribers
Sequoia partner, BC alum, Wyoming native
Videos

Want world class research capabilities, but don’t have the resources of a big lab? At our recent Sovereign AI event, Gabe Pereyra shared Harvey ’s “moneyball” approach. Here’s the playbook: 00:00 Introduction 00:37 Building a research lab on a budget 02:28 Legal Agent Bench, contracting, and the diligence dataset 03:57 Domain experts guiding synthetic data generation 05:23 Why Harvey open sourced its datasets 06:55 Working with the neo labs – and why more than one 08:20 Post-training in-house: building "Associate 1" 09:44 The model serving matrix: 60 countries, fallbacks, SLAs 11:05 Deciding what stays in production 12:29 Simple open source switches and model routing 13:55 Moneyball: "If we win on this budget, we change the game" 14:53 Q&A: Training with sensitive data 17:16 Q&A: Competing for research talent 18:46 Q&A: Designing rubrics that actually challenge frontier models 20:19 Q&A: Where the pipeline breaks — data, research, or infra 22:59 Q&A: The tension in open sourcing a benchmark 25:02 Q&A: Biggest remaining open problems 27:10 Q&A: Competing with horizontal products
Pat Grady179,052 Aufrufe • vor 1 Monat

Most startups celebrate their first couple million of revenue. Factory gave it back. They didn’t have to. They chose to. They decided that the product just wasn’t good enough yet, and they wanted to know that they’d really earned it. They wanted their customers to be obsessed. Two years later they shipped Droid CLI, and they’ve been ripping ever since. The product is fantastic and the team is even better - particularly having been hardened by their “two years in the desert”. Matan Grinberg was one of our very first guests on Training Data, and we are delighted to welcome him back. 00:00 Introduction 01:37 Enterprise Lock In Fears 04:26 No Lock In Promise 05:24 Two Years Early 07:24 Refunding Revenue 13:49 Droid CLI Breakthrough 16:44 Harness Frontier Tactics 26:42 Natural Language Routing 27:32 Open Models Catch Up 29:26 Token Share Forecasts 31:02 Automating Low Leverage Work 32:36 Pricing Beyond Tokens 35:56 Software Factory Vision 43:43 AI Transformation Playbook 46:13 Async Agents And Optimism
Pat Grady229,208 Aufrufe • vor 1 Monat

Intelligence and Experience are orthogonal vectors Terence Tao is perhaps the world’s smartest person, but drop him into an accounting firm or onto a construction site and on day one he’s not going to be very productive Trajectory calls this The Experience Gap, and they have a way to close it Arjun Karanam explained how at our Sovereign AI event: 00:00 Introduction 00:12 Building the platform for continual learning 01:33 The experience gap: models have IQ but no tenure 02:52 Traceability → model spec → better models and harnesses 05:27 Four wishes for the agent ecosystem 06:34 Wish 1: Trace the whole tree — and capture the corrections 08:03 Wish 2: Evals from real traffic, graded in the real harness 09:26 Wish 3: Let the agents cook, and make tool responses informative 10:34 Wish 4: Get comfortable on open weights, experiment with routers 11:51 Why owning your intelligence shouldn't be consulted away 13:15 Demo: import a benchmark, train a model, deploy it 14:28 Q&A: What's the trainable object — weights, harness, or context? 16:08 Q&A: Continual learning without training on customer data 17:13 Q&A: Episodic memory and the hierarchy of feedback 19:37 Q&A: Where continual learning matters most
Pat Grady76,271 Aufrufe • vor 1 Monat

A model is only as good as its data, and we’ve long since exhausted the internet. From here on out, model progress is gated by data production. ’s Brendan (can/do) joined us at our Sovereign AI event to talk about how RL environments get built, and why your data might be your real moat: 00:00 Introduction 00:47 A short history of the data market: crowdsourcing to agentic data 02:29 What an RL environment is: worlds, apps, tasks 03:57 Why only humans can measure the frontier 05:35 Building verifiers is the hard part 06:44 Walkthrough: a real legal RL environment 08:18 Leaderboards — and what open weights change 09:45 Post-training results on Apex Agents 11:17 Three ways companies buy data 12:49 Q&A: How do you price data? 14:17 Q&A: What "data quality" actually means 16:42 Q&A: The misunderstanding about synthetic data 18:17 Q&A: Why RL environments now — and what comes after 21:20 Q&A: Can you scale rubric generation with models? 23:00 Q&A: RL environments for cyber defense 25:33 Q&A: Build data in-house or partner?
Pat Grady66,577 Aufrufe • vor 1 Monat

To date, finding a drug has been a process of guess & check… screening millions of molecules hoping one binds. Chai Discovery is changing the paradigm… describe the molecule you want, and the model designs it. Joshua Meier and Matt McPartlon took antibody hit rates from roughly 1 in 1,000 to 1 in 7. That's the jump from search to engineering. On this episode of Training Data, they get into where the field actually is, why 2024 was the moment to start, what drug discovery looks like when the computer does the design, and more. 00:00 Introduction 01:52 From Discovery to Design 03:25 Protein AI Breakthroughs Timeline 06:04 Why Start in 2024 10:13 Diffusion Models Intuition 11:41 Building the Avengers Team 15:22 Hit Rates and Scaling Laws 25:01 Molecular CAD Vision 25:24 Faster Design Loops 26:32 Future Drug Discovery 28:37 Platform Business Model 31:14 Partnering Reality Check 33:44 Data Flywheel Explained 37:16 Staying Ahead at Scale 39:44 Culture and What's Next
Pat Grady34,554 Aufrufe • vor 1 Monat