
Sonya Huang 🐥
@sonyatweetybird • 29,274 subscribers
funding big computer @sequoia
Videos

What happens when there are more agents than eyeballs browsing the internet? How will the economics of the internet be reinvented? Parag Agrawal thinks Shapley values may hold the answer. Parag built Twitter over a decade, eventually becoming CEO and selling the company to Elon. He's spent the last three years building Parallel Web Systems, a search engine built for agents instead of humans. His core argument: 1) human click data is a bug. Agents aren't just a new technology, they're a distinct customer – and the feedback loop that made Google great is the wrong signal for the thing actually doing the work; 2) the ad-funded web assumed scarce human attention. If agents show up instead of eyeballs, the business model underneath the internet has to be rebuilt, or good content stops being published. The conversation covers: — why he shipped a search agent before a search engine, and how that let him grow the index incrementally instead of buying a full web crawl up front — the billion-to-billion matching problem: pull the right 1,000 tokens out of a trillion web pages, and your agent uses under half the tokens — going from a 3-second compute budget to 200 milliseconds — why he doesn't think Parallel is a neo-lab: "our output is a complement to a model" — the Google Cloud deal — on GCP, your grounding options are now Google Search or Parallel Search — reinventing the economics of the internet, Shapley values as a payment rail for publishers, and why the company was originally incorporated as Shapley Inc — why routing 2-10% of inference spend to web data would dwarf every content business outside the walled gardens — the web going from pull to push: "call me if this happens" 00:00 Introduction 03:25 What Is Web Search 05:17 Why Start a New Index 07:52 Search Agents First 10:17 Not a Neolab 13:14 Agents vs Google Search 19:38 Inside the Search Stack 28:59 Search Multipliers With Agents 30:21 Meeting Prep Agent Workflows 31:46 Quality Cost Latency And Turbo 32:42 Are Agents Overtaking Humans 34:28 Ads Model Meets Agent Web 37:20 New Incentives For Content 40:48 Shapley Values Attribution 47:46 Parallel Web And Future Vision Hosted with my very unwilling co-host Andrew Reed and Sequoia Capital
Sonya Huang 🐥126,158 görüntüleme • 18 gün önce

Two of the people most responsible for scaling the transformer are now betting on a next act. Jerry Tworek ran the Reasoning 🍓 team at OpenAI. rohan anil was a pre-training lead on Gemini after years at Google Brain and Anthropic. They just started to find what comes next. Their core argument: (1) models are trained in the lab but deployed in the real world and can't keep learning once they leave; (2) AI research is done by humans today but models will be able to explore and uncover new advances more rapidly and systematically (controversial but timely w this week's petition). The conversation covers: — why Jerry expected AGI in 2025 and what changed his mind — the two kinds of learning from experience, and why RL only captures one — the computational depth problem baked into today's architectures — why the biggest labs can't afford to look for a transformer replacement — the kernel competition where humans + $100K of coding agents found a 60x speedup no frontier model comes close to — a definition of AGI you can actually test: a model that improves itself with no human in the loop 00:00 Introduction 01:46 Appreciating Transformers 02:44 Scaling Hits Limits 04:54 Why Architecture Matters 05:32 RL Reality Check 07:32 Test Time Learning 09:52 Economics Of Scaling 12:47 Why Start A Company 14:24 Rohan On Transformers 19:11 Computational Depth Problem 20:32 When Transformers Top Out 23:22 Beyond Reinforcement Learning 26:41 Optimization And Efficiency 34:24 Building An Automated Lab 39:45 Kernel Automation Roadmap
Sonya Huang 🐥234,282 görüntüleme • 1 ay önce

Today we release my favorite episode of Training Data yet: the great Rich Sutton. Richard Sutton wrote the textbook, wrote The Bitter Lesson (and many other on-point essays like "Self-Verification, The Key to AI"), and trained a mafia of talented students who went on to change the AI landscape forever including David Silver, inventor of built AlphaGo. Khurram Javed was Rich's PhD student at Alberta and wrote The Big World Hypothesis. They just left academia to start Oak Lab Their core argument: (1) The Bitter Lesson: the world is massively more complex than any model of it, so anything trained on human-curated data has a ceiling (2) Continual Learning: intelligence is continual by definition, and today's models stop learning the moment they ship. The conversation covers: — what The Bitter Lesson actually says, and what people get wrong — why synthetic data is "just a big mistake," and the Big World Hypothesis behind it — how LLMs are both a positive and a negative example of his own essay — why no animal learns by supervised learning, and what squirrels can do that we can't — the cure for catastrophic forgetting: per-weight step sizes and continual backprop — why the biggest labs can't take a path where performance gets worse before it gets better — a trillion parameters on 20 watts, and the Moore's Law math that makes it plausible — why the endpoint isn't one mind but one design, running as many minds It was both a fun generative idea- and debate-filled conversation, and a surprisingly human one too. Rich, thank you for beating cancer and changing the trajectory of AI. 💙 00:00 Introduction 02:10 An AI winter, a cancer diagnosis, and the move to Alberta 07:07 Writing "The Bitter Lesson," and what people get wrong 09:53 Are LLMs a positive or a negative example of it? 11:03 Synthetic data is "just a big mistake," and the Big World Hypothesis 18:01 AlphaGo, human priors, and why prior knowledge and learning should be friends 22:37 "Their weights never change": do LLM assistants actually learn? 26:09 Babies, squirrels, and why no animal learns by supervised learning 32:02 Rockets, imagination, and where paradigm shifts come from 36:42 The Alberta Plan and its 12 steps 38:53 Catastrophic forgetting and the cure 43:43 Oak's biggest ambition: a self-maintaining mind 47:56 Why the big labs are stuck in a local minimum 49:13 If everything goes right: LLMs, many minds, and hiring The man who pioneered reinforcement learning thinks the rest of the field is weird, and lays it all out in today's episode. Together w/ Alfred Lin Sequoia Capital
Sonya Huang 🐥119,115 görüntüleme • 25 gün önce

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existential and strategic topic for our portfolio. Intelligence is the product. Companies want to shape it and own it and let it compound within their own walls. Not your weights, not your product. Now, with frontier open-weight models and fantastic tooling/infrastructure, owning your intelligence at the frontier is finally becoming possible. The result: every application company we work with is embarking on the journey of doing their own research on post-training, evals, harnesses, etc. The hottest neolabs may just be Harvey, Factory, RamPrasad "RamP!" Moudgalya, etc. The list goes on. We held a summit Sequoia Capital to convene our portfolio on this topic, together with Gabe Pereyra (Harvey) on building Harvey Labs, Lin Qiao (Fireworks) on post-training, Harrison Chase (LangChain) on harnesses + evals, Brendan (can/do) () on RL environments and synthetic data, Arjun Karanam (Trajectory) on online continual learning. Opening talk below; rest to come this week! 00:00 What is sovereign AI (and what it isn't) 01:24 Centralized vs. decentralized intelligence 02:54 Four reasons companies own their models: cost, speed, performance, destiny 04:22 "Not your weights, not your product" 05:32 The application companies are the newest neo labs 07:05 Step 1: Deciding what to own vs. rent 09:51 Step 2: Build the team (and don't shoehorn your platform team) 11:17 Step 3: Legibility – why your research has to be visible 12:33 Step 4: The technical roadmap 13:56 The stack: production vs. development 15:16 Opening Pandora's box – base models, harnesses, context
Sonya Huang 🐥128,376 görüntüleme • 1 ay önce

Very special Training Data drops today: Making Cities Awesome ☺️ We haven’t featured an application company on Training Data in a while. The diffusion of AI is the Most Important Thing, and the highest-stakes application category of all is the one sitting near the base of Maslow’s hierarchy: safety. Today I am excited to bring Nick Noone and Ben Rudolph of Peregrine onto the podcast. You may not have heard of Peregrine yet, but their technology is almost certainly keeping you and your loved ones safe, and they are rejecting the surveillance state while doing it. Nick and Ben talk about all the interesting ways that AI is being deployed in customer environments at Peregrine, which include some of the coolest use cases of long-horizon agents and vibe coding that I’ve seen: * a cold-case agent that reproduced an exoneration detectives had reached by hand * a Wisconsin county that placed a suspect using cell records buried in 300GB of evidence * a semantic search tool for identifying active threats to a synagogue * a vibe-coded hurricane simulator to help one Forward Deployed Engineer root-cause an inflection in weather-related incidents Nick formerly ran Palantir’s Special Operations Command business unit and deployed into high-stakes situations in the Middle East. Ben comes from almost the opposite background: building humanitarian technology for refugees and disease prevention in Africa and India. But the two of them converged on a similar thesis: that this technology could be used to make cities work better for the people who live in them. Today, Peregrine powers law enforcement, emergency medical services, fire and rescue, and other services in more than 400 communities globally. Public distrust in surveillance has reached a boiling point, for good reason. Other public safety technology companies grow by collecting more data. Peregrine inverted the model: no sensors, no new data, a business built on connecting the data and information cities already own, with permissions, policy and auditability from day one. Authorized users make sense of existing data when they have a legitimate reason to access it, and not otherwise. Data governance, access controls and audit logs are unglamorous work, but they are the safeguards that let a city protect safety and privacy at once. Nick and Ben’s deep integrity shines through in the episode. I am thankful for their stewardship of this critically important category. 00:00 Introduction 02:07 What Forward Deployed Engineering Means 03:58 What Silicon Valley Gets Wrong 05:23 Humanitarian Work, Refugee Work, And Downstream Data Problems 08:25 Why Cities, Why Safety 10:45 Two Dozen Nos And San Pablo PD 14:19 Building Through Defund The Police 18:16 The Inversion Of The Collection Model 21:20 Data Ownership And Governance 22:57 From Nice Search To Deep Analysis 29:59 Agents Writing The Integrations 31:45 The Cold Case Agent 35:02 The Anti-Network-Effect Proposition 38:40 Facial Recognition And Hard Decisions 40:48 Technology For The Underdogs 42:54 Trusting The Individual Contributor 48:50 Ten Thousand Cities
Sonya Huang 🐥32,836 görüntüleme • 12 gün önce

An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. LangChain founder Harrison Chase joined us at our Sequoia Capital Own Your Intelligence to talk about the piece that often gets the least attention: the harness. He offers a clear heuristic for when to build your own. The more out of distribution you are from what the models were trained on, the more you'll want to customize. And good technical content on how to actually measure performance with evals and langsmith. 00:00 Introduction 00:58 The three parts of an agent: harness, model, context 02:12 What a harness actually does 03:25 Customizing the core loop with middleware 04:41 Sandboxes, file systems, sub-agents, summarization 05:47 Cognitive architectures — and when you still need them 07:03 Build your own harness or use off the shelf? 08:24 In-distribution vs. out-of-distribution: the file-editing example 09:39 Why evals define what "good" means in an organization 11:04 Harbor: what an eval task actually looks like 12:11 Comparing harnesses and models on accuracy, latency, and cost 13:20 Why observability is underrated — it's usually the context 14:34 The data flywheel: traces → curation → experiments 15:42 Getting feedback through UX design and online evaluators 16:51 Demo: LangSmith Engine 19:23 Q&A: Running Engine on Engine, and "codex-ification" 20:44 Q&A: Will harnesses converge or diverge?
Sonya Huang 🐥77,519 görüntüleme • 1 ay önce

"Member of the technical staff" is the hottest job title in SF right now. What's behind the name? OpenAI chose this title deliberately to blow up the previous industry dichotomy between researchers and engineers. The best researchers in AI right now aren't academics in a pure lab environment; they get their hands dirty technically writing code and digging into the data and implementations e.g. Alec Radford Bob McGrew, former Chief Research Officer at OpenAI and one of U.S. Army's newest recruits to Detachment 201, joined us on Training Data to share more about the secret sauce behind leading OpenAI's research org, the three legs of the stool to AGI, and why he thinks we're already there.
Sonya Huang 🐥358,649 görüntüleme • 1 yıl önce

***Lights, inference, action…*** I’m so happy to share that Sequoia Capital has led the Seed in Preview. Developers are flying in magical AI-native editors. But creative tooling is still stuck in the pre-AI stone age. stefo is building the production platform for AI video. It is robust, it is full featured, it is everything a professional videomaker could ask for (as our star-studded alpha client list will attest), and most importantly it is built with care and the utmost attention to craftsmanship <3 Grateful to join many greats of the video and creative tools space on the cap table including @emery_wells of Frameio, Badrul Farooqi Kyle Parrish ex Figma, Burkay Gur batuhan the fal guy of fal, Soleio, and many other friends.
Sonya Huang 🐥29,503 görüntüleme • 1 ay önce

When should you start post-training your own models? Fireworks CEO Lin Qiao’s answer: after product-market fit. Not because it's hard... but because only after PMF is the data coming off your product surface worth training on. Lin joined us for our Sequoia Capital "Own Your Intelligence" event to host a workshop on all things post-training; what works, what breaks, and how not to let the model outsmart you. Must listen!! 00:00 Introduction 00:37 What Fireworks sees across thousands of AI applications 02:47 Off-the-shelf APIs and the problem of keeping your taste 03:58 What "owning your intelligence" actually means 05:43 The progression: prompting → RAG → SFT → preferences → RL 07:20 Why this mirrors how humans learn 09:03 Matching the technique to the problem you actually have 10:46 Where teams get stuck: data quality and vibe evals 12:28 Reward hacking: the model that wrote zero lines of code 13:59 Training-to-serving alignment (and why quality drops) 15:55 Post-training in healthcare and security 17:31 From coding to every co-work domain 19:35 Incumbents, cost burden, and not scaling into bankruptcy 21:26 How much control do you want? 23:24 Q&A: What makes a good reward signal 25:00 Q&A: When to start thinking about post-training
Sonya Huang 🐥28,142 görüntüleme • 1 ay önce

Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring Federico Cassano (Composer Lead at Cursor) and Dmytro Dzhulgakov (Co-Founder at Fireworks). The Cursor team trained Composer 2 on Fireworks by starting with a strong base model (Kimi 2.5) and performing large-scale mid-training on code tokens and web data to learn common patterns and libraries, followed by a large-scale Reinforcement Learning run to learn how to navigate the Cursor harness, call tools, and write correct code. Today's episode dives into the systems and infrastructure challenges of making that large RL run happening, and there were many (!!), from numerical mismatch to global distribution to synchronizing rollouts across asynchronous pipelines to keeping track of expert activation across runs and more. Extremely nerdy in-the-weeds challenges that Federico and Dima were delighted to nerd out on together :) Beyond RL infra, we also discussed Online vs Simulated rollouts, self-summarization for long-horizon agents, environment design ("the most powerful RL environment is the product itself"), and other technical nuggets. PS: We filmed this episode before the SpaceX news, while the Cursor team was still compute-constrained. While Cursor now has *all* the flops, the takeaways and hurdles crossed ring true for any serious application-level company that is racing to post-train their own models. I believe that more serious application companies will go the way of Cursor and post-train their own models. 00:00 Introduction 00:53 Why Cursor Trained Composer 2 04:55 Specialization vs Bitter Lesson 06:16 Composer 2 Training Recipe 16:32 Scaling RL Infrastructure Globally 23:32 Floating Point Drift 25:11 MoE Sensitivity Explained 26:25 Router Replay Fix 27:19 Real Time RL Loop 31:49 Long Horizon Agents 34:29 Why RL Everywhere 37:34 LLM as Judge Rewards 39:14 RL in Hard Domains 40:13 Build Your Own Environments 44:34 Closing Thoughts
Sonya Huang 🐥79,834 görüntüleme • 3 ay önce

The advanced civilizations of sci-fi legend (Banks, Asimov, etc) have some form of simulation to guide society. Joon Sung Park is taking a crack at building that simulator with Simile. As Joon's cofounder Percy Liang puts it: great science starts with a great measurement. From Smallville in 2023 to today, Simile is trying to build the Hubble Telescope equivalent for simulating human behavior. 00:00 Introduction 01:49 Building Generative Agents 02:29 Valentines Day Emergence 03:33 From GPT 3 To Agents 05:03 Social Computing Problem 06:19 Social Simulacra Subreddits 07:57 Models Getting Good Enough 08:57 Humans Are Not Rational 10:04 Turning Research Into Simuli 11:55 Validation And Accuracy Proof 12:43 Customer Workflow CVS Example 16:11 Why Collect Real Data 17:51 Behavioral Signals And RCTs 21:52 Use Cases And Second Order Effects 26:31 Evaluating Convergence Divergence 31:58 Big Societal Simulations Ahead 36:08 Future Of Simulation
Sonya Huang 🐥46,958 görüntüleme • 2 ay önce

Don't sleep on Google DeepMind in AI... This week on Training Data, Google Labs VP Josh Woodward gave us the BTS on Google's imagination playground for AI, from Notebook to Mariner (computer use agent) to Veo (video models). Thanks Josh for the spicy convo and hot takes :)
Sonya Huang 🐥183,357 görüntüleme • 1 yıl önce

Baumol cost disease is real, AI is the solution Take law -- imagine a world where consumers have plentiful access to high-quality legal services Crosby is building an AI-native law firm towards that vision. @Ryanjdaniels John Sarihan share more on Training Data!
Sonya Huang 🐥69,577 görüntüleme • 1 yıl önce

Best technology M&A of all time has to be NVIDIA's $6.9 Billion acquisition of Mellanox in 2020. It was a special treat to interview Michael Kagan, CTO of NVIDIA and co-founder/CTO of Mellanox, for today's Training Data episode. Michael has been driving forward the Compute Frontier for more than 40 years now, first as Chief Architect at Intel in the 90s, then Co-Founder and CTO at Mellanox, and for the last 5 years CTO at NVIDIA. There's nobody better positioned than Michael to share the complete history of the compute frontier and what's ahead, from decades pushing forward Moore's law (squeezing more transistors on a chip) to the last decade of work scaling beyond single chip physics limitations (scaling out to 100K+ GPU clusters). Interconnect is the secret sauce enabling compute to scale beyond chip-level Moore's law. Connecting a fabric of 100K+ GPUs to function as a single unit of compute is the enabling technology for today's intelligence explosion. But other things break at the 100K+ GPU cluster scale: individual chips inevitably fail, power and networking become more complex, etc etc. Net effect of scale out: we've inflected the silicon frontier from Moore's Law (2x every 2 years) to Huang's law (~10x a year). Very excited about today's episode! Learned so much from Michael with Pat Grady
Sonya Huang 🐥56,175 görüntüleme • 10 ay önce

Claude Code and Suno have more in common than you might think: "It's fun to build things, and it's fun to use what you build." AI lets people be creative in almost any domain, from coding to making music. Today on Training Data, Mikey shares his thesis for why generative AI is the newest form of active entertainment (the next 'gaming'), music as a cultural phenomenon vs creative expression platform, and more. My favorite part was Mikey's explanation of why Suno learns music theory implicitly vs explicitly: "In Western music, there are 12 tones. If you tell the model there are 12 tones, it will only ever produce those 12 tones. You will be forever limited. And if you tell the model there's 200 instruments, those are the only sounds that you'll ever be able to make." The more you constrain a model with what humans already know, the less capable it becomes. By treating everything as pure sound, Suno built what Mikey calls a totally generalized "music-making machine." Such is the power of neural nets.
Sonya Huang 🐥22,803 görüntüleme • 4 ay önce

Today on Training Data: Sanjit Biswas, founder & CEO of Samsara (NYSE:IOT) and former Sequoia Capital backed founder of Cisco Meraki Sanjit shares the ups & downs of running neural nets on constrained compute and power footprints in the real world, ~2-10 watts Physical AI is hard 🫡
Sonya Huang 🐥35,714 görüntüleme • 9 ay önce

Some of the most iconic consumer products -- Tide Pods, M&Ms, etc -- were born out of user research studies. LLMs democratize deep user research to every decision. Synthetic audiences will go even further. Alfred Wahlforss of Listen Labs shares more on Training Data cc Konstantine Buhler
Sonya Huang 🐥16,289 görüntüleme • 3 ay önce

Today on Training Data, the OpenAI team behind ChatGPT agent explain how Agent Mode works, combining: 1) Deep Research (text based research agent) 2) Operator (GUI/action based computer agent) 3) Other new tools (terminal, computer apps) 4) Tied together with shared state to create an agent that's highly capable at most tasks that humans do on a computer: data science analysis, analyzing spreadsheets, making slides, etc. Thanks for joining us Isa Fulford Casey Chu Zhiqing Sun Lauren Reeder!
Sonya Huang 🐥44,149 görüntüleme • 1 yıl önce

Our most exciting episode of Training Data yet 🍓🍰 OpenAI’s o1 represents a major leap forward by giving models time to "think." Inference-time compute is the next big research frontier. Thrilled to have Noam Brown, ilge, and hunter on the show Pat Grady Sequoia Capital
Sonya Huang 🐥51,628 görüntüleme • 1 yıl önce