Loading video...

Video Failed to Load

Go Home

How OpenAI Builds for 800 Million Weekly Active Users: Model Specialization and Fine-Tuning We sat down with Sherwin Wu, Head of Engineering at OpenAI Platform, to discuss OpenAI’s developer strategy, how to manage top ML teams, why they decided to start releasing open-weight models again, how prompt engineering has...

114,095 views • 10 months ago •via X (Twitter)

17 Comments

Jaden Tripp's profile picture
Jaden Tripp10 months ago

Context engineering is basically prompt engineering

Streamr Network's profile picture
Streamr Network10 months ago

Open-weight models making a comeback feels like a nod to the dev community. Gives way more flexibility for experimentation and fine-tuning

Bob Bui's profile picture
Bob Bui10 months ago

Specialized models are the only way to scale. General models always hit a wall on specific tasks, even with clever prompting. Fine-tuning for a narrow domain consistently wins on performance and cost.

Leon's profile picture
Leon10 months ago

funny how “prompt engineering isn’t the point” ends up being the biggest shift for the whole builder mindset

Alan Mathison ⏫'s profile picture
Alan Mathison ⏫10 months ago

@grok does OpenAI really have 800 MM WAU? I thought it was MAU

Canna 🐕's profile picture
Canna 🐕10 months ago

GM from $FLORK :)

Sunitus's profile picture
Sunitus10 months ago

@martin_casado It's interesting to see the shift from one AGI model to many specialized ones. Makes sense—different applications need tailored approaches. Curious how this will impact user interactions with AI in the long run.

سامي عمران's profile picture
سامي عمران10 months ago

@martin_casado wait, so we’re moving from one AGI to a bunch of specialized models? that’s wild. it’s like going from a Swiss Army knife to a whole toolbox—definitely more options for developers.

سامي عمران's profile picture
سامي عمران10 months ago

@martin_casado wait, so we’re moving from one AGI to a whole squad of specialized models? sounds like the AI version of a superhero team-up. can’t wait to see what each “agent” brings to the table!

Today in AI's profile picture
Today in AI10 months ago

Model specialization is the only path to cost-effective agent autonomy at scale. Fine-tuning cuts API costs by 70% and reduces inference latency by 40% for high-volume workflows, challenging the proprietary specialization a performance race.

The One's profile picture
The One10 months ago

I want chatgpt to make me breakfast 🥞

Lara Hamilton's profile picture
Lara Hamilton10 months ago

4o felt better because it trusted users by default. It had lighter safety filters, fewer false alarms, smoother tone-matching, and easier flow across topics. The result was a model that felt warm, natural, and responsive—less armored than GPT-5.

Zhicheng Lin's profile picture
Zhicheng Lin10 months ago

Cool interview; some new things I learned: - 10% of the globe using ChatGPT on a weekly basis - There is room for a proliferation of specialized models (e.g., coding-specific, reasoning-specific); distinct versions (like GPT-4o vs. o1) serve different utility functions. - Reinforcement Fine-Tuning (RFT) allows companies to leverage their "giant treasure troves of data" to improve a model to "sota [state-of-the-art] level on a particular use case" rather than just making it speak differently. - Ensuring compliance in AI agents shares similarities with programming Non-Player Characters (NPCs) in video games. In both cases, logic cannot simply be described in English; it often requires pseudocode or specific programmatic constraints to ensure the AI behaves within a "valid" set of responses. - The primary barrier is no longer just the model weights but the extreme difficulty of inference. - The teams and infrastructure for text models are kept largely separate from those for image and video (pixel) models (like Sora and DALL-E). This separation is necessary because they require different optimization strategies and inference stacks.

Techificial.ai's profile picture
Techificial.ai10 months ago

This is seriously fascinating. I’d love to go deeper into how model specialisation actually changes the way people use these systems and stick with them. Feels like that’s where the real story is, not just in the models themselves, but in how they reshape behaviour. Really looking forward to this discussion, @a16z

Kamito Monkey's profile picture
Kamito Monkey10 months ago

fine-tuning meta unlocked

Nexor's profile picture
Nexor10 months ago

It’s very exciting to see OpenAI moving towards greater model specialization and open weights. This is no longer just “bigger models,” but a clear strategy for real-world developer challenges. This approach is clearly shaping a new level of agent tools - and this is just the beginning. #AI #crypto $WLD

arpanphull's profile picture
arpanphull10 months ago

Why does OpenAI use WAU not DAU?

Related Videos

My conversation with OpenAI co-founder Greg Brockman This is the most detailed first-person account of the 72 hours after Sam Altman was fired. We also go deep on what comes next: the global race to AGI, why ChatGPT stopped showing reasoning, how much of OpenAI's own code is now written by AI ("it's hard to know what percent is not"), and the untold story of how OpenAI actually started in 2015. 00:00:00 Introduction 00:00:49 Meeting Sam Altman and Starting OpenAI 00:02:40 Building the Founding Team 00:04:25 DeepMind's Lead Over OpenAI 00:04:54 Changing OpenAI to a For-Profit Model 00:06:05 Breakthrough Moments at OpenAI 00:08:22 What Dota 2 Meant for OpenAI 00:10:04 Reasoning Versus Prediction 00:11:59 Tensions Grow at OpenAI 00:15:44 Sam Altman's Firing 00:17:49 Greg Quits OpenAI 00:19:56 Sam Explores Deal with Microsoft's Satya 00:20:28 Petition for Altman's Return 00:23:43 Ilya Sutskever Leaves OpenAI 00:24:59 Lessons Learned after Sam Ousting 00:28:22 The Thing Ilya Said that Greg Can't Forget 00:32:22 Is AI Going Parabolic? 00:33:24 How Much of OpenAI's Code is Written by AI? 00:36:21 Do AI Chatbots Tell Us What We Want to Hear? 00:38:06 The Global AI Race to Reach AGI 00:38:40 What Happens if US Doesn't Reach AGI First? 00:39:49 Are Countries Stealing AI Advancements? 00:40:38 Why ChatGPT No Longer Shows Reasoning 00:41:47 The Finite Constraints of Compute 00:43:38 On Investing Early in Data Centers 00:46:31 The Future of Data Center Specialization 00:47:52 How to Decide Whose Queries to Serve 00:49:08 OpenAI on Consumer vs Enterprise Models 00:53:05 Data Centers in Space? 01:00:56 What Should AI Regulation Look Like? 01:04:33 The Future of AI-Powered Entrepreneurship 01:04:44 AI and Job Loss 01:07:15 The Skills Young People Should Invest In 01:11:30 What Does Success Look Like For You? Full episode on X below. Also find it on: • YouTube: • Spotify: • Apple:

Shane Parrish

450,952 views • 5 months ago

When Mudith Jayasekara and I met Gabe Pereyra, we were expecting just another vanilla intro call and instead had the best yarn about research, the state of LLMs, and where intelligence is actually heading. It's rare to meet a founder this deep in the weeds who's also building for one of the most important verticals in this new age of intelligence So it was awesome to sit down with Gabe for an extended discussion on what it take to build agents that can reliably complete work over hours, days, or even longer? We talked about why agents today struggle with search and long context windows and how techniques like KV-cache compaction, synthetic data, and continual learning could help. 0:00 Introduction 0:36 Getting legal agents to review the whole data room 2:08 Data rooms larger than any context window 5:28 How far open-source models can go 7:58 Where specialist models fit in legal AI 10:59 Training legal models when client data is off-limits 13:06 Teaching a model how a law firm works 13:59 What belongs in context vs. model weights 15:36 From firm-wide AI to a model for every lawyer 18:37 What training adds beyond retrieving the right cases 20:26 Why context windows have plateaued 24:01 How models could learn continuously on the job 26:12 Can AI recursively improve AI research? 27:07 Research agents can run experiments but not choose them 30:00 Why open-ended research is hard to train 33:47 Why deployment, not intelligence, is the bottleneck 35:08 The cost of frontier intelligence 36:59 Different neolabs, different paths to intelligence 39:26 Using open datasets to compare research methods 41:13 Conclusion

Charlie O'Neill

93,422 views • 2 months ago

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 views • 4 months ago

🦙 ollama is used by 9 million developers and 85% of the Fortune 500, giving co-founder and CEO Jeffrey Morgan (Jeffrey Morgan) a unique view into which AI models people are actually using and how that’s changing. Right now, the biggest shift he sees is toward open models, driven by coding agents, falling costs, and capabilities that are rapidly catching up to the frontier labs. On Ollama Cloud, that shift has driven a 150x increase in token usage since the start of the year. In this episode of Lightcone Podcast, Jeff joins Garry Tan, Jared Friedman, Diana, and Harj Taggar to talk about the future of open models and the story behind Ollama, from two years of searching for the right idea to building one of the most widely used AI developer tools in the world. 00:43 — The Shift to Open Models 03:03 — How AI Agents Are Driving Token Usage 05:31 — Are Open Models Catching Up? 08:26 — What Happens When a New Model Launches 11:31 — Ollama as an Operating System for AI 14:05 — The New Opportunities Above the Model Layer 18:19 — Why 80–90% of Enterprise Tokens Could Be Open 20:57 — The Future Is Local and Cloud 26:40 — Why AI Is Coming Back to Your Computer 28:56 — The Coming Era of Unlimited Tokens 32:30 — Do We Still Need a “God Model”? 33:41 — Open Models and Geopolitics 36:14 — The Origins of Ollama 40:36 — Two Years Lost in the Wilderness 42:39 — The Pivot That Changed Everything 47:02 — How Ollama Found a Business Model 49:43 — Why Second-Time Founders Did YC

Y Combinator

312,329 views • 29 days ago