
Latent.Space
@latentspacepod • 30,804 subscribers
The #1 AI Engineering podcast & newsletter, now covering AI for Science as well. Over 170,000 daily readers. Technical news today you will use at work tomorrow!
Videos

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out to unlock more throughput, why inference teams are still finding 20–200% performance gains, how video generation runs into a quadratic attention wall, and how GLM-5.2 helped rewrite and optimize the GPU kernels serving GLM-5.2 itself.
Latent.Space137,910 Aufrufe • vor 3 Tagen

Gray Swan: Red-Teaming after Mythos & the coming AI security crisis Gray Swan AI cofounders Zico Kolter and Matt Fredrikson explain why AI security is fundamentally different from traditional cybersecurity, how their automated red-teaming system Shade can now outperform humans at breaking frontier models, why prompt injection creates an entirely new class of exploits for agents like Claude Code and Codex, and why the first major AI security breach may be a gray swan event that everyone sees coming.
Latent.Space1,763,741 Aufrufe • vor 1 Monat

From rewriting Google’s search stack in the early 2000s to reviving sparse trillion-parameter models and co-designing TPUs with frontier ML research, Jeff Dean has quietly shaped nearly every layer of the modern AI stack. As Chief AI Scientist at Google and a driving force behind Gemini, Jeff has lived through multiple scaling revolutions from CPUs and sharded indices to multimodal models that reason across text, video, and code. We sat down with Jeff to unpack what it really means to “own the Pareto frontier,” why distillation is the quiet force behind every generation of faster, cheaper models, how energy not FLOPs is becoming the true constraint on AI compute, what it takes to co-design hardware and models 2–6 years into the future, why unified multimodal systems will outperform specialized ones, what it was like leading the charge to unify all of Google’s AI teams, and his prediction that deeply personalized models with access to your full digital context will redefine what useful AI looks like. Jeff Dean Google DeepMind Google
Latent.Space529,014 Aufrufe • vor 5 Monaten

🆕 Marc Andreessen’s 2026 AI Thesis: Agents, Open Source, and Why This Time Is Different Marc Andreessen 🇺🇸 of a16z says AI people keep swinging between utopian and apocalyptic for one simple reason: this field has been “almost here” for 80 years. But now, the breakthroughs are no longer theoretical. Reasoning, coding, agents, and self-improvement are all starting to work at once. This episode goes deep on AI winters, OpenAI + OpenClaw, infrastructure overbuild risk, proof-of-human, why software may soon be written mostly for bots, and why the real bottleneck may be society adopting AI rather than the models improving.
Latent.Space351,146 Aufrufe • vor 4 Monaten

Poolside’s Race to AGI: Model Factory, Laguna S, Open Models, & Beyond MCP Poolside co-founder and CEO Eiso Kant explains why coding is the path to AGI, how Poolside’s Model Factory runs 10,000–20,000 experiments a month and turns new models around in as little as eight weeks, why Laguna S suggests persistence and verification may matter more than raw intelligence, why agents are already helping build the next generation of models, and why the future should have 100 foundation model companies not five.
Latent.Space32,734 Aufrufe • vor 14 Tagen

In this episode, OpenAI Chief Research Officer Mark Chen joins allen to flambé shrimp, cook Korean stew, and chat about being at the frontier of AI research: why scaling laws and pre-training still matter, how OpenAI chooses research bets and allocates compute, what it means to develop research taste, why evals are in crisis, how to avoid benchmark-maxing, and what it will take for models to handle long-horizon real-world work, multimodal reasoning, and eventually end-to-end AI research. Timestamps: 0:00 Intro 0:28 The Soup Story 1:52 From Trading to AI Research 3:21 How to Develop Research Taste 5:23 RL, Evals, and Superhuman Benchmarks 8:17 Cooking Begins on the Impulse Stove 8:53 Scaling Laws, Pre-Training, and Reasoning 12:33 OpenAI’s Research Roadmap and Compute Allocation 15:48 What Makes a Great Researcher 19:33 The Evals Crisis and Benchmark-Maxing 24:34 Jagged Intelligence, Context, and Long-Horizon Learning 27:14 Shrimp Flambé and New Research Bets 31:32 Multimodal Models and One Architecture 32:36 Vibe Researching and End-to-End AI Research 34:36 Failed Bets, Postmortems, and OpenAI’s Alpha 37:07 Final Taste Test 37:53 Overrated vs. Underrated AI Research 41:00 Closing
Latent.Space62,796 Aufrufe • vor 1 Monat

Why the Frontier Ecosystem must be Open — Matei Zaharia and Reynold Xin, Databricks Databricks cofounders Matei Zaharia and Reynold Xin explain why Databricks is moving into the infrastructure layer for enterprise agents, how Omnigent creates a shared harness for coding agents and custom agents, why LTAP and Lakebase rethink the split between operational and analytical databases, why agent security needs contextual policies and spend controls, and why the future of software may be as simple as getting the right data in place and putting agents on top.
Latent.Space61,250 Aufrufe • vor 1 Monat

In this episode, Engram co-founder and CEO Dan Biderman joins allen to cook Mediterranean meatballs with yellow rice and talk about building AI that actually learns from you: why long context, RAG, and compaction eventually break down, how Engram compresses knowledge into cartridges and model weights, what continual learning could unlock for long-horizon agents, why token efficiency is inseparable from intelligence, how personal models could improve like Tamagotchis, and what it takes to build the research and infrastructure for millions of continuously updated AI memories. Timestamps: 0:00 Intro 0:26 Engram’s $98M Launch and Meatballs 1:45 From Naval Special Operations to AI Research 4:32 Israeli Military Culture and Founder Maturity 7:12 Why Engram Is Betting on Context and Continual Learning 9:14 Knowledge Cartridges, Compression, and Model Intuition 14:10 Trillion-Token Company Knowledge and Context Rot 18:05 Long-Context Limits, Compaction, and Neural Memory 22:20 Test-Time Training and “Destroying Prefill” 24:31 Harvey and Holistic Enterprise Queries Beyond RAG 27:02 Personal AI Models and Tamagotchi Weights 30:00 What Belongs in Weights vs. Text 32:25 Autonomous Memory and User-Specific Feedback Loops 34:20 Token Efficiency, Model Routing, and Harder Tasks 38:03 Engram’s Research Team and Product Culture 43:02 Hiring Researchers and Infrastructure Engineers 45:25 Doing More With Less 47:41 Where to Find Engram 48:19 Final Taste Test
Latent.Space33,729 Aufrufe • vor 24 Tagen

From early engineer at Glean to Partner at Menlo Ventures co-leading the Anthology Fund with Anthropic, Deedy Das has gone from shipping “boring” enterprise search to backing frontier labs, infra, and research plays like Anthropic, OpenRouter, Goodfire, Prime Intellect, and Whisper. We sat down with Deedy to unpack how Glean quietly built a real AI moat before LLMs were cool, why enterprise search is significantly different than consumer search, where value actually accrues in the model vs app stack, Anthropic’s rise and the importance of Claude Code, the Anthology Fund thesis, OpenRouter’s wedge, mechanistic interpretability, the coming compute arms race, coding agents and “LLM psychosis,” and what all of this means for AI engineers deciding what to build. Deedy swyx 🔜 NeurIPS + #DevWritersRetreat Alessio Fanelli
Latent.Space269,074 Aufrufe • vor 8 Monaten

"Projects like the New Deal, the Apollo program pale in comparison to what we're doing right now." 🆕 Greg Brockman (Greg Brockman) joins us to talk GPT-5, GPT-OSS, and what's next on OpenAI's road to crystallizing all of human intelligence! “Energy turns into compute, turns into intelligence… crystallizing compute into potential energy you can release again and again.” 0:00:04 - Introductions 0:01:04 - The Evolution of Reasoning at OpenAI 0:04:01 - Online vs Offline Learning in Language Models 0:06:44 - Sample Efficiency and Human Curation in Reinforcement Learning 0:08:16 - Scaling Compute and Supercritical Learning 0:13:21 - Wall clock time limitations in RL and real-world interactions 0:16:34 - Experience with ARC Institute and DNA neural networks 0:19:33 - Defining the GPT-5 Era 0:22:46 - Evaluating Model Intelligence and Task Difficulty 0:25:06 - Practical Advice for Developers Using GPT-5 0:31:48 - Model Specs 0:37:21 - Challenges in RL Preferences (e.g., try/catch) 0:39:13 - Model Routing and Hybrid Architectures in GPT-5 0:43:58 - GPT-5 pricing and compute efficiency improvements 0:46:04 - Self-Improving Coding Agents and Tool Usage 0:49:11 - On-Device Models and Local vs Remote Agent Systems 0:51:34 - Engineering at OpenAI and Leveraging LLMs 0:54:16 - Structuring Codebases and Teams for AI Optimization 0:55:27 - The Value of Engineers in the Age of AGI 0:58:42 - Current state of AI research and lab diversity 1:01:11 - OpenAI’s Prioritization and Focus Areas 1:03:05 - Advice for Founders - It's Not Too Late 1:04:20 - Future outlook and closing thoughts 1:04:33 - Time Capsule to 2045 - Future of Compute and Abundance 1:07:07 - Time Capsule to 2005 - More Problems Will Emerge
Latent.Space305,090 Aufrufe • vor 11 Monaten

Priscilla Chan and Mark Zuckerberg co-founded the Chan Zuckerberg Initiative (CZI) in 2015, committing 99% of their Meta shares to advance science, education, and opportunity. As a pediatrician and CEO of Meta respectively, they've built CZI into one of the most ambitious technology-driven philanthropic organizations with a moonshot goal to help cure, prevent, or manage all diseases by the end of the century. In this episode, we sit down with Priscilla and Mark to unpack how CZI bridges state-of-the-art AI with open biomedical research, their investments in open-source tools like the Human Cell Atlas, their philosophy on building teams across different disciplines, and what they envision the future of healthcare. Chan Zuckerberg Initiative swyx Alessio Fanelli
Latent.Space166,190 Aufrufe • vor 9 Monaten

Noam Brown from OpenAI just dropped a truth bomb: "Your fancy AI scaffolds will be washed away by scale" Routers, harnesses, complex agentic systems... all getting replaced by models that just work better out of the box The reasoning models already proved this
Latent.Space208,301 Aufrufe • vor 1 Jahr

From applied cryptography and offensive security in France’s defense industry to optimizing nuclear submarine workflows, then selling his e-signature startup to Docusign and now running AI as CTO of Superhuman Mail (Superhuman, recently acquired by Grammarly), Loïc Houssier has lived the full arc from deep infra and compliance hell to obsessing over 100ms product experiences and AI-native email. We sat down with Loïc to dig into how you actually put AI into an inbox without adding latency, why Superhuman leans so hard into agentic search and “Ask AI” over your entire email history, how they design tools vs. agents and fight agent laziness, what box-priced inference and local-first caching mean for cost and reliability, and his bet that your inbox will power your future AI EA while AI massively widens the gap between engineers with real fundamentals and those faking it. Ordinary Extrare Superhuman swyx Alessio Fanelli
Latent.Space126,794 Aufrufe • vor 7 Monaten

Long Live Outputmaxxing: AI compute grids, Anthropic’s coding takeoff, data center backlash, & frontier systems AMP PBC founder Anjney Midha explains why 95% GPU utilization was considered an outage at Google, why the AI race is no longer just about buying more GPUs, how AMP is trying to make FLOPs flow like megawatts, why data center backlash could become one of AI’s biggest bottlenecks, how Anthropic cracked coding through culture and preparation, why DeepMind research hoarding creates a market failure, and why the next frontier may belong to teams that can “output max” across compute, capital, culture, and science.
Latent.Space30,621 Aufrufe • vor 1 Monat

Abridge: 100M+ medical conversations, real-time prior auth, and the clinical intelligence layer Abridge is building the clinical intelligence layer for healthcare. In this episode, Janie Lee and Chaitanya Asawa explain why ambient documentation was only the first wedge, how Abridge is turning patient conversations into real-time clinical decision support, why healthcare may become one of AI’s most important proving grounds, and how 100M+ medical conversations, specialty-specific evals, and deep EHR integrations create a moat for AI-native healthcare.
Latent.Space42,612 Aufrufe • vor 2 Monaten

🆕 Claude Cowork, Skills, and the Future of AI Coworkers Felix Rieseberg has spent years working at the interface layer, from Electron and the Slack desktop app to now helping build Claude Cowork. In this episode, Felix explains why execution is getting so cheap that teams can “build all the candidates,” why Anthropic is betting on local-first agent workflows, and why the future of AI products may belong less to chatbots and more to systems that can actually do knowledge work.
Latent.Space65,812 Aufrufe • vor 4 Monaten

For our first episode of our new show In-Context Cooking, we have the Founder & CEO of SemiAnalysis Dylan Patel. We talk about: • Taiwan endgame scenarios & TSMC risk • AI export controls + Chinese talent flight • $180–200B hyperscaler capex (is this a bubble?) • Nvidia vs vertical integration • What actually bottlenecks AI (power? fabs? chips?) • Why the public might turn anti-AI Dylan Patel SemiAnalysis allen
Latent.Space66,930 Aufrufe • vor 5 Monaten

Modal's Agent-Native Cloud: DX→AX, sandboxes, elastic inference, and 100,000 rollouts Modal CTO Akshat Bubna explains why developer experience is becoming agent experience, why agents need infra they can operate instead of YAML they have to reason through, how sandboxes turn the agent loop into something real, why elastic inference and GPU snapshotting matter for production AI, how RL rollouts can require 100,000 sandboxes, and why Modal’s $355M Series C marks a new phase for AI-native cloud infrastructure.
Latent.Space15,359 Aufrufe • vor 28 Tagen

From scaling startups in sales and public cloud (including building Palo Alto Networks’ central U.S. cloud business) to joining Kleiner Perkins to help technical founders turn product edge into repeatable revenue, Joubin Mirzadegan built a career around one obsession: distribution. That obsession led him to start the podcast Grit as a hiring wedge, work alongside breakout companies like Glean and Windsurf, and now incubate Roadrunner which is an AI-native rethink of CPQ as SaaS pricing explodes from “seats” into consumption, bundles, renewals, and SKU sprawl. We sat down with Joubin to dig into the behind-the-scenes craft of great interviews (i.e. why he never sends questions), what Windsurf got right about “Google-class product + Salesforce-class distribution,” how to hire early sales leaders without getting fooled by shiny logos, why legacy CPQ data models are quietly breaking modern revenue teams, and his bet that rebuilding quoting and approvals from the ground up then layering LLMs on top will eliminate the deal-desk Slack chaos and become the backbone for how enterprise AI companies sell in the next decade. Joubin Mirzadegan swyx Roadrunner Kleiner Perkins
Latent.Space77,604 Aufrufe • vor 7 Monaten

🆕Scaling Test Time Compute to Multi-Agent Civilizations, with Noam Brown We're excited to publish our full conversation with Noam Brown on the frontiers of the new reasoning paradigm at OpenAI! - first principles for starting the "Multi-Agents" team - what's not captured by the "System 1/System 2" analogy for inference time compute - how Ilya Sutskever convinced him that reasoning was closer than he thought - Deep Research is existence proof that RL generalizes beyond verifiable rewards - the relationship between AI for imperfect information games (like Poker, Stratego, Diplomacy) and reasoning Enjoy! on youtube, or wherever fine podcasts are sold.
Latent.Space105,923 Aufrufe • vor 1 Jahr