
Julia Turc
@juliarturc • 24,161 subscribers
Explaining AI on YouTube • YC S24 Founder • Ex-Google Research • Eastern-European nihilist & American optimist
Videos

How far are we from "Her"? I couldn't help but wonder, given the rapid fire of GPT-Live and Thinking Machines. Turns out, full duplex speech-to-speech is not particularly new. Moshi has been around for more than 2 years. Many thanks to Neil Zeghidour for bringing so much nuance. Check out kyutai who are doing god's work and open-sourcing it (not sponsored!). 00:00 Intro 01:24 ChatGPT Standard & LLM-based cascades 04:08 ChatGPT Advanced & end-to-end speech models 05:35 ChatGPT Live & Her 07:38 Thinking Machines & multi-stream modeling 09:13 Moshi & full duplex assistants 13:24 Moshi’s inner monologue 15:38 Why aren’t full duplex models everywhere? 17:08 Tool calling in full duplex models 18:00 Moshi RAG 19:06 Is the future full duplex? 20:24 How far are we from Her?
Julia Turc70,096 views • 9 days ago

This is what happens when you plug LLMs into voice assistants, instead of a decade of handwritten rules. This video dissects Voxtral (a family of OSS speech models) and the foundational work behind it (audio tokenization, semantic/acoustic disaggregation, etc). Thank you Mistral AI for your collaboration and for your detailed technical reports in an increasingly opaque industry! 00:00 Intro 01:03 Modular vs end-to-end speech models 03:30 Speech-to-Text 06:07 Delayed Streams Modeling (DSM) 09:41 Whisper Streaming 10:33 Voxtral Realtime 13:07 Voxtral Text-to-Speech 14:28 Throwback: WaveNet 15:24 Audio tokenization 20:39 The Voxtral Codec 21:49 Back to Voxtral TTS 25:30 Outro
Julia Turc297,055 views • 1 month ago

"World models" is one of the buzziest yet ambiguous terms in AI right now. I started this video with many questions: - How are they different from video generation? - Can they do more than AI slop? - Can LeCun be trusted given that he wears knee-high white socks? Many thanks to TJ Galda and NVIDIA AI for helping me answer (most) of these questions!
Julia Turc248,035 views • 2 months ago

This is why we need to gate-keep science. Two pseudo-intellectuals thinking they discovered something deep, conflating >social attention >transformers attention >quantum physics observer (attention) These have nothing in common, other than the ambiguity of English language. Naked ladies on Instagram have nothing to do with a weighted average followed by softmax. But they're both so mind-blown by their discovery. Dunning–Kruger will only get amplified by AI sycophancy. Please call me out if you see me going beyond my own DK threshold.
Julia Turc119,551 views • 2 months ago

Early on, Jonathan Ross foresaw the unprecedented demand for compute of the 2020s. After pioneering Google's TPUs, he founded Groq, a "Language" Processing Unit dedicated for inference. After joining NVIDIA, he is betting on the GPU+LPU combo to power the agentic chains of AI calling AI calling AI. This interview is part of my larger series on dedicated AI chips, so drop any suggestions/questions below. Thank you NVIDIA for hosting us at your HQ! 00:00 Intro 01:00 The Google TPU & Groq origin story 02:41 How is Groq different from a GPU? 04:43 Static scheduling makes Groq faster 05:47 Does Groq work with Mixture-of-Experts? 09:27 Are LPUs limited to text models? 11:03 Diffusion models 13:41 NVIDIA Vera Rubin: the GPU+LPU combo 15:26 Will Groq still be sold as a standalone chip? 16:49 How does agentic AI impact inference economics? 19:21 Will AI replace CUDA kernel engineers? 21:08 Will AI democratize hardware design? 26:11 Jevon's paradox: An endless demand for compute 29:04 What should kids learn in the AI age?
Julia Turc46,967 views • 1 month ago

Diffusion models clicked for me when I started seeing them through the lens of particle motion. I built this interactive playground where you too can clickety-clack to understand how drift, noise, and other hyperparams control diffusion. I hereby submit this as penance for the sin of YouTube edu-tainment 😇 Link in the first comment.
Julia Turc32,207 views • 4 months ago

An H100 GPU can run 10^15 FLOPs/s. But somehow, a model like Gemini Flash-Lite is stuck at 200 tokens/s. These numbers show just how much of a bottleneck memory bandwidth is. In this video, we look at how LLMs are evolving to work around this limitation. Thanks Inception for sharing your insights!
Julia Turc26,110 views • 5 months ago

Diffusion clicked for me when I read about score-based models, a line of work pioneered by Stefano Ermon (et al.) at Stanford. So it was a full-circle moment to collab with him and Inception on a video about training & sampling techniques for making diffusion LLMs faster.
Julia Turc27,186 views • 5 months ago

Just published a new Flow Matching tutorial (link below). It's a little unorthodox, so I want to hear your feedback before turning it into a video. On the surface, the FM algorithm looks totally arbitrary: 🫤Pick two arbitrary points 🫤(Arbitrarily) decide to draw a straight line 🫤(Arbitrarily) model particle moving at constant speed 🫤Pick an arbitrary point in time 🫤Evaluate the velocity and use it as training label 🤯Mathematically totally valid Hopefully this tutorial will turn all of these arbitrary-looking things into intuition.
Julia Turc17,970 views • 3 months ago

When reading diffusion papers, my most common reaction is "mkay the math works out but WHY WOULD A SANE PERSON CHOOSE TO DO THIS". This included the (foundational) DDPM formulas. The saving grace is that you can visualize it as 2D particle motion and get a solid intuition.
Julia Turc20,528 views • 4 months ago
No more content to load