Loading video...

Video Failed to Load

Go Home

Gavin Baker on the one variable that now dominates all others in AI compute: "the more memory you put with flop for a given unit of compute, the more tokens you get out -- it's the single most important thing you could do to increase token output per unit...

18,521 views • 4 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 views • 1 month ago

The creator of High Bandwidth Memory (HBM) put a number on the AI build that should stop every infra investor cold. A cluster of a million GPUs runs at roughly 10-20% utilization (Save this). Kim Jung-ho spent thirty years building what feeds the GPU, and his claim is that the GPU is barely working. Here is what is actually happening. Every time a model generates output, the data has to be read out of memory, computed, and written back. The read and the write swallow almost the entire cycle. While that data moves, the GPU does nothing. It sits there, fully powered, fully paid for, waiting. By Kim's estimate the memory is doing only about 30 percent of the work it needs to do. The processor idles the rest. So a million installed GPUs run at 10 to 20 percent. You are not compute constrained. You are memory constrained, and the expensive part is standing around. Adding more GPUs does not fix this. It gives you more processors starving for the same data. Here is the part that decides the next decade. Memory can grow. When a cell cannot shrink any further, you stack it into a high-rise, layer on layer. A GPU cannot be stacked. It runs too hot and needs a cooler bolted to its back, so the one move that rescues memory is closed to the processor. The thing that can keep stacking compounds. The thing that cannot plateaus. The marginal dollar in an AI build now buys more by fixing the memory path than by bolting on another idle GPU. Which is why the companies that control memory bandwidth and supply are not suppliers to the AI trade. They are the AI trade.

Fireside Alpha

38,370 views • 1 month ago

GAVIN BAKER: NVIDIA IS NO LONGER JUST A CHIP COMPANY. Gavin Baker recently laid out one of the sharpest observations about the current AI hardware cycle, and it explains why Nvidia's multiple looks strange relative to the strategic position the company actually occupies. The starting point is a shift in how semiconductor supply is being allocated. Historically, buyers with enormous volume like Apple could break long-term agreements with suppliers whenever they wanted. There was no meaningful consequence, because the supplier had nowhere else to go. That world is over. In the current cycle, there are at least four major memory buyers with genuine scale, plus a wave of AI startups adding to demand. The dynamic has flipped. If a hyperscaler breaks an LTA on price, the supplier can now retaliate by reallocating volume to a competitor. In a cyclical industry where oversupply is always followed by undersupply, breaking an LTA today can cost you your entire allocation the next time capacity gets tight. That leverage is one reason Nvidia's position is so unusual. Every constraint in the AI buildout tilts in its favor. Nothing on earth is more financeable than an Nvidia GPU. And on land and power, Nvidia is playing the matchmaking chess game between customers, energy suppliers, and infrastructure partners. The most interesting development Baker highlighted is Nvidia's new business model. He described it as a credit wrapper with a revenue share above a certain GPU price floor. Nvidia is financing customer deployments while retaining upside if GPU prices stay elevated. This converts Nvidia from a pure hardware vendor into something structurally closer to a cloud franchise via royalties on the compute being deployed on its chips. Nvidia could end up with an effectively massive cloud business very quickly, not through building data centers but through collecting royalties on the compute those data centers produce. The market is still pricing $NVDA as a semiconductor company. The company itself has already moved beyond that model. The gap between the two will eventually close. Invest Like the Best Patrick OShaughnessy Gavin Baker

Lumida Wealth Management

184,590 views • 4 days ago

The AI boom just hit a wall nobody saw coming. And it's not software. It's not regulation. It's not even energy... It's memory chips. Right now, Dell is raising PC prices by 30%. Intel can't ship chips. Nvidia is slashing GPU production by 40%. And almost nobody understands why. Here's the "hidden" crisis the AI industry is trying to hide: AI data centers are hoarding memory. Not GPUs. Not processors. MEMORY. Every AI server needs massive amounts of high-bandwidth memory (HBM) to run those models everyone's hyping. One problem: There are only 3 companies in the world that can make it. Samsung. SK Hynix. Micron. That's it. And all 3 just diverted their entire production capacity away from normal RAM to feed AI data centers. The math that breaks everything: 1 gigabyte of HBM takes 4X the manufacturing capacity of regular DRAM. AI will consume 20% of global DRAM production in 2026. But the thing is, consumer demand for RAM didn't disappear. PCs still need memory. Phones still need memory. Cars still need memory. But there's no capacity left to make it. The price explosion: RAM prices are up 246% in the last 6 months. DDR5 contract prices jumped 100% month-over-month in some cases. Dell's CFO said he's "never witnessed costs escalating at this pace." SK Hynix and Micron? Sold out through all of 2026. Micron straight up EXITED the consumer memory market entirely to focus on AI customers. If you're not building an AI data center, you're not getting memory chips. AI data centers pay 3-5X margins compared to consumer products. So memory manufacturers are rationally choosing: Serve Microsoft and Google's AI buildout, or serve Dell's laptop business? Easy choice. Every wafer allocated to an Nvidia H100 GPU is a wafer DENIED to your next laptop. It's a zero-sum game. And consumers are losing. The dangerous cascade effect: Nvidia is cutting RTX 50-series GPU production by 30-40% because they can't get GDDR7 memory. Dell, Lenovo, HP are all raising PC prices 15-30% in early 2026. Xiaomi and other smartphone makers are cutting shipment targets. Even Intel's crash last week? Partially driven by memory shortages limiting chip production. This is a PERMANENT reallocation of the world's silicon capacity. Not a temporary supply hiccup. For decades, consumer electronics (phones, PCs, laptops) drove memory production. Now? AI data centers are the priority customer. And that priority shift is reshaping the entire tech economy. The timeline Is worse than you think: Industry analysts project shortages lasting through 2027, maybe 2028. Why? Because building new memory fabs takes 3-5 YEARS. Micron's new Idaho fab won't meaningfully impact supply until 2028. Samsung and SK Hynix are too busy ramping up HBM4 production to expand consumer DRAM. So we're stuck. AI companies need memory to scale. But producing that memory DESTROYS the supply chain for everything else. My question here: Everyone's betting on AI scaling infinitely. But what if the AI boom STALLS because there's not enough memory to support it? What if we're not in an "AI supercycle" but a "memory shortage that kills the AI buildout"? Intel crashed 17% because they can't manufacture enough chips. The root cause though? Memory shortages limiting what they can even produce. Nvidia is cutting GPU production by 40%. AMD is struggling to get GDDR6 for Radeon cards. This isn't just a consumer problem. It's an AI infrastructure problem. And if memory doesn't scale, AI doesn't scale. The AI industry sold you on infinite scaling. But they forgot to mention the part where there's only 3 companies making the memory chips that power everything. And all 3 just chose AI data centers over you. Even Nvidia can't make enough GPUs to meet demand. Not because of energy. Not because of regulation... But because the memory supply chain is BROKEN. And it won't be fixed until 2028.

Ricardo

594,643 views • 6 months ago

Jensen Huang just replaced the most important metric in global economics. Not trade volume. Not oil output. Not manufacturing. Compute. Huang: “Compute equals GDP. I know that for certain.” He did not say probably. He said certain. If your nation does not produce compute, it does not produce intelligence. If it does not produce intelligence, it does not produce revenue. Two links in the chain. Miss one and the whole thing breaks. Huang: “Not one country in the future will say, ‘Guess what, we’re gonna opt out on intelligence.’” Because opting out of compute is not a strategic decision. It is an extinction schedule. Every country that does not build its own inference capacity becomes a tenant in someone else’s infrastructure. Not an ally. Not a partner. A dependent. And dependents do not negotiate terms. They accept them. But this is not just a story about nations. Huang: “The entire software industry will be token-driven.” Every product. Every platform. Every service you touch. The entire business model of software is about to be measured in tokens consumed. Not seats sold. Not licenses renewed. Tokens burned. Software used to be a thing you bought. Now it is a thing that thinks. And thinking costs compute. Every query. Every action. Every decision the machine makes on your behalf. The meter is always running. Huang: “The entire internet industry could take 100% of their CapEx and make it AI because it’s better.” Not ten percent. Not a pilot program. One hundred percent. The moment any internet service rebuilds itself on generative intelligence, it outperforms every version that came before it. Search. Ads. Recommendation. Infrastructure. All of it. Better on contact. CapEx follows. All of it. Trillions moving in one direction with no offramp. The companies still budgeting AI as a line item are telling you exactly how much they understand. AI is not the line item. AI is the budget. The global economy is being re-denominated in a currency most people have not even heard of yet. Tokens. Whoever controls the supply of that currency is not playing in the new economy. They are the house. And the house does not lose.

Dustin

43,505 views • 4 months ago