Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Microsoft sold every spare CPU it had to Anthropic and OpenAI. Amazon tripled its CPU buys year over year and still can't keep up. Two of AWS's biggest customers asked Andy Jassy if they could buy the entire 2026 production run of Graviton chips. He said no. The ratio...

290,767 Aufrufe • vor 4 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Jensen Huang just identified the next $200 billion market (Save this). The shift starts with a observation about agentic AI that changes everything about infrastructure. In the era of training and inference, the GPU was everything while CPU was a traffic cop, scheduling work, managing memory, dispatching tasks while the GPU did the heavy lifting. Agentic AI breaks that model entirely. An AI agent does not just run a single inference pass but rather it plans, calls tools, executes code in sandboxes, retrieves data from multiple sources and loops through complex multi-step reasoning sequences often thousands of times per second at scale. Every one of those operations runs through the CPU and the GPU sits idle waiting for the CPU to prepare the next task, supply the right context and execute the retrieval and tool calling logic fast enough to keep the accelerators fed. The CPU is now the conductor and the GPU is the orchestra and the bottleneck is the conductor falling behind. This is showing up in production AI factory utilization right now, which is exactly why Jensen built Vera from scratch rather than licensing x86. Vera achieves 40% lower peak memory latency than x86, 50% faster core to core communication, and 1.8 times the agentic sandbox performance of current x86 processors on a purpose-built architecture designed around the agentic loop. Now here is where the investment thesis gets interesting. The obvious beneficiary is Nvidia itself, and that thesis is real. Nvidia's CFO has guided for nearly $20 billion in Vera CPU revenue this fiscal year alone, a market Nvidia had zero presence in just three years ago. Intel held 60% of server CPU market share as recently as Q4 2025 and that transition is now happening at a pace Intel structurally cannot respond to. But the deeper question is, what architecture is Vera actually built on? Vera's Olympus cores are ARM compatible and every single Vera CPU deployed in every Vera Rubin rack in every data center in the world runs on ARM architecture. And ARM Holdings collects a royalty on every one of them. ARM does not make chips but rather licenses the instruction set architecture and CPU core designs that others build on top of. Every time Nvidia ships a Vera CPU, every time a hyperscaler deploys a Vera Rubin rack, every time an enterprise qualifies Vera for their AI factory, ARM earns a royalty. The secular tailwind here is almost perfectly constructed for ARM's business model. Amazon's Graviton, Microsoft's Cobalt, Google's Axion, Apple's silicon stack, and Qualcomm's data center push all run on ARM. And now Nvidia's Vera, which is projected to displace Intel as the largest server CPU supplier by revenue in a single fiscal year, is ARM. ARM's royalty rate on high end server chips is estimated at roughly 1 to 2% of chip selling price. At $5,000 per Vera CPU and 4 million units projected for FY2027, that is a royalty line growing from near zero to potentially $400 million to $800 million annually from Nvidia's data center CPU business alone before counting Amazon, Microsoft, Google, Apple, and Qualcomm. The total ARM addressable royalty base across all the silicon it already licenses is compounding at a rate that the current $130 billion market cap does not fully reflect. Jensen's CPU thesis is the most underappreciated catalyst in ARM's fundamental story, and the royalty compounding has barely started. Come join Milk Road Pro and get our full ARM royalty model and our entire AI trade thesis. Link below!

Milk Road AI

11,819 Aufrufe • vor 2 Monaten

Cathie Wood just flagged the sleeper trade inside the AI boom that most people are completely missing. Everyone has been chasing GPUs. Nvidia, the data center buildout, the chip arms race. That trade has been obvious for two years. But OpenAI's CFO Sarah Fryer said something quite different: people are going to be really shocked by how agentic AI activates CPUs. Right now, for every CPU in an AI workload, there are 4 to 5 GPUs. That's the current ratio. Wood thinks that ratio is going to 1 to 1. Think about what that means. AI inference at scale, agents running autonomously, pipelines executing tasks across systems. The compute mix shifts dramatically away from pure GPU dominance. CPUs become a first-class citizen in the AI stack. Cathie called it going "back to the future." Intel has taken off. Flex (formerly Flextronics) is booming. Stocks that were giants in the dot-com bubble are resurging because the underlying demand for their products is real again. The GPU trade made sense at the training stage. You need massive parallel compute to train frontier models. But agentic AI runs differently. Agents are constantly orchestrating, reasoning, calling APIs, executing workflows. That workload looks a lot more like traditional computing. And traditional computing runs on CPUs. If Cathie Wood is right about the ratio collapsing to 1:1, the CPU demand signal embedded in the AI buildout is orders of magnitude larger than the market is currently pricing.

Milk Road AI

234,897 Aufrufe • vor 3 Monaten

Rene Haas just confirmed the Vera CPU thesis on yesterday’s Arm Q4 call. He didn’t mean to His framing: GPUs are reticle-limited. CPUs are not. The ratio shift is happening in core count, not chip count His exact words: “256 Vera CPU chips, 88 cores per chip, a 200-kilowatt liquid-cooled rack designed to sit in a data center adjacent to a Vera Rubin system” That is not a host CPU. That is a dedicated agentic orchestration Two days ago NVIDIA’s own engineers published the receipt. They traced a real 33-minute Claude Code session: 283 inference requests 58 main-agent turns coordinating 225 sub-agent invocations Context grew from 15K to 156K tokens before compaction dropped it to 20K Main agent alone processed ~3.5 million input tokens in the first 40 turns Anthropic’s own number: agentic systems consume up to 15x more tokens than chat. Coding agents sustain 95 to 98 percent prompt cache hit rates. Without caching, costs would be 6x higher This is what’s happening between GPU calls. File reads. Tool invocations. Sub-agent spawns. Compaction. KV cache management. None of it runs on the GPU That’s why 12,000 GPUs need 400,000 CPU cores. The 33-to-1 ratio isn’t a forecast. It’s a measurement NVIDIA states it in the blog directly: this won’t be resolved by adding more compute FLOPs and memory capacity Translation: the GPU-only path is exhausted. The agentic chapter requires a platform, not a chip Their seven-chip answer: Vera Rubin NVL72 —capacity and prefill Vera CPU — tool execution, KV cache offload Groq 3 LPX — SRAM-first decode, low-jitter generation NVLink 6, ConnectX-9, BlueField-4, Spectrum-X — fabric Result they claim: 400+ tokens per second per user on trillion-parameter MoE at 400K context. Vera spec: 88 Olympus cores, 176 threads, 1.8 TB/s NVLink-C2C, 1.2 TB/s LPDDR5X, 227 billion transistors. A 256-CPU rack delivers 45,056 threads and 400 TB of memory One detail nobody is talking about. The blog’s second author was previously Head of Agents at Groq. The third was previously at Groq Inc and Intel. NVIDIA didn’t license the LPX architecture. They absorbed the team that built it Haas isn’t pitching a competing thesis. He’s confirming this one from the other side of the table. Arm data center royalties doubled year-on-year. He expects them to double again Things feel slow right now because we’re between platforms. The speedup ships in H2 2026. The architectural argument is over. Deployment is the only variable left I cover this in The Quiet Architect and The Fourth Piece $arm $NVDA

Ben Pouladian

62,986 Aufrufe • vor 3 Monaten

September 2009. Jensen Huang walks onto a small stage at the Fairmont hotel in San Jose. About 1,500 people are in the room. He runs a company that makes chips for video games. He spends the next 8 minutes doing math on a whiteboard, explaining why the future of computing won't come from making CPUs faster. He calls it "CEO math" and apologizes in advance to every computer science professor in the audience. Then he lays out an argument that almost nobody took seriously at the time: the way to make computers dramatically faster is to pair a regular CPU with hundreds of tiny parallel processors, the kind that already exist inside graphics cards. One CPU for the sequential stuff. Hundreds of GPU cores for everything else. He calls it "heterogeneous computing." He shows the math. A workload that can be split into many pieces at once gets up to 200x faster on this combined system. A workload that has to run one step at a time loses nothing. "The most important thing in creating a new architecture," he says, "is to make sure it does no harm." This was the first GPU Technology Conference. NVIDIA had launched a software platform called CUDA three years earlier, in 2006, to let developers write programs that run on graphics cards instead of just regular processors. Almost nobody cared. GPUs were for rendering Call of Duty, not for scientific computing. The academic world was polite but skeptical. The enterprise world ignored it entirely. By this point, Huang had been making this argument for years. NVIDIA was a $7 billion company. It competed with AMD and Intel for market share in the graphics market. That was the whole business. Jensen kept saying the GPU wasn't just a gaming chip; it was a computing platform. He kept saying parallel processing would reshape every industry from medicine to finance to physics simulations. People kept nodding, then doing nothing. Then deep learning happened. Around 2012, AI researchers discovered that training a neural network, which means teaching a computer to recognize patterns by running the same calculation millions of times across huge datasets, was exactly the kind of workload Jensen had been describing. GPUs can train AI models 10 to 50 times faster than CPUs. The architecture he outlined in this 2009 talk, with one CPU handling step-by-step tasks while hundreds of GPU cores crunch through massive amounts of parallel data, is now the literal blueprint for every AI data center on earth. ChatGPT runs on NVIDIA GPUs. Claude runs on NVIDIA GPUs. Gemini, Llama, Midjourney, nearly every major AI model you've heard of was trained on NVIDIA hardware using CUDA, the software platform Jensen built for a market that didn't exist yet. NVIDIA was worth about $7 billion when Jensen gave this talk. It is worth over $4.4 trillion today. That's a 600x increase. Jensen Huang, who founded the company at a Denny's in 1993 with two friends, now has a net worth of over $160 billion. He made Forbes' list of the 10 richest people for the first time this year. GTC 2026 is currently ongoing. 17,000 people are packing a hockey arena to watch the same guy explain what comes next. In 2009, 1,500 people showed up at a hotel ballroom, most of them for gaming graphics.

Anish Moonka

413,645 Aufrufe • vor 5 Monaten

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

318,151 Aufrufe • vor 4 Tagen

$AMD is easily a $1,200 stock IMO| CPUs TAM 🧵 Not Financial Advice! DYOR! In this thread, I want to discuss the actual TAM for CPUs data center for just 2026, where many are giving different ranges, where I don't agree with. I will explain in detail why I disagree with these research firms and financial analysts using Math. And this thread should not be treated as Financial Advice. I'm just explaining my research and thought process so we can have a discussion. In 2024/2025, I gave out $620 PT for FY2026 was too conservative for AMD potential. At the time, It was early and many were just laughing, that PT was unrealistic and the AI world is run on GPUs only. Today, most of these folks are laughing with me. That is ok, I dont offer financial advice, and I do not need everyone to agree with me. I respect other opinions. If you enjoy this kind of thread, slap the like/repost/bookmark. If you want to support my work further and gain more in-depth analysis, consider subscribe! In early 2026, hyperscalers, enterprises, and OEMs are scrambling as Intel and AMD server CPUs are largely sold out for the year, with prices jumping 10–20% and lead times stretching from weeks to months (or longer for certain SKUs). What was once a GPU dominated story has flipped: the shift to explosive Agentic AI with its multi-step reasoning loops, tool calling, multi-agent orchestration, real-time data movement, and reinforcement learning, is dramatically tightening CPU:GPU ratios from the old training-era 1:4–8 all the way to 1:1 to 5:1 or even CPU-heavy configurations. CEOs across NVIDIA, AMD, Intel, Google, Meta, Microsoft, and public companies have been sounding the alarm on CNBC, Bloomberg, and earnings calls. CPUs are “cool again,” and in many agentic deployments they are becoming the new bottleneck alongside (or even ahead of) GPUs and custom ASICs. In 2025, roughly 12-15m AI GPUs + AI ASICs GPUs shipped, and is expect to be 15-20m units by 2026, where it suggesting Training demand is not going away. The actual TAM is structural, multiplicative demand that has already forced AMD to double its long-term server CPU TAM forecast to >$120 billion by 2030 (>35% CAGR), with Dr. Lisa Su noting Q2 2026 server CPU sales expected to surge 70%+ year-over-year and demand “far exceeding expectations.” At the same time, AMD’s secured 30–40% share of TSMC’s initial 2nm capacity (behind only Apple’s >50%) positions it to ramp Zen 6-based EPYC Venice exactly when this agentic wave hits hardest but even that aggressive five-fab 2nm expansion (with plans scaling toward 11 total advanced facilities) cannot instantly close the gap in the near-term. Supply constraints on wafers, advanced packaging, and power are compounding the squeeze, just as hyperscalers forward-buy and lock in long-term deals. 1. The actual potential TAM Various sources and institutions are giving $50-$160-$200B CPUs TAM toward 2030, and i disagree, where supply is severely behind vs Demand by at least 2-3 years or even longer by some estimates. The actual TAM will probably be 15-20m for FY2026. The typical average selling price from low to high end is $5,000 to $15,000, but due to rising memory, and different inflationary pressures on Semi, it would be more logical to think between $7,000-17,000. A. CPU:GPU Ratio at 1:1 A basic calucation at mid range =12,000 x 15-20m CPUs= $180-$240B TAM B. CPU:GPU Ratio at 5:1 = $12,000 x 75m-100m CPUs= $900B-$1.2T TAM Of course TSMC cannot even supply 20% of this massive inflection TAM in 2026. But do we think of Demand for TAM or Supply for TAM? Hence we are seeing massive 2nm Ramp from TSMC for $AMD. IMO, conservatively, I would take down 15-20% on 1:1 or $135-$192B TAM for just 2026. Im not even talking about 2030. We are just months into this, it is impossible to estimate Cagr atm, but this is 1-5 agents running tasks, I wrote a thread on 24/7 autonomous agents thread, where companies could use 50-250 agents to run tasks for them 24/7. It would require a different structural CPU:GPU to bring down the cost of token as well as handling the Orchestration bottleneck. GPUs would be useless and sit idle waiting for CPU due to highly CPU-intensive nature. The cost per Million tokens must come down more rapidly for this 50-250 autonomous agents to work, otherwise the token cost would be too enormous. Helios Rack is estimated to bring inference cost down to $0.0003-$0.0005/M tokens with 18 EPYC Venices along with 72 MI455x and other chips+ Components. A heavier or CPUs dense rack would bring down inference cost further. EPYC Verano(2027 gen 7 AI-optimized) is expected to drive inference costs meaningfully lower than the Venice baseline likely to the $0.00002–$0.00025 per million tokens range (or even sub-$0.00015 in highly optimized agentic/batch workloads). Verano have higher core counts than Venice, LPDDR5X SOCAMM2 memory support, more AI optimized and Next-Gen rack density & efficiency. 2. $AMD secured at least 30-40% of TSMC 2nm capacity and Memory from Samsung through 2028-2030. 2 2nm fabs are entering ramping phase toward 60-65k wafers per months and 5 dedicated 2nm fabs entering mass production/ramp in 2026. Will link sub threads below if you are interest for full detail. Apple is reported to secure 50%+ 2nm capacity for Iphone 18 and Mac chips and AMD secured at least 30-40% capacity while $NVDA $AVGO $ARM $AMZN $GOOGL and others are on 3nm. This broader aggressive ramp from TSMC to target up to 11 fabs is to address $AMD massive growth ahead. Where $ARM is facing massive CPUs supply constraints as they have to compete with other Mega Cap players on 3nm allocation. And $INTC is also facing supply constraints for data center CPUs and PC per management with lead times extrended to longer than 12 weeks. Dr. Su is aiming for higher than 50%+ Market share, and I believe it is achievable in 2026 or 2027 as AMD has the strongest CPUs offerings. Dr. Su did not want to take advantage of the shortage and she said during the Q1 earning call, AMD is prioritizing Units shipped while guiding margin to be inching 60%. If Jensen were in charge, I'm sure margin would be 70-75% in this kind of severe CPUs shortage condition. But that is not how Dr. Su operates for more than a decade. She wants most market share. So we will see it in revenue growth, but as TSMC ramps faster and faster, AMD Operating and FCF margin will massively improve vs prior decade. A significantly higher margin profile than before. 3. How I came up with $1,200 withint 12-18 months? At $1,200/ share, that would be around $2 Trillion MC. I expect FY2027 revenue to be $124-$144B where data center revenue dominates overall revenue. AI GPUs: I will stick to the lowest end so show u that I'm conservative at $18B for each GW vs $NVDA Rubin is $30B+ (most likely Helios Rack in the $20B+ due to memory price rising). We know deals with OpenAI and Meta are around 12GW and additional multi-customers at multi-GW scale were hinted and will be revealed as we get to July 22-23 2026 Advancing AI event. For now I will conservatively add a bit more to this model. (3-6GW Helios Rack Range) EPYC Venice is reported to be in $15,000-$20,000. However large customers will likely to enjoy $10-$12k discount. I expect AMD to be able to ramp 7m EPYC Venice for entire 2026 and 3-4m of EPYC Verano(higher price than Venice). If we take an average selling price of $10,000 to be on the conservative side. Take down another 30% to be even more conservative on projection. I like to be conservative. That would be ~ 7m EPYC CPUs(Venice + Verano) for FY2027 or 583,000 units per month or 15,000 additional 2nm wafers per month which is completely reasonable for current TSMC Ramp, and I may be too conservative here. EPYC Verano and MI500 series will also be on 2nm. AI GPUs: 3GW x $18B= $54B EPYC CPUs: $10k x 7m CPUs= $70B = Data center revenue alone is $124B Other segments= probably in the $20-$25B FY 2027. FY2027 revenue = $124-$149B At 7m EPYC CPUs for entire 2027, that would be more than 50% market share when we comp it to availability from supply side, not from total Demand. It is possible that TSMC could significantly ramp even more capacity in 2027, so we will see. Metric Q1 2026 FY2027 Gross Margin 55-56% 60-62% Operating Margin 25-26% 32-35% Net Income Margin ~22% 26-30% FCF Margin 25% 28-30% At $124-$149B Revenue FY 2027 Net Income would be $32-$44B EPS would be $20-$27 (GAAP) Non-GAAP would be $25-$31 At $1,200 a share or $2T valuation that would be: 13.4-16x Price to Sales (P/S) 38-48 P/E At this kind of growth of AI SuperCycle, I think it is very reasonable valuation. If we use today at $406/share or $661B MC: 2027 P/S = 4.4x-5.3x 2027 P/E = 13x-16x Is AMD today expensive or cheap to you? Above is already a very conservative where I trimmed 20-30% of doable units. Meaning, there could be upside if TSMC is able to ramp meaningfully like they are planning. Conclusion: A $1,200 per share valuation IMO for AMD in FY2027 is not expensive at all; it is, in fact, conservative when viewed against the structural explosion in agentic AI demand we have mapped out. With server CPU TAM potentially scaling into the $100–$200B+ range in just CPU:GPU 1:1 Ratio for just 2026. AMD positioned to capture 50%+ share thanks to its 2nm TSMC allocation advantage and full-stack leadership, the company could realistically deliver $124–149B in total revenue and $25–$31+ non-GAAP EPS. At those levels, $1,200 implies a 2027 P/E = 13x-16x. Entirely reasonable for a company that will have become the clear Inference Queen (and in many workloads the preferred) AI infrastructure provider, with operating margins expanding above 30% and tens of billions in high-margin rack-scale AI revenue. Dr. Lisa Su was right presciently so about the Agentic AI inflection all the way back to her early 2022–2023 commentary on the coming shift from pure training to inference and orchestration-heavy workloads. While the broader market only fully woke up to this in 2026 when she doubled AMD’s long-term server CPU TAM forecast to >$120B by 2030 (with >35% CAGR), Dr. Su and her team have consistently positioned the company at the center of the CPU renaissance. The explosive demand we are seeing today, sold-out lines, rising ASPs, and hyperscalers forward-buying entire gigawatts of Helios-class systems is exactly the outcome she forecasted years ago. Not Financial Advice! DYOR!

Mike

301,322 Aufrufe • vor 3 Monaten

I had to test it myself to believe this unreal inference speed. 3,000 tokens/s for 1 user on standard datacenter GPUs. They leveraged a hidden efficiency gap in how GPUs generate tokens. Kog just achieved 3,000 tokens/s on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200 (FP16, no speculative decoding). Their tech preview is on a 2B model, and they show how their techniques will scale to large frontier MoE models at similar speeds. That's a huge number because normal low-batch GPU decoding for 2B to 8B models is usually closer to 100 to 300 tokens/s per request, so Kog is claiming something like a 10X to 30X jump in the speed one user actually feels. Their trick: they are getting the speed by treating LLM decoding as a memory streaming problem, not mainly a math problem. For 1 user at batch size 1, the GPU is not doing big, efficient matrix-matrix work like in training or large-batch serving; it is repeatedly pulling the model’s active weights from high-bandwidth memory for each new token, so speed depends on how smoothly those weights keep flowing. Normal inference stacks keep breaking that flow. They run many separate GPU programs for different parts of the model, move intermediate results through memory, wait at synchronization points, talk back to the CPU for scheduling or sampling, and then repeat this token after token. Kog’s answer is to co-design 3 things that are usually tuned separately: the runtime, the low-level GPU code, and the model architecture. The biggest engineering move is the monokernel, where the whole decode pass runs as 1 persistent GPU-resident program, including sampling, so the system does not keep stopping for kernel launches, CPU scheduling, and intermediate memory round trips. They also rebuilt synchronization, because their own measurements say grid sync was eating around 35% of token-generation time; instead of making every compute unit wait at a broad barrier, each unit waits only for the exact data it needs. On AMD MI300X, they also map memory access around the chiplet layout, because memory latency changes depending on which die makes the request. Then their Laneformer model uses Delayed Tensor Parallelism, which lets cross-GPU communication happen in the background instead of blocking every layer.

Rohan Paul

13,244 Aufrufe • vor 2 Monaten