Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I think today’s $NVDA surprise factor could be memory driven price hikes expected in early 2027 pulling Vera Rubin orders forward and accelerating the next AI infrastructure upgrade cycle. That demand would hit just as Rubin ramps with higher rack density and better energy efficiency helping hyperscalers squeeze more...

144,183 görüntüleme • 10 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

$AMD's heading to $5T MC LT| Lowest $/M tokens 🧵 The real reason why Institutions are FOMOing into AMD while other Semi stocks are underperforming ($NVDA $AVGO) Not Financial Advice! DYOR! Under Dr. Lisa Su’s leadership, AMD has transformed from a distant challenger into a formidable force in AI infrastructure, delivering the industry’s most compelling TCO story for high-volume inference. Her clear vision open ecosystems, aggressive annual roadmaps, rack-scale innovation, and relentless focus on tokens-per-dollar has positioned AMD’s Helios racks as the go-to solution for hyperscalers and AI natives struggling with exploding token costs, collapsing the cost down to $0.0003-$0.0005/M tokens. I will link various threads on this analysis to supply chain and wafer ratio if you are interested in understanding the full picture. In the last 3-4 months, explosive Agentic AI demand significantly increased Inference demand for Agentic AI models with 5-10 agents. If you are a listener of CNBC or Bloomberg, u should know enterprises and companies are complaining abt cost of token, and how it starts to spike up way too much to make sense. The fact that most data center today are run by $NVDA Chips, where the cost is way too high for Training or Inference. 1. Token cost Here are some quick comp, so u understand why $META OpenAI Anthropic $MSFT $AMZN Softbank $GOOGL and many more small to medium AI Natives are buying AMD CPUs and GPUs as much as they want, or pretty much AMD chips are sold out for the next 3-5 years. Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens 2. Why Hyperscalers and AI Natives Are Choosing AMD Token consumption (especially Agentic) is outpacing even NVIDIA’s efficiency gains, making diversification mandatory for economic viability. Massive deals reflect this reality like $META, OpenAI, $MSFT, Softbank, $AMZN, Oracle, LumaAI, G42... Dr. Lisa Su’s Vision in Action: Since taking the helm, Su has driven AMD’s turnaround with disciplined execution, annual GPU cadence (MI300 → MI350 → MI400), full-stack software (ROCm 7), open ecosystems (UALink, OCP designs), and customer-centric rack-scale solutions like Helios. Her emphasis on “tokens per dollar” and TCO has turned AMD into the pragmatic choice for sustainable AI scaling. Power/Energy Efficiency: ~Helios Rack-level is estimated at 120kW-140kW with 50% more HBM4 where Inference and Training cost matter ~Rubin Rack-Level is estimated at 160kW-230kw AMD Helios shines in owned TCO, memory density, and energy flexibility at hyperscale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B 3. Superior CPUs to pair with GPUs on massive scale 5-10-20GW Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. Conclusion: NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always-on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. Not Financial Advice! DYOR! Video source: Microsoft Build 2026

Mike

145,992 görüntüleme • 3 ay önce

The CEO of the world's largest asset manager just said something that should reframe how every investor thinks about the AI trade. Larry Fink, managing $11.5 trillion at BlackRock, stood at the Milken Institute Global Conference and said four words that matter, "We just don't have enough compute." "The United States is short power. We're short compute. We're short chips. And there's going to be shortages in all three and memory, four things. I actually believe a new asset class will be buying futures of compute." Think about what that means. Fink is predicting that compute becomes a tradable commodity like oil, like grain, like natural gas where investors buy forward contracts on future capacity because the shortage is so structural and so predictable that a derivatives market will emerge to price it. That is not a minor observation from a finance executive but rather the chairman of the most powerful capital allocator on the planet telling you that compute scarcity is a multi-year, investable megatrend. The data backs him up completely. Data centers will consume 70% of all memory chips produced globally in 2026. Advanced HBM production from Samsung, SK Hynix, and Micron is sold out through 2026 and into 2027 and a single AI server consumes 10-20x more memory than a conventional workload server. DRAM supply growth is running at just 16% annually while AI infrastructure demand is growing at 80%+. The chip crunch, the power crunch, and the compute crunch are not temporary dislocations, they are structural, and they will get worse before they get better. Fink also said something the bears keep getting wrong: "There is not an AI bubble. There is the opposite. We have supply shortages. Demand is growing much faster than anyone has ever anticipated." This is why the Milk Road Pro portfolio is built the way it is, long the companies producing and supplying the constrained resources: chips, memory, compute infrastructure, and power. Check out Milk Road Pro, link below to access our full thesis and plays.

Milk Road AI

419,755 görüntüleme • 4 ay önce

The $1 trillion narrative around NVIDIA becomes clearer when structured against actual numbers and timelines. 1. 𝗧𝗵𝗲 𝗯𝗮𝘀𝗲𝗹𝗶𝗻𝗲 In 2020, NVIDIA generated ~$10.9B in annual revenue 2. 𝗧𝗵𝗲 𝗴𝗿𝗼𝘄𝘁𝗵 𝗰𝘂𝗿𝘃𝗲 FY2022 → ~$26.97B FY2024 → ~$60.9B FY2025 → ~$130.5B FY2026 → ~$215.9B ~𝟮𝟬𝘅 𝗴𝗿𝗼𝘄𝘁𝗵 𝗶𝗻 𝘀𝗶𝘅 𝘆𝗲𝗮𝗿𝘀 🤯 3. 𝗧𝗵𝗲 𝗱𝗲𝗺𝗮𝗻𝗱 𝗲𝘅𝗽𝗮𝗻𝘀𝗶𝗼𝗻 At GTC, Jensen Huang outlined: ~$1 trillion in cumulative demand for Blackwell and Vera Rubin through 2027 up from ~$500B projected through 2026 just one year ago Demand expectations doubled within 12 months 4. 𝗧𝗵𝗲 𝘀𝗰𝗮𝗹𝗲 𝗰𝗼𝗻𝘁𝗲𝘅𝘁 ~$1T across two product lines ≈ ~4.6x FY2026 revenue This excludes gaming, automotive, and other segments Even ~50% conversion implies a pathway toward ~$500B annual revenue within this decade 5. 𝗧𝗵𝗲 𝗶𝗻𝗱𝘂𝘀𝘁𝗿𝘆 𝗯𝗮𝗰𝗸𝗱𝗿𝗼𝗽 Top cloud companies are expected to invest ~$700B in data center infrastructure in 2026 Long-term projections point toward ~$3T–$4T annually by 2030 6. 𝗧𝗵𝗲 𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗮𝗹 𝗹𝗼𝗼𝗽 AI adoption → rising inference demand Inference demand → infrastructure investment Infrastructure → higher capability Capability → increased usage 𝐀 𝐜𝐨𝐦𝐩𝐨𝐮𝐧𝐝𝐢𝐧𝐠 𝐜𝐲𝐜𝐥𝐞 7. 𝗧𝗵𝗲 𝘀𝗵𝗶𝗳𝘁 Compute evolves into a primary economic layer Every AI interaction consumes it Every enterprise integrating AI scales it 𝗙𝗶𝗻𝗮𝗹 𝘁𝗮𝗸𝗲𝗮𝘄𝗮𝘆: The next decade aligns growth with access to compute And the companies controlling it define the pace and limits of that growth

Jaymin Shah

306,435 görüntüleme • 5 ay önce

Nvidia's Jensen Huang on Fox Business: * Asked for a single word on the rest of the year, Huang says "demand accelerating" * Says the reason is AI does productive work and the tokens labs generate are "profitable tokens," so what's holding it back now is just compute. * Next year is already sold out and supply-constrained at 70% growth. * "Every single [chip] that we can make has already been sold." * Customers commit a year-plus ahead because "you're not buying a computer, you're building a factory," at $50-60 billion each. * With memory prices way up, Nvidia reset margin guidance this quarter to "rip the band-aid off." * Margins fall from 75% to 72 - 73% next year after absorbing memory cost increases and repricing, to clear what he called "a cloud on our stock." * Next year's Vera Rubin will ramp even faster than Blackwell did, which Huang calls "the fastest product launch we've ever had in history." * Bottlenecks "literally everywhere," the biggest buildout in history. * Constraints span TSMC CoWoS packaging, wafers, memory, networking, silicon photonics, power generation, and "high-quality land that's powered," plus a skilled-labor shortage in the US * Says scarcity is not a bad thing. This quarter's 100% growth and next year's 70% are both supply-constrained, which can be healthy because it "causes the best technologies to emerge and rise" and makes everybody sharper. * AWS alone is set to add 2 million GPUs across 2027 and 2028. Half of revenue comes from about six hyperscalers, much of it rented onward to "Nvidia developers all over the world" * Says on hyperscalers, the Nvidia platform is "the most rentable" and "captures the highest pricing" of anything hyperscalers own. * After Bill Gates suggested taxing tokens over AI job losses, Huang says: "I see things very differently than he does." * Argues AI is "a net job creator at a scale that we've never seen" because productive companies hire more people not lay people off Plus, more. See video link in reply. Bookmark & watch: _____ More on why Nvidia needed to "rip the band-aid" due to memory makers:

Fireside Alpha

13,018 görüntüleme • 8 gün önce

Jensen Huang just doubled NVIDIA's demand forecast to $1 Trillion through 2027 🤯 Then spent two hours explaining why that number is conservative… Here's everything today from GTC: - NemoClaw: NVIDIA's open-source enterprise AI agent stack built around OpenClaw. Jensen called OpenClaw "the operating system for personal AI" and said every company needs a strategy for it. - Space-1: NVIDIA is putting Vera Rubin data centers in orbit. Not a concept. An actual system being designed for space deployment right now. - DLSS 5: 3D-guided neural rendering that blends raw graphics with generative AI. Jensen called it the future of real-time rendering. - AWS: Deploying 1 million+ NVIDIA GPUs starting this year. Azure was the first hyperscaler to power up Vera Rubin. - Vera Rubin: NVIDIA's next-gen AI supercomputer. 10x more performance per watt than Blackwell, 700 million tokens per second, shipping later this year. - Groq 3 LPU: First chip from NVIDIA's $20B Groq acquisition. A purpose-built inference accelerator that ships Q3. NVIDIA now owns training AND inference. -Feynman: The architecture after Rubin, coming 2028. New GPU, new LPU, new CPU. NVIDIA is on a 12-month chip cadence and the treadmill never stops. - Autonomous driving: BYD, Hyundai, Nissan, and Geely building Level 4 vehicles on NVIDIA. Uber deploying NVIDIA-powered robotaxis across 28 cities by 2028. The man doubled his demand forecast to a trillion dollars, announced data centers in space, and closed the show with a robot singing country music. This is NVIDIA's world. Everyone else is just renting compute in it.

Josh Kale

45,875 görüntüleme • 5 ay önce

This is the next big plan for SpaceX: AI Data Centers in Space. • To achieve even a small fraction of a Kardashev Type II civilization (harnessing the full energy of the Sun), AI compute will require orders of magnitude more energy than Earth can ever provide. • Earth only intercepts about 1–2 billionths of the Sun’s total energy output. • Massive-scale AI (e.g., a million times more energy than Earth could produce) can only be powered by capturing far more solar energy in space. • Space-based solar-powered AI satellites/compute clusters are therefore inevitable. • In space, sunlight is continuous (no night, no clouds, no atmosphere), so no batteries are needed. • Solar panels in space can be extremely lightweight and cheap (no glass, no storm-proof framing required). • Cooling in space is dramatically easier and simpler: just radiate heat directly into the cold vacuum — no water, no fans, no liquids, no massive cooling infrastructure. • Most of the mass/volume of current supercomputer racks (e.g., GB300) is cooling hardware; in space that largely disappears. • The cost-effectiveness of electricity and compute in space will soon be overwhelmingly better than on Earth. • Elon’s Prediction: within ~5 years (by ~2030), the lowest-cost way to run large-scale AI will be solar-powered satellites in space. • A terawatt/year of AI compute is essentially impossible on Earth with any realistic build-out of power plants. • Scaling both power generation and cooling on Earth at the required rate is physically and politically unfeasible.

Nic Cruz Patane

49,035 görüntüleme • 8 ay önce