Загрузка видео...

Не удалось загрузить видео

На главную

🔥 Tyler on $CIFR’s Direct Compute Vision When asked about direct compute, Tyler revealed they’re exploring a 56 MW GPU pilot program. He’s evaluating options to rent, lease, or return GPUs for credit — a model that keeps hardware fresh every few years. But he emphasized one thing: it...

13,441 просмотров • 10 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Demis Hassabis just explained why the real AI bottleneck has nothing to do with training runs. Most people picture the AI arms race as who can build the biggest model. GPT-4 or Gemini Ultra style training runs, a few hundred million in compute, fired once or twice a year. The constraint sits somewhere else. Every time a researcher has a new algorithmic idea, a new architecture, a new training technique, they can't just test it on a laptop. They have to run it at the scale where it would actually be deployed, because ideas that look promising at small scale fall apart completely when you put them into a real system. Every research hypothesis burns significant compute before a single line of production code gets written. At a lab like DeepMind, hundreds of researchers are running hundreds of ideas simultaneously. The demand for experimental compute is continuous. It never stops. Now layer the hardware reality on top. GPU lead times are currently 36 to 52 weeks for data center hardware. Global AI data centers are already drawing 29.6 gigawatts, equivalent to the peak power demand of the entire state of New York, and they still can't meet demand. Companies willing to pay any price can't just buy more compute. They wait in line. The speed of scientific discovery in AI is now gated by hardware availability. The next breakthrough is sitting in a researcher's head right now. Whether it gets validated fast enough to matter depends entirely on whether the compute is there when they need it. The AI race gets won by whoever can run the most experiments per month.

Aakash Gupta

32,150 просмотров • 3 месяцев назад

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 просмотров • 5 месяцев назад

Extra outtake clip from latest Bg2 Pod with Jensen Brad Gerstner Contrary to popular hypersensationalist rhetoric -- that we are in a massive AI glut -- we are likely in a stretch of structural compute shortage. Google announced in May that tokens had grown 50x y/y, and doubled again by July 2025 (100x) to 1 quadrillion monthly tokens. In that period - algorithmic and hardware advances improved efficiencies by ~10-15x - which means Google had to increase accelerated compute dedicated to token generation by 3-10x. Our estimate is that Google increased accelerated compute by ~3x during the period - which means that they had to pull compute from training, recommenders, etc to allocate to token generation. Significant algorithmic advancements (Flash Attention, quantization, MoE), and infrastructure investments (prompt caching, batching) have driven much of that efficiency, but counting on the hardware to get better is something the industry is counting / relying on. We used to be able to ride Moore's Law / Dennard scaling to improve compute per watt. But now... we have to rely on $NVDA / hardware ecosystem (Google, $AMD, etc) to drive 2-4x improvement per generation (Huang's Law). People underestimate the strain of exponentials on human systems - that are hard coded to think linearly... The total global accelerated compute base is probably ~8 GW of installed capacity on my math and analysts have estimated growth to increase to 10x to ~80 GW globally by 2030. Even assuming all of that capex gets done, and Nvidia continues to push yearly roadmap (generating 2x y/y performance uplift), the 10x power increase should equate to ~50x increase in compute. We just had 100x AI usage increase in 1 year from Google's testimony. OpenAI, Google, Anthropic are in structural shortage of compute - each bit they bring online is fully invested in serving their users. And that is without even expanding into true video / world models, robotics, or long horizon thinking to find novel breakthroughs. What could change this trajectory? If the algorithmic efficiencies we are gaining from new breakthroughs outstrip the exponential increase in current demand -- and FUTURE demand. Investors hyperventilated at DeepSeek's release earlier this year, but their gains were outweighed by the increase in demand created by reasoning.

Clark Tang

176,622 просмотров • 10 месяцев назад

🚨ALERT: 50% of Data Centers will NEVER connect to the grid. Half of the data centers announced in the last 24 months will NEVER connect to the grid. Kevin O’Leary said it. The data proves it. While everyone’s chasing “paper capacity,” $CIFR and $IREN are sitting on EXECUTED grid connections that can’t be replicated. Here’s why they’re untouchable: 266 GW of power projects canceled in 2025 alone. That’s 2.4x the cancellations from 2024. Why? Because the U.S. grid is facing a structural deficit that nobody wants to talk about. • Data centers need 18-36 months to build • Grid connections take 5-7 YEARS (sometimes 12) • Interconnection queues in PJM and ERCOT now average 7 years • Average interconnection cost in MISO: $753,116 per MW Translation: You can announce a data center tomorrow. But you CAN’T connect it to power until 2032. The math doesn’t work. The timeline doesn’t work. The physics don’t work. $CIFR - The Fixed-Price Power Moat: Cipher control one of the lowest-cost power portfolios in North America. > Power cost: $0.027/kWh (fixed, long-term PPAs) > Debt: $0 > Portfolio: 2.2 GW across Texas But here’s what everyone’s missing: Their 1-gigawatt Colchis site has a FULLY EXECUTED Direct Connect Agreement with American Electric Power. Not “in the queue.” Not “under study.” EXECUTED. Energization: 2028. While competitors are stuck waiting 7+ years for interconnection approvals, $CIFR already has a Tier 1 grid connection locked in. And they just signed: • $5.5 billion, 15-year lease with AWS for 300 MW • 10-year hosting deal with Google/Fluidstack for 168 MW That’s $8.5 billion in contracted lease payments for AI infrastructure. $IREN - The Microsoft Validation: $IREN didn’t just secure power. They secured the ONLY thing that matters: a hyperscaler willing to pre-pay billions. November 2025: $9.7 billion AI Cloud contract with Microsoft. Let me repeat that. Microsoft PRE-PAID for capacity that doesn’t exist yet. Deal structure: • 200 MW of liquid-cooled AI capacity • $1.94 billion annual recurring revenue (once online) • 20% prepayment to fund $5.8 billion GPU purchase from Dell • Four “Horizon” data centers at their 750 MW Childress campus But the real alpha? Their 2.91 GW portfolio of GRID-CONNECTED power. Not speculative. Not “in the queue.” Connected. Energized. Operating. > Sweetwater 1: 1.4 GW (energization accelerated to April 26) > Childress: 750 MW (operating) > Prince George: 160 MW hydro (23k GPUs for AI) $IREN is scaling to $3.4 billion in AI Cloud ARR by end of 2026 using only 16% of their total power capacity. The Peer Comparison Nobody’s Talking About: Everyone’s excited about $RIOT, $MARA, $CORZ, and $WULF. Here’s the problem: $RIOT: 1.7 GW portfolio, mostly Bitcoin-focused. 25 MW HPC lease with AMD ($311M over 10 years). That’s 1/30th the size of IREN’s Microsoft deal. $MARA: Building “behind-the-meter” natural gas generation to BYPASS the grid entirely. Smart strategy, but they’re starting from scratch. 1.8 GW capacity, mostly mining. $CORZ: $10B+ contract with CoreWeave sounds massive. But they’re CONVERTING old mining infrastructure. Not purpose-built for AI. Currently unprofitable. $WULF: 750 MW at Lake Mariner. Zero-carbon hydro/nuclear. Clean energy story is strong. But only 72.5 MW of HPC capacity by Q2 2025. Meanwhile: • $CIFR has 2.2 GW with executed grid agreements and $8.5B in hyperscaler contracts • $IREN has 2.91 GW of energized capacity and a $9.7B Microsoft deal The Cooling Bottleneck: Secured power means NOTHING without secured cooling. November 2025: CyrusOne data center in Illinois went down for 10 hours because ONE chiller failed. This facility handles TRILLIONS in CME trading volume. Energy, agriculture, crypto derivatives markets frozen globally. Why? Because AI racks now consume 600 kW of power (enough to power 500 homes). A single rack failure creates catastrophic heat buildup. $IREN’s solution: Liquid-cooled infrastructure at all Horizon facilities. $CIFR’s solution: Turnkey air-and-liquid cooling delivery for AWS. Hyperscalers aren’t paying billions for “power connections.” They’re paying for THERMAL RELIABILITY. The Numbers That Matter: > PJM capacity prices: 10x increase from 2024 to 2025 (extreme scarcity signal) > Interconnection costs in Louisiana/Missouri: $900,000+ per MW > $64 billion in U.S. data center projects blocked or delayed in 2024-2025 > 25+ major data center projects canceled in 2025 alone The grid is saturated. The timeline is broken. The infrastructure doesn’t exist. But $CIFR and $IREN? They already own the infrastructure. They already have the grid connections. They already have the hyperscaler contracts. The Bottom Line: > AI demand is doubling every 90 days. > Grid capacity takes 5-7 years to build. > You can’t close that gap with announcements. You close it with EXECUTED agreements and ENERGIZED megawatts. $CIFR: $0.027/kWh power, $8.5B in contracts, 1 GW Tier 1 grid connection $IREN: $9.7B Microsoft deal, 2.91 GW energized portfolio, $3.4B ARR target by 2026. While half the industry fights over interconnection queues, these two are already plugged in. The power crunch isn’t coming. It’s here. And the only winners will be the ones who secured their megawatts BEFORE the grid broke. Bullish $CIFR and $IREN. Note: This is NOT financial advice.

Black Panther Capital

347,528 просмотров • 6 месяцев назад

At LinqAI, we’ve been mission-driven since day one—building competitive, revenue-generating SaaS applications. As our product line has expanded and our customer base has grown, one thing has become clear: we have an opportunity to forge a deeper connection between our SaaS innovations and the Web3 ecosystem. As we scale our SaaS offerings, our demand for compute power continues to grow. Traditionally, paying for compute means sending value outside our ecosystem—resources spent elsewhere, never to return. But what if every compute cycle contributed to the LinqAI ecosystem itself? LinqProtocol is our answer. By transforming idle computers (PCs, laptops, servers, and in the future even phones) into a global compute network, we not only secure scalable and efficient infrastructure but also strengthen the utility of $LNQ—ensuring that every workload contributes to our own decentralized future. LinqProtocol allows for arbitrary compute which means that there are virtually no limits to what can be done on LinqProtocol. From AI Training, to Synthetic Data Generation, 3D Rendering and even hosting a Website - It's all possible on LinqProtocol. ✅ On-demand, scalable compute ✅ Transparent, fair-market pricing ✅ Trustless execution with real-time tracking ✅ Governance-based dispute resolution The future of compute isn’t just more efficient. It’s owned by those who use it. It’s decentralized. And it’s coming soon. $LNQ #LinqProtocol

LinqAI

239,673 просмотров • 1 год назад

The neocloud category may be the most misunderstood corner of the AI trade because the market still treats these names as one uniform GPU-hours bet when they are actually very different business models: 1. $NBIS (Cloud Utility for the Agentic AI Age) $NVDA just chose Nebius as an architecture partner for the agentic AI era by co-designing AI factories with them, and the Rubin GPU access that comes with this partnership means Nebius gets the next-generation inference stack before almost anyone else in the market. At a $28B market cap, a 5GW power target and Nvidia’s engineering team embedded in the stack.. this is my favorite name in the neocloud category. 2. $IREN (Energy-to-Compute Engine of the AI Era) The dilution fear is real but the market is misreading it. IREN is not diluting to survive but diluting to scale into a $3.7B ARR target and the $9.3B in funding already secured through customer prepayments and GPU financing means the $6B ATM is optionality capital. The real bottleneck in AI infrastructure right now is power and IREN controls ~4.5GW of secured capacity while needing only ~500MW to support its ARR target by year-end. That 10x ratio of power capacity to near-term need is something no competitor can replicate quickly. 3. $CIFR (Landlord of the AI Utility Era) Cipher is not a pure neocloud but is a hyperscale infrastructure landlord signing decade-long leases to $AMZN AWS and $GOOGL while they fill the shells with compute. The AWS lease alone is expected to generate ~$700M in average annualized NOI for the next decade at nearly 100% NOI margins. Power-rich land is the scarcest resource in AI infrastructure and Cipher controls it with 600MW fully contracted, both facilities fully funded through non-recourse fixed-rate project debt and a 3.4GW development pipeline. 4. $CRWV (The Fragile Giant) CoreWeave’s demand backlog and revenue growth are very real but none of that matters if the capital markets close for even one quarter. Interest expense hit $388M in Q4 and management guided Q1 2026 interest expense to ~$550M which implies an annualized run rate above $2B before a single new data center comes online. The bull case requires capital markets to stay open, rates to cooperate, hyperscalers to honor take-or-pay contracts in full and construction to stay on time. That is a lot of dependencies in a macro environment where oil is approaching $100 and private credit is already showing signs of stress.

Shay Boloor

1,097,755 просмотров • 4 месяцев назад

Nebius will be the first neocloud to hit $1 trillion dollar company and here is exactly why (Save this). As dylan patel says Jensen Huang absolutely hates a world where the hyperscalers have all the power. A world where Microsoft, Amazon, and Google are the only ones building compute is a world where Nvidia is slowly being squeezed by a handful of customers all simultaneously developing custom chips to replace Nvidia GPUs entirely. Google's TPU, Amazon's Trainium and Microsoft's Maia all exist for one reason, to cut Nvidia out of the stack and Jensen knows it so he is playing a long game most investors haven't registered yet. By funding NeoClouds and NeoLabs at scale, Jensen is deliberately engineering a multipolar compute world where no single hyperscaler can dictate terms and where Nvidia hardware remains the default infrastructure layer regardless of which model or platform ultimately wins. Nvidia has deployed roughly $40 billion in AI ecosystem investments across OpenAI, Anthropic, CoreWeave, Nebius, xAI, and dozens of infrastructure companies, all running almost exclusively on Nvidia chips, cementing GPU dependency across the entire AI stack.sedaily Every neocloud that survives and scales becomes a permanent Nvidia GPU customer structurally opposed to the hyperscalers building custom silicon expanding Nvidia's market while simultaneously weakening its biggest competitive threat. Dylan Patel described the neocloud ecosystem as throwing bait into the water and letting the best fish survive, warning that many heavily-backed teams will fail, but the ones that emerge will pull hundreds of millions in ARR right out of the gate. Nebius is that fish because it's the only neocloud operating at hyperscaler scale while remaining fully purpose-engineered for AI workloads from silicon to software. The numbers confirm Nebius has already cleared the survival bar that will eliminate most of the 200+ neoclouds competing right now. Revenue hit $399 million in Q1 2026, up 684% year-over-year, backed by $46 billion in contracted backlog, 3.5 GW of contracted power across seven site and a target of $7–$9 billion in annualized revenue by year-end. When Google approached neoclouds about deploying TPUs, Nebius said no, its Chief Revenue Officer noting that demand is 99% for Nvidia GPUs and that TPU interest comes almost entirely from former Google employees rather than the actual market. That alignment with Nvidia's ecosystem, at this scale, with this backlog, and this level of strategic backing is why Nebius sits in a category of one among the neocloud field. Patel framed the broader play correctly, every neocloud that survives makes Google's TPU and Amazon's Trainium structurally weaker simply by existing and five years from now, the winners will have reshaped the entire compute landscape in Nvidia's favor. Nebius is already hundreds of millions in ARR ahead of the competition while most of the field is still treading water. Milk Road subscribers are already up massively on the Nebius trade, and we are tracking the neocloud buildout as Nvidia works to reshape the entire compute market. Come join Milk Road Pro for our full Nebius breakdown, the valuation framework, the revenue targets we are watching, and the AI infrastructure names we like next for just $1. Link below!

Milk Road AI

92,650 просмотров • 1 месяц назад

David Sacks just said what every honest analyst in Silicon Valley is already thinking (Save this). Nobody has ever seen anything like this. Anthropic has grown at 10x per year for three straight years and going into 2026, the conventional wisdom was that the rate of growth had to slow at this level of scale but then the numbers came in. Q1 alone is $10B ARR to $30B, in April, $30B to $44B and that's $96 million in new ARR added every single day. Inference margins are now above 70%, up from 38% last year and the only thing holding them back was compute. That's solved now, the SpaceX deal and others Anthropic has been quietly signing unlocks the supply side. This is exactly why we are bullish on Nebius and AMD. When a single company is adding nearly $100M in ARR per day, the real trade isn't the frontier lab but rather the infrastructure underneath it. Nebius, one of the fastest-growing neoclouds on the planet posted 547% YoY revenue growth in Q4 2025, exited the year with $1.25B ARR, and is guiding for $7–9B ARR by year-end 2026. Their revenue backlog has reached $46B, with projections of $16B in revenue by 2028 and NVIDIA locked in a $2 billion stock buy agreement with them giving Nebius early access to cutting-edge chips while every other cloud scrambles for supply. AMD is the other side of the same coin. Data center revenue hit $5.78B in Q1, up 57% year-over-year with total company revenue at $10.25B, up 38%. Meta has committed to deploying up to 6 gigawatts of AMD Instinct GPUs. Data center GPU revenue is forecast to surge 114% year over year to $15B in 2026. MI400-series chips hit the market in H2 and analysts project segment operating margins climbing to 31% as the next generation ramps. The model is simple, Anthropic is printing revenue and that that revenue pays for compute. That compute flows through companies like Nebius and AMD. This is why Milk Road PRO remains bullish on them and our positions are up massively. Our analysts have broken down the full thesis, the allocations, and the price targets. Go PRO at Milk Road to see everything, link below!

Milk Road AI

184,643 просмотров • 3 месяцев назад

THE 5 BIGGEST BOTTLENECKS POWERING THE AI ECONOMY The way I'm thinking about AI winners today is that the market is moving beyond the simple question of which mega-cap company spends most on AI and has been rewarding the companies that control the scarce inputs, contracted capacity, data movement, power infrastructure, edge compute and workflow layers that make the AI economy function. These are the five bottlenecks I'm watching most: • Memory | $MU, Samsung, SK Hynix Memory is the clearest scarce input because HBM feeds the accelerator, only a few companies can make it at volume and buyers are locking in supply through long-term agreements that create revenue visibility through the end of the decade. • Connectivity | $AVGO, $MRVL, $ALAB, $CRDO, $AAOI, $ANET Connectivity determines whether AI clusters can move data fast enough because training runs span tens of thousands of chips that need to act like one machine. Once copper runs out of reach that causes the cluster depends on optics, retimers, switches and custom silicon to keep the system moving. • Power | $CEG, $VST, $GEV, $FPS, $VRT, $NVTS, $TLN, $ON Power determines whether new AI data centers can actually come online because the binding constraint is shifting from getting chips to getting megawatts so the value flows to the companies that control generation, grid equipment, power delivery, thermal management and efficiency. • Compute | $NBIS, $CIFR, $IREN, $APLD, $WULF, $CORZ, $CRWV Compute capacity is overflow layer when hyperscalers are sold out. Capital alone doesn't guarantee GPU access which is why buyers are signing multi-year contracts for clusters before they are even fully built. • CPU | $NVDA, $AMD, $INTC, $ARM, $QCOM On-device CPU (edge compute) becomes next bottleneck as AI moves into inference, agents, PCs, phones, vehicles and physical devices.

Shay Boloor

108,564 просмотров • 1 месяц назад

If intelligence is the log of compute… it starts with a lot of compute! And that’s why we’re scaling our GPU fleet faster than anyone else. Just last year, we added over 2 gigawatts of new capacity – roughly the output of 2 nuclear power plants. And today we’re going further, announcing the world's most powerful AI datacenter, located in southeastern Wisconsin. Fairwater is a seamless cluster of hundreds of thousands of NVIDIA GB200s, connected by enough fiber to circle the Earth 4.5 times. It will deliver 10x the performance of the world’s fastest supercomputer today, enabling AI training and inference workloads at a level never before seen. For AI training workloads, you need compute at exponential scale. That’s why we designed the datacenter, GPU fleet, and network together as one integrated system. This ensures a single job can run from day 1 at exponential scale across thousands of GPUs. Fairwater uses a liquid-cooled closed-loop system for cooling GPUs that requires zero water for operations after construction. And we’re matching all of the energy that is consumed with renewable sources. And of course, it is just one of several similar sites we’re lighting up across our 70+ regions. We have multiple identical Fairwater datacenters under construction in other locations across the US, in addition to our AI infrastructure already deployed in over 100 datacenters around the world, powering model training, test-time compute, RL tuning, and real-time inference at global scale. Too often during times like this, people go with the current and only later wonder, how did we get here? With Fairwater, we're charting a new path: doing the hard engineering work, bringing compute, network, and storage into one highly scaled cluster, and designing closed-loop energy systems to meet real-world computing needs. And partnering with local communities to ensure it's thoughtfully done in a way that is sustainable, creates new jobs, and expands opportunity. We are thrilled to see this take hold in Wisconsin, and we are just getting started.

Satya Nadella

2,022,693 просмотров • 10 месяцев назад

Elon Musk was asked how fast AI is moving, and answered by correcting the most aggressive forecaster in technology for not being aggressive enough. Musk: “I have to give credit to Ray Kurzweil in being actually remarkably accurate in his predictions. If anything, I think he was perhaps a bit conservative in his predictions.” Kurzweil spent thirty years publishing timelines that were treated as proof he had lost the plot. He hit nearly all of them. The word Musk used for that record was conservative. Kurzweil built his forecasts by reading the rate of change. Musk is reading it now. Musk: “The dedicated AI compute appears to be growing by a factor of 10 every six months.” That is not a projection. “Appears to be growing” describes something that has already happened. Musk: “Almost a 100x improvement per year, at least for the next few years.” Moore’s Law was 2x every two years. That one curve produced the internet, the smartphone, the cloud, and every instinct you have about how fast the world can change. Your sense of normal was calibrated on 2x. The system is running at 100x. The new world is not being built next to the old one. Musk: “Probably a lot of the data centers, maybe most of the data centers that currently do conventional compute, will transition to AI compute.” The infrastructure holding up the current economy is being converted into what replaces it. The old world is the raw material for the next one. Every generation assumes the speed it grew up with is the speed of history. Ours is the first to be handed the actual number and still round it down. Everything you build takes time. A child, a trade, a life’s work. All of it rests on the assumption that the world will still be recognizable when the work is finished. That assumption has held for every human who has ever lived. It was never a law of nature. It was a side effect of how slowly things moved. The most extreme forecast on record came in low. Reasonable is no longer the safe position. It is only the comfortable one.

Dustin

163,648 просмотров • 3 дней назад

How could you possibly be bearish on compute right now? (Save this). Every 10 seconds in 2026, the world generates 31.7 billion tokens and by 2030, that number hits 1.27 trillion, every 10 seconds. That's a 40x increase and that's before the full agent economy comes online. The Qualcomm CEO said total token demand by 2030 is in the quintillions. Here's what most people miss because when you use ChatGPT, you generate tokens one conversation at a time but agents don't sleep. ] They run 24/7, spawning sub-agents, carrying context, updating memory, catching mistakes and every single one of those actions burns tokens. The shift from human paced to agent paced activity is the single biggest structural change in compute demand we've ever seen. You don't need a perfect forecast but rather just need to believe agents become persistent and if they do, compute demand goes vertical. The infrastructure has to be built before the demand fully arrives, which means the window to own the picks and shovels is right now. That's where neoclouds like Nebius come in. Nebius isn't trying to be AWS, it is a pure-play AI cloud, GPU clusters, inference infrastructure, and developer tooling built from scratch for AI workloads. Q1 2026 revenue hit $399M, up 684% year over year and they're guiding for $7–$9 billion annualized run rate by end of 2026. Analysts are modeling roughly 2,000% total revenue growth from end of 2025 to end of 2027. They already have contracts with Microsoft and Meta already signed. Capex guidance raised to $20–$25 billion because customer commitments justified it. They are sold out of capacity because the constraint isn't customers, it's how fast they can build. Adjusted EBITDA margin on the core AI business hit 45% in Q1 and Jensen Huang called Nebius a close partner at GTC 2026. And in a world where GPU access is the single biggest competitive moat, that relationship matters more than most people realize. The bear case on compute requires you to believe the agent economy stalls and that's a very lonely bet to make right now. Bullish on Nebius and Milk Pro subscribers are already up massively on this trade, come join us using the link below to get our full AI trades and we have a HUGE 33% off right now!

Milk Road AI

16,246 просмотров • 1 месяц назад

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 просмотров • 1 месяц назад

Greg Brockman, President of OpenAI, said there is not enough compute in the world to satisfy AI demand, and OpenAI itself cannot launch products it has already built because it cannot find the infrastructure to run them (Save this). OpenAI is spending $50 billion on compute in 2026 alone and it still is not enough. That is the setup but here is the trade. Nebius is one of the most asymmetric infrastructure plays in public markets right now, and most people have never heard of it. Q1 2026 revenue came in at $399 million, up 684% year over year, with AI cloud revenue specifically growing 841% in a single quarter. The company entered 2026 with an exit ARR of $1.25 billion and is targeting $7 to $9 billion by year end, a number that would make it one of the fastest revenue ramps in the history of public infrastructure companies. The contracted backlog sits at $50 billion anchored by a $17.4 billion agreement with Microsoft through 2031 and a $27 billion five-year deal with Meta. They are decade-scale infrastructure commitments from the two largest enterprise AI spenders on earth, signed before the demand curve has even reached its steepest point. Nvidia took a direct equity stake in Nebius, one of only two neoclouds it has invested in alongside CoreWeave. That relationship is not just financial but rather means Nebius gets preferential access to GPU allocation at a moment when every lab and every hyperscaler is competing for the same constrained supply. Contracted power capacity now exceeds 3.5 gigawatts, with expansion plans targeting 5 to 6 GW by mid-2029. And power is the other binding constraint in AI infrastructure, you cannot build a data center without it and Nebius has already secured the capacity that competitors are still fighting to acquire. At full ramp, analysts project revenue in the $15 to $25 billion range by 2029, against a current market cap the contracted backlog alone already dwarfs. Come join Milk Road Pro and get our full Nebius deep-dive, the exact price levels we are watching, how we are sizing the position against the backlog and power capacity timeline, and our full AI thesis. link below!

Milk Road AI

14,578 просмотров • 1 месяц назад

Chamath Palihapitiya just dropped the number that explains the entire AI infrastructure trade (Save this). A gigawatt of compute now costs $100 billion and when he started his Arizona data center project it was $4 to $5 billion, it has gone up 20x in a single investment cycle. The implication is not just that AI infrastructure is expensive but rather that the capital barrier to owning meaningful compute has become so high that only a handful of entities in the world can actually build it and the companies who got there early are sitting on what may be the most durable pricing power in the history of the technology industry. This is the neocloud trade. The neocloud market, purpose-built GPU cloud providers like CoreWeave, Nebius, and Lambda Labs was worth $35 billion in 2026 and is projected to reach $236 billion by 2031, compounding at 46% annually. For context, that is faster growth than cloud computing itself posted in its first decade. The reason is very simple, hyperscalers like AWS, Azure, and Google are building for everything, storage, databases, enterprise software, networking and their GPU pricing reflects the overhead of that full-stack infrastructure. Neoclouds build for one thing only, AI compute. The result is a 60% to 85% cost advantage on the same Nvidia silicon, bare metal H100s at $0.78 to $2.79 per GPU-hour on a neocloud versus $3.43 to $5.07 per GPU-hour on a hyperscaler. That spread does not close as AI demand scales but rather it widens, because hyperscalers have to amortize legacy infrastructure and margin expectations that neoclouds do not carry. Gartner projects that by 2030, neoclouds will capture 20% of the $267 billion AI cloud market, and Vultr's own analysis says at least 80% of GPU market share by end of 2026 will be held by a small group of scaled neocloud providers. Now zoom into Nebius specifically, because it is the most interesting publicly traded proxy for this trade. Nebius is the infrastructure arm of the former Yandex Russia's equivalent of Google rebuilt from the ground up after Russia's invasion of Ukraine by Arkady Volozh and relisted on Nasdaq in October 2024. The team that built it already knew how to run internet-scale infrastructure at the lowest possible cost, which is exactly the operational DNA a neocloud requires. In Q1 2026, Nebius reported revenue of $399 million and already generating serious cash on a young business with revenue growing nearly eightfold year-over-year. Then in March 2026, Meta signed a five-year infrastructure agreement with Nebius worth up to $27 billion, $12 billion in committed dedicated GPU capacity deployments beginning early 2027, plus up to $15 billion more tied to Meta purchasing Nebius's unsold third-party capacity. The deal will be executed on one of the first large-scale deployments of Nvidia's Vera Rubin platform, the next-generation architecture after Blackwell making Nebius one of a tiny number of operators in the world with confirmed priority access to the most advanced AI hardware available. Following the contract, Nebius guided to $7 to $9 billion in annualized recurring revenue for 2026 representing 540% year-over-year growth. Chamath Palihapitiya point about the $100 billion capital moat is the bear case for new entrants and the bull case for incumbents. No one can afford to build the next CoreWeave or Nebius from scratch at current hardware and power costs. The companies that are already built, already contracted, and already deploying Nvidia's latest silicon have a moat that compounds with every GPU generation cycle because they get allocations first, they deploy fastest, and their customers re-sign rather than wait for a new operator that does not yet exist. Come join Milk Road Pro for our full breakdown, the complete neocloud competitive landscape, how to think about Nebius's valuation versus CoreWeave and AI entire thesis. Link below.

Milk Road AI

139,047 просмотров • 1 месяц назад

$AMD| $META is using $GOOGL to negotiate 🧵 The Ironwood pod is 5.1–10x more expensive annually ($148.3 million ÷ $14.87–$29.04 million) and 5.1–10x more expensive monthly ($12.36 million ÷ $1.24–$2.42 million) than renting 15 MI450 racks for equivalent compute. The rapidly evolving landscape of artificial intelligence infrastructure presents a complex interplay of technological innovation, market dynamics, and strategic maneuvering among major players. Recent leaked information suggesting that Meta Platforms ($META) might work with Google's Tensor Processing Unit (TPU) in 2027 has sparked speculation about its true intent. This leak is likely a strategic move by Meta to negotiate more favorable terms with AMD , leveraging the competitive dynamics of the AI hardware market to optimize its substantial investment in AI infrastructure. By examining the key elements of this scenario Meta's investment strategy, the comparative advantages of AMD's MI450 and Google's Ironwood TPU, and the broader market context; we can discern the potential beneficiaries and the strategic implications of this information. Meta's aggressive pursuit of AI capabilities is underscored by its planned expenditure of $66-72 billion on AI infrastructure in 2025, with expectations to escalate significantly in 2026. This investment is part of a broader strategy to build "titan clusters" like Prometheus, which are projected to reach 1 gigawatt of compute power by 2026. Such a scale of investment reflects Meta's recognition of the critical role that AI will play in its future growth, particularly in enhancing its social media platforms and developing new AI-driven applications. However, the financial burden of this infrastructure buildout necessitates a careful consideration of cost-effectiveness and scalability, which brings us to the leaked information about potential collaboration with Google's Ironwood TPU. Google's Ironwood TPU, introduced as the seventh-generation ASIC optimized for TensorFlow-based inference, represents a high-cost, cloud-locked solution priced at $445 million per pod (9,216 chips) over three years. This model, while offering significant performance gains and power efficiency, is tailored for pod-scale deployment and integrated with Google's cloud services, limiting flexibility and increasing costs for customers. In contrast, AMD's MI450 GPU, priced at $30,000–$40,000 per unit, provides a modular, open ROCm ecosystem that delivers comparable compute capacity at a fraction of the cost. Renting 15 MI450 racks could achieve similar 42+ exaFLOPS inference compute at 5–10x lower cost than renting a single Ironwood pod, underscoring AMD's competitive edge in terms of total cost of ownership (TCO). The leaked information about Meta's potential TPU deployment in 2027, therefore, can be interpreted as a negotiating tactic rather than a definitive shift in strategy. By signaling interest in Google's solution, Meta may be attempting to pressure AMD into offering more favorable terms/prices for 5-10GW. This tactic aligns with Meta's broader goal to finance most of its AI spend internally while exploring partnerships that can reduce costs and enhance flexibility. The post's emphasis on MI450's TCO advantage and its partnerships with major players like OpenAI, Microsoft, and Meta itself suggests that AMD is a critical component of Meta's AI infrastructure strategy. The threat of working with Google's TPU could prompt AMD to reassess its pricing, provide additional support, or offer incentives to retain Meta as a customer, thereby securing or expanding its market share. From a logical standpoint, Meta stands to benefit the most from this strategy. As a major buyer in a high-stakes market projected to surpass $1 trillion in annual spending by 2030, Meta's negotiating power is significant. The leaked information could lead to substantial cost savings on its $66-72 billion investment, enhancing its financial flexibility and allowing for further investment in AI capabilities. Moreover, this tactic reinforces Meta's position as a leader in the AI infrastructure race, potentially attracting more external financing for its data center projects and strengthening its competitive stance against other hyperscalers like Amazon and Microsoft. AMD could also benefit from this scenario. The negotiation pressure might lead to small short-term concessions, but it could also solidify long-term partnerships with Meta, ensuring continued demand for MI450 and other AI hardware solutions. Initially Meta's 42% allocation to AMD MI300X and its partnerships with Oracle, Dell, and HP indicates a deep integration of AMD's technology into Meta's infrastructure, which could be leveraged to maintain this relationship. For AMD, retaining Meta as a large key customer is crucial to capturing a larger share of the rapidly growing data center infrastructure market, driven by the insatiable demand for AI compute power. Google, on the other hand, faces a more limited benefit from this leaked information. While securing Meta as a customer would reinforce its position in the AI hardware market, the high cost and ecosystem lock-in of the Ironwood TPU might deter Meta from fully committing to this solution. The leaked information could prompt Google to reconsider its pricing or ecosystem strategy to remain competitive, but the immediate impact is likely to be minimal compared to the potential gains for Meta and AMD. Investors and market analysts also stand to benefit from this information, as it provides insights into the competitive dynamics of the AI hardware market. Adjustments in portfolios based on anticipated shifts in market share and profitability could lead to opportunities for those who correctly anticipate outcomes. The negotiation dynamic might introduce volatility, but it also highlights the strategic importance of cost-effective solutions in the AI infrastructure space. Lastly, the leaked information about Meta potentially working with Google's TPU in 2027 is likely a strategic move to negotiate with AMD, leveraging the competitive landscape to optimize its AI infrastructure investment. Meta, as the primary negotiator, stands to gain the most by securing better terms from AMD, reducing costs, and enhancing its financial flexibility. AMD, while initially at risk, could benefit from retaining a key customer and solidifying its market position. Google faces limited immediate benefits but may need to adapt its strategy to remain competitive. This scenario underscores the complex interplay of technology, market dynamics, and strategic maneuvering in the AI hardware market, where cost-effectiveness and scalability are paramount. As the data center infrastructure market continues to grow, the outcomes of such negotiations will shape the future of AI development and deployment.

Mike

182,225 просмотров • 8 месяцев назад