Загрузка видео...

Не удалось загрузить видео

На главную

SITUATION EXPLAINED: Dwarkesh argues compute could get 10X more expensive. • Anthropic's revenue has been 10X-ing year over year while lab compute only 3X's • Three ways that gap can close: margins rise, compute gets more expensive, or labs shift compute to inference • All three are already happening,...

10,617 просмотров • 4 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 просмотров • 5 месяцев назад

Extra outtake clip from latest Bg2 Pod with Jensen Brad Gerstner Contrary to popular hypersensationalist rhetoric -- that we are in a massive AI glut -- we are likely in a stretch of structural compute shortage. Google announced in May that tokens had grown 50x y/y, and doubled again by July 2025 (100x) to 1 quadrillion monthly tokens. In that period - algorithmic and hardware advances improved efficiencies by ~10-15x - which means Google had to increase accelerated compute dedicated to token generation by 3-10x. Our estimate is that Google increased accelerated compute by ~3x during the period - which means that they had to pull compute from training, recommenders, etc to allocate to token generation. Significant algorithmic advancements (Flash Attention, quantization, MoE), and infrastructure investments (prompt caching, batching) have driven much of that efficiency, but counting on the hardware to get better is something the industry is counting / relying on. We used to be able to ride Moore's Law / Dennard scaling to improve compute per watt. But now... we have to rely on $NVDA / hardware ecosystem (Google, $AMD, etc) to drive 2-4x improvement per generation (Huang's Law). People underestimate the strain of exponentials on human systems - that are hard coded to think linearly... The total global accelerated compute base is probably ~8 GW of installed capacity on my math and analysts have estimated growth to increase to 10x to ~80 GW globally by 2030. Even assuming all of that capex gets done, and Nvidia continues to push yearly roadmap (generating 2x y/y performance uplift), the 10x power increase should equate to ~50x increase in compute. We just had 100x AI usage increase in 1 year from Google's testimony. OpenAI, Google, Anthropic are in structural shortage of compute - each bit they bring online is fully invested in serving their users. And that is without even expanding into true video / world models, robotics, or long horizon thinking to find novel breakthroughs. What could change this trajectory? If the algorithmic efficiencies we are gaining from new breakthroughs outstrip the exponential increase in current demand -- and FUTURE demand. Investors hyperventilated at DeepSeek's release earlier this year, but their gains were outweighed by the increase in demand created by reasoning.

Clark Tang

176,622 просмотров • 10 месяцев назад

David Sacks just said what every honest analyst in Silicon Valley is already thinking (Save this). Nobody has ever seen anything like this. Anthropic has grown at 10x per year for three straight years and going into 2026, the conventional wisdom was that the rate of growth had to slow at this level of scale but then the numbers came in. Q1 alone is $10B ARR to $30B, in April, $30B to $44B and that's $96 million in new ARR added every single day. Inference margins are now above 70%, up from 38% last year and the only thing holding them back was compute. That's solved now, the SpaceX deal and others Anthropic has been quietly signing unlocks the supply side. This is exactly why we are bullish on Nebius and AMD. When a single company is adding nearly $100M in ARR per day, the real trade isn't the frontier lab but rather the infrastructure underneath it. Nebius, one of the fastest-growing neoclouds on the planet posted 547% YoY revenue growth in Q4 2025, exited the year with $1.25B ARR, and is guiding for $7–9B ARR by year-end 2026. Their revenue backlog has reached $46B, with projections of $16B in revenue by 2028 and NVIDIA locked in a $2 billion stock buy agreement with them giving Nebius early access to cutting-edge chips while every other cloud scrambles for supply. AMD is the other side of the same coin. Data center revenue hit $5.78B in Q1, up 57% year-over-year with total company revenue at $10.25B, up 38%. Meta has committed to deploying up to 6 gigawatts of AMD Instinct GPUs. Data center GPU revenue is forecast to surge 114% year over year to $15B in 2026. MI400-series chips hit the market in H2 and analysts project segment operating margins climbing to 31% as the next generation ramps. The model is simple, Anthropic is printing revenue and that that revenue pays for compute. That compute flows through companies like Nebius and AMD. This is why Milk Road PRO remains bullish on them and our positions are up massively. Our analysts have broken down the full thesis, the allocations, and the price targets. Go PRO at Milk Road to see everything, link below!

Milk Road AI

184,643 просмотров • 2 месяцев назад

NVIDIA is about to grow a new moat that has the potential to be the company’s most impenetrable moat yet: The Compute Futures Market One of the reasons why there’s so much demand for U.S. T-bills is because the market for them is incredibly deep, and large pools of capital can come in and out freely without much disturbance. That attribute is attractive, which leads to more demand for T-bills Liquidity begets liquidity Another example is what Bill Ackman is describing in the attached clip - the more valuable a company becomes, the easier it is for said company to raise capital to fund expansion, which in of itself is virtuous and valuable The same concept is going to apply to compute, and in some ways- it already is. NVIDIA GPUs are already the most financeable (in many cases the only financeable) form of compute for various parties to deploy in their data centers. But the introduction of a forward curve that facilitates financial expression like hedging takes that concept to a different level Companies who consume compute (structurally short compute) will want to be able to hedge their input costs, and companies who produce compute (structurally long compute) will want to be able to hedge their output price. The market will coalesce even more around the deepest / most liquid pools to facilitate this - which will be NVIDIA based futures (WTI Crude) This liquidity itself will become part of the value of purchasing NVIDIA equipment, it will be part of the justification to pay the “Nvidia tax”. And when the liquidity reaches a certain depth, it will be almost impossible to break - because to break it would involve coordinating between massive amounts of misaligned parties- impossible So NVIDIA, by no additional virtue or effort of their own will likely absorb another massive structural advantage that may be the most difficult advantage to overcome for any competitor out of anything that has been built so far. The rich really do get richer.

Nick Dorsey

33,944 просмотров • 21 дней назад

The CEO of the world's largest asset manager just said something that should reframe how every investor thinks about the AI trade. Larry Fink, managing $11.5 trillion at BlackRock, stood at the Milken Institute Global Conference and said four words that matter, "We just don't have enough compute." "The United States is short power. We're short compute. We're short chips. And there's going to be shortages in all three and memory, four things. I actually believe a new asset class will be buying futures of compute." Think about what that means. Fink is predicting that compute becomes a tradable commodity like oil, like grain, like natural gas where investors buy forward contracts on future capacity because the shortage is so structural and so predictable that a derivatives market will emerge to price it. That is not a minor observation from a finance executive but rather the chairman of the most powerful capital allocator on the planet telling you that compute scarcity is a multi-year, investable megatrend. The data backs him up completely. Data centers will consume 70% of all memory chips produced globally in 2026. Advanced HBM production from Samsung, SK Hynix, and Micron is sold out through 2026 and into 2027 and a single AI server consumes 10-20x more memory than a conventional workload server. DRAM supply growth is running at just 16% annually while AI infrastructure demand is growing at 80%+. The chip crunch, the power crunch, and the compute crunch are not temporary dislocations, they are structural, and they will get worse before they get better. Fink also said something the bears keep getting wrong: "There is not an AI bubble. There is the opposite. We have supply shortages. Demand is growing much faster than anyone has ever anticipated." This is why the Milk Road Pro portfolio is built the way it is, long the companies producing and supplying the constrained resources: chips, memory, compute infrastructure, and power. Check out Milk Road Pro, link below to access our full thesis and plays.

Milk Road AI

419,283 просмотров • 2 месяцев назад

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 просмотров • 1 месяц назад

.Dylan Patel lays out how we know the hard upper bound on how much compute can be produced annually by 2030: around 200 GW/year. That’s a crazy number (there’s about 20 GW of AI deployed in the world right now), but it’s nowhere near enough to satisfy Sam/Elon/Dario/Demis’s ambitions. Lots of things in the supply chain can be scaled up over 4 years, including things that other people think are bottlenecks, like datacenter power or fab clean room space. But the thing that’s inflexible over that timeline is the number of EUV tools. Dylan forecasts that production of ASML’s EUV tools will scale from 60 per year now to about 100 per year by the end of the decade - which means something like 700 total machines running in 2030. For a fab to make a GW worth of the Rubin chips that NVIDIA is deploying later this year, it needs to make 55,000 3nm wafers, 6,000 5nm wafers, and 170,000 memory wafers. Each 3nm wafers needs about 20 EUV passes, so about 1.1 million passes per GW. Adding on 5nm and memory, you need two million passes. Each tool can do 75 passes per hour, so with 90% uptime that’s around 600k passes per year - so a single machine can make less than a third of a GW in a year. So in 2030, we have 700 total machines, each making 0.3ish GW a year, which means we can produce 200 GW of compute a year. That’s a lot. But Sam Altman wants a gigawatt a week by the end of the decade. Anthropic and Google will be wanting about the same. And Elon wants to be putting 100 GW in space every year. Any one of these players could maybe get what they need, but not all of them.

Dwarkesh Patel

114,538 просмотров • 4 месяцев назад

Demis Hassabis just explained why the real AI bottleneck has nothing to do with training runs. Most people picture the AI arms race as who can build the biggest model. GPT-4 or Gemini Ultra style training runs, a few hundred million in compute, fired once or twice a year. The constraint sits somewhere else. Every time a researcher has a new algorithmic idea, a new architecture, a new training technique, they can't just test it on a laptop. They have to run it at the scale where it would actually be deployed, because ideas that look promising at small scale fall apart completely when you put them into a real system. Every research hypothesis burns significant compute before a single line of production code gets written. At a lab like DeepMind, hundreds of researchers are running hundreds of ideas simultaneously. The demand for experimental compute is continuous. It never stops. Now layer the hardware reality on top. GPU lead times are currently 36 to 52 weeks for data center hardware. Global AI data centers are already drawing 29.6 gigawatts, equivalent to the peak power demand of the entire state of New York, and they still can't meet demand. Companies willing to pay any price can't just buy more compute. They wait in line. The speed of scientific discovery in AI is now gated by hardware availability. The next breakthrough is sitting in a researcher's head right now. Whether it gets validated fast enough to matter depends entirely on whether the compute is there when they need it. The AI race gets won by whoever can run the most experiments per month.

Aakash Gupta

32,150 просмотров • 3 месяцев назад