Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🔥 NVIDIA GTC 2026: The Vera Rubin platform just dropped, introducing 10x inference throughput per watt, at 1/10th the cost per token vs. Blackwell. Jensen Huang called it: every company will now run a "token factory." Here's the thing: Aethir has been running distributed token factories since day one,...

15,420 Aufrufe • vor 5 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Jensen Huang just doubled NVIDIA's demand forecast to $1 Trillion through 2027 🤯 Then spent two hours explaining why that number is conservative… Here's everything today from GTC: - NemoClaw: NVIDIA's open-source enterprise AI agent stack built around OpenClaw. Jensen called OpenClaw "the operating system for personal AI" and said every company needs a strategy for it. - Space-1: NVIDIA is putting Vera Rubin data centers in orbit. Not a concept. An actual system being designed for space deployment right now. - DLSS 5: 3D-guided neural rendering that blends raw graphics with generative AI. Jensen called it the future of real-time rendering. - AWS: Deploying 1 million+ NVIDIA GPUs starting this year. Azure was the first hyperscaler to power up Vera Rubin. - Vera Rubin: NVIDIA's next-gen AI supercomputer. 10x more performance per watt than Blackwell, 700 million tokens per second, shipping later this year. - Groq 3 LPU: First chip from NVIDIA's $20B Groq acquisition. A purpose-built inference accelerator that ships Q3. NVIDIA now owns training AND inference. -Feynman: The architecture after Rubin, coming 2028. New GPU, new LPU, new CPU. NVIDIA is on a 12-month chip cadence and the treadmill never stops. - Autonomous driving: BYD, Hyundai, Nissan, and Geely building Level 4 vehicles on NVIDIA. Uber deploying NVIDIA-powered robotaxis across 28 cities by 2028. The man doubled his demand forecast to a trillion dollars, announced data centers in space, and closed the show with a robot singing country music. This is NVIDIA's world. Everyone else is just renting compute in it.

Josh Kale

45,875 Aufrufe • vor 5 Monaten

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 Aufrufe • vor 2 Monaten

AMD might have disrupted Nvidia's entire cloud GPU rental business. In January at CES, AMD CEO Lisa Su demonstrated a $1,499 mini PC running the same class of AI model that currently costs companies $2,500 to $3,000 every month to rent from Nvidia-powered cloud servers. AMD's own branded version opened pre-orders this month at $3,999. Third party manufacturers have been selling the same chip since 2025 starting at $1,499. Here is exactly why this is dangerous for Nvidia. Nvidia's $75 billion quarterly revenue is built almost entirely on one business model, companies rent access to Nvidia GPUs through cloud providers like AWS and Lambda Labs to run AI. They pay monthly. Nvidia gets paid every time someone runs an AI model in the cloud. That recurring rental income is what turned Nvidia into a $5 trillion company. The AMD box eliminates that monthly fee permanently. One AI consultant switched from $2,800 per month in Nvidia cloud rental costs to $8 per month in electricity. The hardware paid for itself in 11 days. Over 8 months he generated $47,000 running the same AI workloads that previously left him paying Nvidia's ecosystem $2,800 every single month. Multiply that across thousands of enterprise customers and the revenue erosion becomes structural. Every business that buys this box stops paying cloud rental fees forever. Lawyers, doctors, banks, accountants, and financial advisors, businesses with sensitive data that cannot legally go to a cloud server represent billions in annual cloud GPU fees that Nvidia is now at risk of losing permanently. The threat is also closing in from the top. Google signed deals worth tens of billions with Anthropic and Meta to replace Nvidia with its own chips. Amazon built its own AI chips across AWS. Apple trained its AI on Google's chips, not Nvidia's. Custom silicon has grown from 21% of the AI chip market in 2025 to 28% in 2026. Nvidia's rental model only worked because serious AI compute had no alternative.

Bull Theory

26,765 Aufrufe • vor 2 Monaten

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 Aufrufe • vor 6 Monaten

Introducing KausaCompute. Running AI models privately is expensive and complicated. Traditional cloud providers require credit cards, KYC verification, and complex setup processes. For many developers, especially in emerging markets, accessing GPU compute remains out of reach. KausaCompute changes that. Deploy any Docker container with NVIDIA GPU, pay with USDC from Maze Pocket, no KYC, no credit card, no cloud provider account needed. GPU pricing starts at $0.47/hr with competitive rates across all tiers. KausaCompute lives inside KausaLayer Pocket. Open a pocket, swap SOL to USDC using the built-in swap feature, and start deploying GPU containers right away. Everything stays within one ecosystem. What it actually does: Private LLM endpoints. Deploy Llama 3, Mistral, CodeLlama, or any open-source model as a personal API. No one logs the prompts. No one reads the data. Full control. AI coding assistants that never see external servers. Speech-to-text processing for thousands of audio files in minutes instead of hours. Sentiment analysis across millions of data points. Document summarization at scale. Multi-modal AI that processes images and text together. All running on dedicated NVIDIA GPUs. A6000, T4, L4, L40, A100, H100, and more. Pick the hardware, pick the duration, deploy in one click. The billing is straightforward. USDC is deducted from Maze Pocket before deployment. Pro tip: Maze Pocket supports multiple pockets per wallet. Create a dedicated pocket just for KausaCompute to keep GPU spending separate from other activities like trading or transfers. Clean separation, easy tracking. KausaCompute is not another cloud provider. It is the fastest path from USDC to a running GPU, with zero identity requirements. Live now at

KausaLayer

10,914 Aufrufe • vor 3 Monaten

Jensen Huang just admitted the biggest AI labs can't borrow money like normal companies. So Nvidia signs for them, and they spend it on Nvidia chips. Nvidia reported Wednesday and the numbers are absurd: Revenue of $96.2 billion, up 106%, with net income of $59.7 billion, the most profitable quarter any public company has EVER posted. And Huang just told Fox Business that every chip Nvidia can make next year is already sold. Here's why this matters the most: Huang wrote this himself about his own customers: "Frontier AI labs have extraordinary demand for training and inference compute, but many are growing faster than their balance sheets and long-term credit profiles can support." Then: They "still lack the decades-long infrastructure contracts and investment-grade financing capacity needed to secure the AI factory infrastructure independently." Put simply: His customers can't get the loans. So Nvidia signs for them. There's a compute campus going up in Ohio with OpenAI as the tenant. Nvidia has tied roughly $105 billion in commitments to it. OpenAI's existing and planned commitments now come to about 12 gigawatts of Nvidia compute. CFO Colette Kress told analysts Nvidia will also provide selective credit enhancement for nearly 2 gigawatts of compute at a second frontier lab. She wouldn't say which one. Nvidia put up to $10 billion into Anthropic in November at a valuation near $350 billion, and Anthropic agreed to buy up to a gigawatt of Grace Blackwell and Vera Rubin systems in the same deal. And Nvidia isn't only guaranteeing these companies. It OWNS pieces of them. This week's filing shows $18 billion committed to equity investments for the rest of the fiscal year, and $47.9 billion already sitting in private companies as of late July. Now here's where it gets really insane: Last week, Huang sat on a CNBC set surrounded by six of Wall Street's biggest firms. Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR. They signed a memorandum to mobilise up to $500 billion in outside capital for AI data centres. Nvidia kept the option to backstop up to a quarter of those deals. And Huang used that stage to announce that Nvidia GPUs are now an asset class. Pension and credit funds can now lend against graphics cards the way they lend against office towers. Kress saw the accusation coming and got ahead of it on the earnings call: "We recognise the scale of this support, and we know some will call this circular financing. We see it differently." But look at the two things Huang says about the same companies. On the earnings call he said AI has hit its inflection point, that the tokens are productive and profitable, and that compute is now revenue. But he also said those same labs can't secure investment-grade financing on their own. A business that's inflecting into profit is exactly the business a bank lends to. Banks lend against cash flow every day. But Nvidia‘s guarantee exists because something in that first story isn't landing with the people whose job is pricing risk. Kress does have a real answer to this though. She said the second lab's credit support only complements capacity it already secured on its own, without Nvidia backing it. Vendor financing is also old and legal. Cisco did it and GE built a finance arm on it. Huang's case is that Nvidia understands these businesses better than any lender could, and he says the risk is low and his only regret is not investing more and sooner. He may be completely right. But one thing is certain: Nvidia guarantees the paper. The paper buys the chips. Nvidia books the sale. Then Nvidia tells you the order book is full for a year. That order book is the entire argument for a $5 trillion company. And Jensen Huang just explained, in his own words, that his customers couldn't have written those orders without him. Isn’t this suspicious?

Ricardo

62,605 Aufrufe • vor 6 Tagen

Jensen Huang just BROKE the most important rule in the industry. And it explains why Nvidia controls 95% of the AI chip market. Last night at CES, he unveiled Vera Rubin - the new AI supercomputer that's shipping right now. Full production started weeks ago. But here's the part that made every semiconductor engineer in the room go crazy: Reuben GPU is 5x faster than Blackwell. But only has 1.6x the transistors. That should be physically impossible. Moore's Law says you get maybe 25% more performance per transistor generation. Jensen just delivered 300%. How? He BROKE the most sacred rule in chip design. The rule every company follows: "Never redesign more than 1-2 chips per generation." Nvidia redesigned all six chips simultaneously. Vera CPU. Reuben GPU. Connect X9 networking. Bluefield 4 DPU. MVLink switches. Spectrum X Ethernet. Every. Single. Component. From scratch. He calls it "extreme co-design." The industry calls it insane. One rack now moves 240 terabytes per second. That's TWICE the entire global internet bandwidth. In a single rack. And it runs on 45°C water - no chillers needed. Which saves 6% of global data center power. But the real story isn't the hardware... It's what they're doing with it. Nvidia just open-sourced Alpha Mayo. The world's first reasoning autonomous vehicle AI. Mercedes-Benz CLA launches with it in Q1. Europe Q2. Asia by year-end. Not a concept car. Not a limited release. Full production vehicles. And the AI will even explain its reasoning out loud. "I'm slowing down because the truck ahead is braking and there's a cyclist merging." It thinks. Then tells you what it's thinking. Then executes. Jensen drove it through San Francisco for an hour yesterday. No hands. No interventions. Through heavy Sunday traffic. The whole thing is open source now. Every line of training code. Every data source. The entire stack. But why would Nvidia give this away? Because they learned something from the last year: Open models activated the entire world. DeepSeek R1 proved open source can hit the frontier. Downloads exploded. Every country, every startup, every researcher can now build AI. And they all need Nvidia hardware to train it. That's the strategy. Give away the recipes. Sell the kitchen. The partnerships tell you where this is going: Siemens is integrating Nvidia into every industrial design tool. Cadence and Synopsys are rebuilding chip design around Nvidia. Palantir, ServiceNow, Snowflake - their entire platforms now run on Nvidia's agentic AI stack. This isn't just selling chips anymore. Nvidia is rebuilding the entire computing stack. From design to manufacturing to deployment. Every layer of the trillion-dollar AI infrastructure buildout runs through them. And now they're 18 months ahead of everyone else. Again. The competition is still trying to match Blackwell. Nvidia's already shipping the thing that makes Blackwell look slow. What do you think - is anyone catching them? The only company capable of this might be Google.

Ricardo

580,166 Aufrufe • vor 7 Monaten

Chamath Palihapitiya just dropped the number that explains the entire AI infrastructure trade (Save this). A gigawatt of compute now costs $100 billion and when he started his Arizona data center project it was $4 to $5 billion, it has gone up 20x in a single investment cycle. The implication is not just that AI infrastructure is expensive but rather that the capital barrier to owning meaningful compute has become so high that only a handful of entities in the world can actually build it and the companies who got there early are sitting on what may be the most durable pricing power in the history of the technology industry. This is the neocloud trade. The neocloud market, purpose-built GPU cloud providers like CoreWeave, Nebius, and Lambda Labs was worth $35 billion in 2026 and is projected to reach $236 billion by 2031, compounding at 46% annually. For context, that is faster growth than cloud computing itself posted in its first decade. The reason is very simple, hyperscalers like AWS, Azure, and Google are building for everything, storage, databases, enterprise software, networking and their GPU pricing reflects the overhead of that full-stack infrastructure. Neoclouds build for one thing only, AI compute. The result is a 60% to 85% cost advantage on the same Nvidia silicon, bare metal H100s at $0.78 to $2.79 per GPU-hour on a neocloud versus $3.43 to $5.07 per GPU-hour on a hyperscaler. That spread does not close as AI demand scales but rather it widens, because hyperscalers have to amortize legacy infrastructure and margin expectations that neoclouds do not carry. Gartner projects that by 2030, neoclouds will capture 20% of the $267 billion AI cloud market, and Vultr's own analysis says at least 80% of GPU market share by end of 2026 will be held by a small group of scaled neocloud providers. Now zoom into Nebius specifically, because it is the most interesting publicly traded proxy for this trade. Nebius is the infrastructure arm of the former Yandex Russia's equivalent of Google rebuilt from the ground up after Russia's invasion of Ukraine by Arkady Volozh and relisted on Nasdaq in October 2024. The team that built it already knew how to run internet-scale infrastructure at the lowest possible cost, which is exactly the operational DNA a neocloud requires. In Q1 2026, Nebius reported revenue of $399 million and already generating serious cash on a young business with revenue growing nearly eightfold year-over-year. Then in March 2026, Meta signed a five-year infrastructure agreement with Nebius worth up to $27 billion, $12 billion in committed dedicated GPU capacity deployments beginning early 2027, plus up to $15 billion more tied to Meta purchasing Nebius's unsold third-party capacity. The deal will be executed on one of the first large-scale deployments of Nvidia's Vera Rubin platform, the next-generation architecture after Blackwell making Nebius one of a tiny number of operators in the world with confirmed priority access to the most advanced AI hardware available. Following the contract, Nebius guided to $7 to $9 billion in annualized recurring revenue for 2026 representing 540% year-over-year growth. Chamath Palihapitiya point about the $100 billion capital moat is the bear case for new entrants and the bull case for incumbents. No one can afford to build the next CoreWeave or Nebius from scratch at current hardware and power costs. The companies that are already built, already contracted, and already deploying Nvidia's latest silicon have a moat that compounds with every GPU generation cycle because they get allocations first, they deploy fastest, and their customers re-sign rather than wait for a new operator that does not yet exist. Come join Milk Road Pro for our full breakdown, the complete neocloud competitive landscape, how to think about Nebius's valuation versus CoreWeave and AI entire thesis. Link below.

Milk Road AI

139,047 Aufrufe • vor 2 Monaten