Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

The defining differentiator in AI right now isn't model performance, it's cost per token. Because in this market, efficiency is what turns demand into actual growth. Roman Chernin of Nebius outlines why inference economics are becoming the central challenge, as companies operating on thin margins need infrastructure that can...

294,803 Aufrufe • vor 4 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Interview with Nebius Co-Founder Roman Chernin Please like & share this video so that all $NBIS investors on X will see it! :) If you prefer watching on YouTube: Timestamps: 00:00 - Why AI Infrastructure Is So Hard to Understand 00:24 - Market Fragmentation and What Actually Differentiates Providers 01:30 - Consolidation, Segmentation, and the Future AI Cloud Landscape 02:56 - What Analysts and VCs Still Get Wrong About AI Infrastructure 05:34 - Nebius Cloud: Product Readiness and Customer Proof Points 07:42 - Why Inference Workloads Are Exploding 09:11 - Training vs. Inference: How AI Models Actually Reach Production 10:10 - Why Inference Market Share May Concentrate Around a Few Winners 12:36 - Customer Use Cases: Coding, Enterprise AI, and Real-World Adoption 14:01 - Why Integrated Training and Inference Matter Strategically 16:01 - Building Scalable AI Infrastructure With High Utilization 18:24 - Token Factory: Inference as a Managed Service 20:24 - Revolut Case Study: AI-Driven Product Enhancements 22:56 - Token Factory Performance Optimization and Competitive Advantage 25:07 - Scale, Capacity, and Efficiency as Growth Drivers 28:36 - Why Inference Capacity Could Become the Next Major Bottleneck 30:10 - How Nebius Benchmarks Performance Across Providers 33:14 - The Future Size and Shape of the Inference Market 36:38 - Value-Based Pricing: Moving Beyond Cost per GPU Hour 40:55 - How Nebius Wins Deals: Quality, Performance, and Customer Experience 44:53 - Autonomous AI Platforms and the Rise of Agent-Based Models 47:28 - Tavily, Agentic Applications, and the Next Layer of the AI Stack 50:45 - Strategic Trade-Offs: Scaling, Product Roadmap, and Customer Relevance 55:40 - Final Thoughts: Adapting to the Next Shift in AI Workloads Nebius Roman Chernin

Daniel Koss

203,706 Aufrufe • vor 3 Monaten

Mark my words, Nebius will be the first Trillion dollar Neo-cloud company and here is why (Save this). Roman Chernin, CEO of Nebius just said on 20VC that Nebius raised prices and demand didn't move. When a company can raise prices and still have more demand than supply, that's the opportunity. Chernin also explained why he is deliberately not charging the maximum. As AI shifts from training, a one time cost to inference, which is the ongoing cost of serving every user and every query, compute pricing becomes the cost structure of the entire AI economy. If Nebius prices customers out, those customers cannot grow, and Nebius cannot grow with them. That is the compounding flywheel built directly into the revenue model. The numbers are already confirming it. Q1 2026 revenue came in at $399 million, up 684% year over year. The AI cloud segment grew 840% and represented 98% of total revenue. Adjusted EBITDA flipped positive to $129.5 million. And Nebius signed a long-term agreement with Meta worth up to $27 billion over five years, a hyperscaler outsourcing its own AI compute stack to a neocloud, which tells you that even companies with $50 billion capex budgets cannot build fast enough. Goldman Sachs says the consensus is underestimating 2027 hyperscaler capex by $500 billion. Every dollar hyperscalers cannot provision themselves flows to neoclouds like Nebius. As that gap widens, Nebius captures the overflow with 3 gigawatts of contracted power already secured and a CEO who just told you raising prices did not dent demand. Our subscribers are already up massively on Nebius and come join Milk Road Pro for our full breakdown, how to size Nebius against the broader neocloud opportunity, and our full AI thesis. Link below!

Milk Road AI

15,677 Aufrufe • vor 2 Monaten

This is the biggest irony in tech history. Microsoft beat revenue estimates. Stock plunged 11%, wiped out $400 BILLION in market cap. Salesforce reported growth. Stock fell 5.6%. ServiceNow beat earnings. Stock crashed 11%. SAP beat projections. Stock dropped 16%. Entire software sector entered bear market territory. Down 22% from peak. These are the companies everyone said would WIN from AI. They spent billions BUYING AI companies. ServiceNow: $7.75 billion for Armis. Salesforce: $8 billion for Informatica. They launched AI products. Built AI workflows. Hired AI teams. And the market said: You're all dead. Because investors just realized something nobody wanted to admit: AI doesn't make software companies stronger. AI makes software companies OBSOLETE. Morgan Stanley: "In an environment of heightened investor skepticism, stable growth falls short of shifting the narrative." Good earnings aren't enough anymore. The market is pricing in a world where AI replaces the software these companies sell. ServiceNow CEO tried defending on the earnings call: "AI needs workflow orchestration. ServiceNow is the gateway to this shift." Market response: 11% crash. Because here's what he didn't say: If AI can write code, automate workflows, and generate apps at a fraction of the cost, why would anyone pay $50,000 per year for enterprise software licenses? The per-seat pricing model that made SaaS companies rich is getting murdered by AI efficiency. One AI agent replaces 10 seats. One prompt replaces months of custom development. One LLM call replaces entire software categories. Klarna already proved it. CEO said they pulled Salesforce out of their stack. Built everything themselves using AI. And that's just the beginning. The software apocalypse hit hardest on companies that INVESTED IN AI: Atlassian: down 12.6% Intuit: down 7.8% HubSpot: down 11.5% Zscaler: down 6.3% Meanwhile, the companies ENABLING AI made money: Nvidia: up Semiconductor stocks: surging Memory firms: rallying The divide is brutal. Hardware companies print cash. Software companies get destroyed. Because in an AI-first world, you need GPUs to build the models. But you don't need software subscriptions when the AI builds the software for you. Jim Cramer called it the "P/E multiple compression crisis." Translation: Investors don't care about earnings anymore. They care about whether your business model survives the next 5 years. And right now software business models look doomed. They're literally stuck: If they DON'T invest in AI, they fall behind. If they DO invest in AI, they cannibalize their own products. It's a death spiral with no exit. ServiceNow spent $12 BILLION on acquisitions in 2025 alone. Trying to buy their way into relevance. And yesterday the market cooked them. The craziest thing to me tho... Most software companies beat earnings. Revenue was solid. Growth was fine. But it didn't matter. Because the market stopped pricing software on what it earns TODAY. It's pricing software on what it's worth in a world where AI does the job for free. And in that world these companies are worth nothing. This is the biggest sector repricing since 2008. $500 billion in market value gone in ONE DAY. And it's not stopping. Because every company watching this is thinking the same thing: "If I can replace ServiceNow with 3 AI agents and save $10 million per year, why wouldn't I?" The answer used to be: "Because you need enterprise-grade reliability." But now? AI agents are getting reliable. Fast. Software companies just realized they're competing with open-source models that cost $0.02 per 1,000 tokens. You can't win a pricing war against free. The companies that spent BILLIONS preparing for AI are getting killed BY AI. What an irony.

Ricardo

1,815,322 Aufrufe • vor 6 Monaten

Mark Zuckerberg is explaining one of the most misunderstood dynamics in AI and it has direct investment implications (Save this). The concept he's describing is model distillation, and it's one of the most important techniques to emerge in AI over the past year. Here's how it works. You train a massive, enormously expensive model, in Meta's case, Llama 4 Behemoth, a 2 trillion parameter teacher model and then you use that model to teach a much smaller, cheaper model. The smaller model inherits roughly 90 to 95% of the intelligence of the giant while running at 10% of the cost and on a fraction of the compute. Meta already did this with the Llama 4 family and Behemoth serves as the teacher. Llama 4 Scout and Maverick, the publicly released open-source models were distilled from it. Scout runs on a single H100 GPU with a 10 million token context window and outperforms models that cost far more to operate. Maverick, at 17 billion active parameters, rivals DeepSeek V3 in coding at half the parameter count and beats GPT-4o on multimodal benchmarks. Both are completely free for commercial use. What Zuckerberg is pointing at is a structural shift in how AI gets deployed in the real world. Companies aren't taking a frontier model off the shelf and running it as-is but rather taking open-source models, fine-tuning them on their own proprietary data, distilling them into even smaller custom models tailored to their specific use case, and running them on infrastructure they control at a fraction of the cost of a closed frontier API. The investment implication of this is significant and runs in two directions. For Meta specifically, this is a strategic masterstroke. Every company that builds on Llama, fine-tunes it, distills it, or deploys it through their infrastructure is pulling into Meta's orbit while Meta builds the most powerful open teacher model. The ecosystem of companies using it grows and that ecosystem generates commercial activity across Meta's platforms and data services. Meta's AI research benefits from billions of real world deployment signals and it's a flywheel that closed model providers cannot replicate because their strategy requires charging per token, which is now a 65x cost disadvantage against the open-source alternative. For the broader market, distillation changes the economics of inference in a way that has barely been priced in. As intelligence becomes extractable into smaller and cheaper models, the absolute demand for compute doesn't decline but rather it explodes, because now the number of applications that are economically viable expands by orders of magnitude. Every task that was previously too expensive to automate at $3.25 per call becomes viable at $0.05 that means more total token usage, more total GPU utilization, and more demand for the infrastructure companies, the Nebiuses, the GE Vernovas, the Constellation Energies that supply the underlying compute and power.

Milk Road AI

27,908 Aufrufe • vor 1 Monat

Bloomberg warns that China’s AI price war may make profitability difficult for years. They should be more worried about Wall Street. If the price war continues, the real casualty may not be AI. It may be the entire valuation structure built around it. OpenAI and Anthropic are valued as though frontier intelligence will remain scarce, expensive, and capable of producing software-like margins. Hyperscalers are spending hundreds of billions on infrastructure based on that assumption. Nvidia’s valuation depends on those capital expenditures continuing for years. But Chinese models are driving token prices toward commodity levels. Usage can explode while revenue per token collapses. The servers become busier. The models become cheaper. The profits fail to appear. Nvidia has real revenue and real margins, so it is not the weakest link. But even Nvidia is priced on the belief that today’s extraordinary AI capital expenditure will remain economically justified. If cheaper models, better efficiency, and open weights destroy the expected returns on that infrastructure, Wall Street will not merely reprice OpenAI and Anthropic. It will reprice the entire AI chain. America built the AI bubble on Nvidia. Nvidia built its empire on TSMC. Taiwan built its economic security around TSMC. If Chinese efficiency destroys the economics of brute-force AI, the repricing will not stop in Silicon Valley. It will cross the Pacific and land in Hsinchu. AI may continue transforming the world while the AI bubble collapses. The internet survived 2000. The valuations did not.

𝘊𝘰𝘳𝘳𝘪𝘯𝘦

13,927 Aufrufe • vor 19 Tagen

$PLTR $AMD | Dr. Karp and Dr. Su were right! ✍️ Companies are now fighting back. Dr. Karp, Palantir CEO, recently told CNBC that enterprises are privately "unhappy" with frontier AI labs like OpenAI and Anthropic, accusing them of prioritizing "tokenmaxxing" or maximizing AI token consumption to signal activity over delivering real business value and understanding customer needs. Uber, Coinbases routing to capping token usage or routing to cheaper models to keep cost under control. or Microsoft revoked Claude Code licenses companywide, Priceline imposed token limits after sharp cost spikes, and reports cite Meta, Salesforce, and multiple unnamed firms facing 3x+ budget overruns or $ hundreds of millions in unexpected spend by mid-2026. Analysts note this as an emerging industry pattern, with FinOps and executives describing "existential crises" over token bills; dozens of enterprises are now adding guardrails, though public complaints remain concentrated among high-profile tech firms experimenting at scale. Dr. Lisa Su anticipated the pivot to inference economics and CPU-dense systems for agentic AI, correctly predicting that token costs, power efficiency, and deployability on standard platforms would determine scalable adoption long before the current enterprise pushback. Dr. Alex Karp accurately diagnosed the disconnect in frontier labs' approach, calling out "tokenmaxxing" as activity without outcomes; enterprises are indeed demanding real implementation and business-specific value rather than raw volume that inflates bills without proportional ROI. Together, their independent foresight validates the maturing AI thesis, efficient infrastructure (AMD Helios/EPYC optimized for lowest TCO & $/M Tokens) paired with outcome-focused platforms (Palantir AIP/Foundry) positions both companies to benefit as the market shifts from hype-driven consumption to sustainable, value-driven deployment. Yes it may look good on the revenue growth for AI Labs to show off on IPOs investors/bankers, but the customers have to find value in those tokens spent where $NVDA & In-house chips on inference claims are just false. At the end of the day, ~Token cost needs to go down more & more particularly inference by owning more AMD chips/racks. In-house chips can make all kind of claims for years, but the bills enterprises paid have to obey economic. ~Enterprises want a thick software OS or solution focused, they do not want to have unlimited budget for "tokenmaxxing" where it is leading to high costs with limited business transformation; success increasingly depends on implementation layers that route tasks, enforce policies, and connect AI to existing workflows. Not Financial Advice! DYOR!

Mike

253,557 Aufrufe • vor 1 Monat

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 Aufrufe • vor 6 Monaten