Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Most AI projects don't fail because models aren't good enough. They fail because inference economics don't scale. During #CXOSpice, Roman Chernin Nebius broke it down: ✅Closed models kill margins ✅Open models unlock control, if you can run them well ✅DataData only becomes a moat when inference is sustainable One...

13,662 Aufrufe • vor 6 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Interview with Nebius Co-Founder Roman Chernin Please like & share this video so that all $NBIS investors on X will see it! :) If you prefer watching on YouTube: Timestamps: 00:00 - Why AI Infrastructure Is So Hard to Understand 00:24 - Market Fragmentation and What Actually Differentiates Providers 01:30 - Consolidation, Segmentation, and the Future AI Cloud Landscape 02:56 - What Analysts and VCs Still Get Wrong About AI Infrastructure 05:34 - Nebius Cloud: Product Readiness and Customer Proof Points 07:42 - Why Inference Workloads Are Exploding 09:11 - Training vs. Inference: How AI Models Actually Reach Production 10:10 - Why Inference Market Share May Concentrate Around a Few Winners 12:36 - Customer Use Cases: Coding, Enterprise AI, and Real-World Adoption 14:01 - Why Integrated Training and Inference Matter Strategically 16:01 - Building Scalable AI Infrastructure With High Utilization 18:24 - Token Factory: Inference as a Managed Service 20:24 - Revolut Case Study: AI-Driven Product Enhancements 22:56 - Token Factory Performance Optimization and Competitive Advantage 25:07 - Scale, Capacity, and Efficiency as Growth Drivers 28:36 - Why Inference Capacity Could Become the Next Major Bottleneck 30:10 - How Nebius Benchmarks Performance Across Providers 33:14 - The Future Size and Shape of the Inference Market 36:38 - Value-Based Pricing: Moving Beyond Cost per GPU Hour 40:55 - How Nebius Wins Deals: Quality, Performance, and Customer Experience 44:53 - Autonomous AI Platforms and the Rise of Agent-Based Models 47:28 - Tavily, Agentic Applications, and the Next Layer of the AI Stack 50:45 - Strategic Trade-Offs: Scaling, Product Roadmap, and Customer Relevance 55:40 - Final Thoughts: Adapting to the Next Shift in AI Workloads Nebius Roman Chernin

Daniel Koss

203,706 Aufrufe • vor 3 Monaten

✈️ Starting the day in Istanbul. Let's talk AI. The future of AI won't be shaped by size, but by precision. Lately I've been diving into what OpenLedger has been building, and I think we're witnessing one of the most important hard forks in AI: 👉 From giant generalist models 👉 To focused, hightrust AI agents Here's why that shift matters , and why OpenLedger's vision makes perfect sense: ✅ Specialized > Generalized General AI can do many things. But specialized AI? It does one thing extraordinarily well. Tailored models don't waste compute on irrelevant context, every parameter is purposedriven. ✅ Explainability isn't optional anymore In highstakes sectors like finance or healthcare, because the model said so won't cut it. We need transparent reasoning paths. Models must show how they reached conclusions , not just what they concluded. ✅ Trust comes from traceability With OpenLedger's Proof of Attribution, each AI decision is traceable, verifiable, and tamperproof. We're talking onchain records of who contributed what , accountability by design. ✅ Less hallucination, more signal Smaller, specialized models trained on clean, domainspecific data are far less prone to hallucinations. Clear boundaries = higher reliability. ✅ Efficiency is the real scalability Deploying one massive model for everything? Expensive, slow, and unsustainable. Specialized AI is leaner, faster, and far more costeffective. In short, OpenLedger isn't just following a trend. They're laying the rails for the infrastructure layer of verifiable, domainaware AI. And in a space flooded with blackbox models and hype, that clarity hits different.

Crypto Sinan

13,636 Aufrufe • vor 11 Monaten

Mark my words, Nebius will be the first Trillion dollar Neo-cloud company and here is why (Save this). Roman Chernin, CEO of Nebius just said on 20VC that Nebius raised prices and demand didn't move. When a company can raise prices and still have more demand than supply, that's the opportunity. Chernin also explained why he is deliberately not charging the maximum. As AI shifts from training, a one time cost to inference, which is the ongoing cost of serving every user and every query, compute pricing becomes the cost structure of the entire AI economy. If Nebius prices customers out, those customers cannot grow, and Nebius cannot grow with them. That is the compounding flywheel built directly into the revenue model. The numbers are already confirming it. Q1 2026 revenue came in at $399 million, up 684% year over year. The AI cloud segment grew 840% and represented 98% of total revenue. Adjusted EBITDA flipped positive to $129.5 million. And Nebius signed a long-term agreement with Meta worth up to $27 billion over five years, a hyperscaler outsourcing its own AI compute stack to a neocloud, which tells you that even companies with $50 billion capex budgets cannot build fast enough. Goldman Sachs says the consensus is underestimating 2027 hyperscaler capex by $500 billion. Every dollar hyperscalers cannot provision themselves flows to neoclouds like Nebius. As that gap widens, Nebius captures the overflow with 3 gigawatts of contracted power already secured and a CEO who just told you raising prices did not dent demand. Our subscribers are already up massively on Nebius and come join Milk Road Pro for our full breakdown, how to size Nebius against the broader neocloud opportunity, and our full AI thesis. Link below!

Milk Road AI

15,677 Aufrufe • vor 1 Monat

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 Aufrufe • vor 1 Monat

Nebius will be a trillion dollar company (Save this). The neocloud market, purpose-built AI cloud infrastructure, separate from legacy hyperscalers generated roughly $25 billion in revenue in 2025, up 223% year over year. Synergy Research projects it will approach $400 billion by 2031, compounding at 58% annually one of the fastest sustained growth rates ever recorded for an infrastructure category of this scale. The CEO's explanation for why they win is worth understanding in detail. GPU compute is scarce and that part everyone knows but Nebius is not simply renting GPUs by the hour and marking them up, which is what most neocloud imitators do. They have built their own physical capacity for inference, optimized the full technology stack from the software layer all the way down to the rack hardware and recently acquired a company called Agen specifically to push inference latency even lower and throughput even higher. The CEO frames the core problem directly that in 2026, every product you build is powered by tokens, AI intelligence and while you can get those tokens from OpenAI or Anthropic via a simple API call, the moment you want to run open source models, specialized vertical models, or anything other than the two dominant frontier labs, you run into a wall. You can download the weights from Hugging Face and assemble the pieces. But getting those workloads to run at scale, at the economics you need, with the reliability your product requires, is an extraordinarily complex engineering challenge that most companies cannot staff or afford to solve in-house. That is the problem Nebius is solving, and that is why their inference product called Token Factory exists. The financial results are among the most dramatic growth numbers reported by any public company this year. In Q1 2026, Nebius posted $399 million in revenue, a 684% increase from the same quarter a year earlier. In the span of twelve months, the company swung from a $104 million net loss to $621 million in net income. Cash from operations went from negative $184 million to positive $2.26 billion in the same period meaning this is not growth funded by burning investor capital, it is growth that is now generating its own fuel. For the full year 2026, Nebius is guiding for an annualized revenue run rate of $7 billion to $9 billion, with pipeline creation tracking to surpass $4 billion. The contracted backlog sits at $49 billion, anchored by a $27 billion agreement with Meta, a deal worth up to $19.4 billion with Microsoft, and a public endorsement from Jensen Huang at NVIDIA's GTC conference in 2026. The current market cap is approximately $56 billion. A company with $7 to $9 billion in annualized revenue, growing at 684%, turning cash-flow positive, sitting on $49 billion in contracted backlog, operating in a market compounding at 58% annually toward $400 billion, that company has a credible path to 20x from its current valuation if execution holds. That is the trillion dollar case, and it does not require any heroic assumptions and it requires Nebius to keep doing what it is already demonstrably doing. Milk Road Pro called this one early. Our analysts added Nebius to the portfolio when it was still flying under the radar, and we are sitting on a massive gain on that position right now. If you want to see what else we are building conviction on before the rest of the market catches up, come join us at Milk Road Pro using the link below!

Milk Road AI

28,622 Aufrufe • vor 2 Monaten

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,354 Aufrufe • vor 22 Tagen