Загрузка видео...

Не удалось загрузить видео

На главную

Most AI projects don't fail because models aren't good enough. They fail because inference economics don't scale. During #CXOSpice, Roman Chernin Nebius broke it down: ✅Closed models kill margins ✅Open models unlock control, if you can run them well ✅DataData only becomes a moat when inference is sustainable One...

13,662 просмотров • 8 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

My full interview with Roman Chernin, Co-founder & Chief Business Officer of Nebius Nebius just signed a $17 billion deal with Microsoft and a $3 billion deal with Meta. These are two of the biggest tech companies on Earth - and they're coming to Nebius for AI infrastructure. But the company does so much more than that: - Full-stack AI cloud (data centers → software → managed services) - Nebius Token Factory Nebius Token Factory for managed inference & post-training - Partners include Shopify, Higgsfield, Jetbrains My 5 takeaways with Roman Chernin below - thanks for the conversation! Full interview also available on YT ⬇️ Timestamps: 00:00 - Introduction 01:47 - Overview of Nebius $NBIS 03:23 - Roman's background and role at Nebius 04:10 - The $17B Microsoft & $3B Meta deals 05:29 - Why hyperscalers trust Nebius over building in-house 07:06 - Nebius's full-stack approach: Data centers to managed services 08:39 - Customer segmentation: Hyperscalers & Frontier AI Labs vs AI Startups vs Enterprises 14:41 - Token Factory explained: Managed inference & post-training 17:41 - When to switch from closed-source to open-source models 20:59 - Real results: Process achieves 26% cost reduction 25:05 - Why developers love building on Nebius 28:40 - AI agents and the future of infrastructure interaction 32:22 - NVIDIA's Groq acquisition: What it means for inference 35:38 - Why AI adoption will surprise us

Elliot Garreffa

29,888 просмотров • 7 месяцев назад

Interview with Nebius Co-Founder Roman Chernin Please like & share this video so that all $NBIS investors on X will see it! :) If you prefer watching on YouTube: Timestamps: 00:00 - Why AI Infrastructure Is So Hard to Understand 00:24 - Market Fragmentation and What Actually Differentiates Providers 01:30 - Consolidation, Segmentation, and the Future AI Cloud Landscape 02:56 - What Analysts and VCs Still Get Wrong About AI Infrastructure 05:34 - Nebius Cloud: Product Readiness and Customer Proof Points 07:42 - Why Inference Workloads Are Exploding 09:11 - Training vs. Inference: How AI Models Actually Reach Production 10:10 - Why Inference Market Share May Concentrate Around a Few Winners 12:36 - Customer Use Cases: Coding, Enterprise AI, and Real-World Adoption 14:01 - Why Integrated Training and Inference Matter Strategically 16:01 - Building Scalable AI Infrastructure With High Utilization 18:24 - Token Factory: Inference as a Managed Service 20:24 - Revolut Case Study: AI-Driven Product Enhancements 22:56 - Token Factory Performance Optimization and Competitive Advantage 25:07 - Scale, Capacity, and Efficiency as Growth Drivers 28:36 - Why Inference Capacity Could Become the Next Major Bottleneck 30:10 - How Nebius Benchmarks Performance Across Providers 33:14 - The Future Size and Shape of the Inference Market 36:38 - Value-Based Pricing: Moving Beyond Cost per GPU Hour 40:55 - How Nebius Wins Deals: Quality, Performance, and Customer Experience 44:53 - Autonomous AI Platforms and the Rise of Agent-Based Models 47:28 - Tavily, Agentic Applications, and the Next Layer of the AI Stack 50:45 - Strategic Trade-Offs: Scaling, Product Roadmap, and Customer Relevance 55:40 - Final Thoughts: Adapting to the Next Shift in AI Workloads Nebius Roman Chernin

Daniel Koss

204,842 просмотров • 5 месяцев назад

Quick chat with dylan ツ (Dylan Bristot, GTM @ $NBIS). Also on YouTube (link in first comment) for those who prefer to watch/listen there. Timestamps 00:00 – Dylan's role at Nebius and Nebius Token Factory 01:48 – Dylan's investing philosophy and portfolio approach 05:07 – How working in AI infrastructure influences his investing 08:29 – Training vs. inference and why inference demand could explode 13:22 – Enterprise AI adoption: from POCs to production 18:02 – Open-source vs. closed/frontier models 24:44 – The economics of open vs. closed AI models 29:27 – Where the next AI infrastructure bottlenecks could emerge 31:22 – Dylan's AI Bottlenecks project and approach to stock selection 34:06 – Closing thoughts Key Insights (AI Summary, so you don't have to copy paste and prompt for exactly that ;D) “I seem to like areas where the demand really looks kind of secular, but the supply is genuinely hard to create.” → Implication: The most attractive AI trades may sit in physical bottlenecks where supply cannot quickly respond to demand. “The bottleneck is who has the pricing power and kind of what might get commoditized and where the concentrate might move next.” → Implication: Value capture across the AI stack will keep shifting as individual layers become scarce or commoditized. “Training creates the intelligence and then the inference actually monetizes and distributes.” → Implication: Training and inference are complementary, rather than one ultimately replacing the other. “One user action can become dozens or hundreds of model calls, tools calls, and like verification steps, retries.” → Implication: Agentic AI can drive token consumption far faster than user growth alone would suggest. “The best infra for making any model and the best infra for serving a billion interactions are not necessarily the same.” → Implication: Training and inference could increasingly require different hardware and infrastructure architectures. “The Frontier Labs might be incentivized to run more and more of the inference of these models for internal research instead of providing it to external people.” → Implication: The most capable models and their compute could increasingly be used internally to accelerate frontier research rather than monetized externally. “Enterprise AI adoption is actually much further along than a lot of people kind of think. But probably less mature than the headlines suggest.” → Implication: Enterprise demand is real, but deployment maturity still has significant room to improve. “The POC problem might be solved for a lot of companies, but the production problem isn’t yet.” → Implication: The enterprise bottleneck is shifting from proving AI works to deploying it reliably, securely and economically at scale. “They feel like it’s time for them to actually not only integrate AI, but build some sort of moat out of the AI.” → Implication: Enterprises increasingly want proprietary AI systems built around their own data rather than simply consuming generic models. “The more autonomous the software becomes, the more infra discipline you need underneath it.” → Implication: Agents increase the importance of inference cost, reliability and infrastructure optimization. “Maybe I have fifteen different versions of very different LLMs, fine tuned on fifteen different kinds of tasks that I’m operating across my business, instead of having a one model fits all.” → Implication: Enterprise AI could evolve toward many specialized models rather than one frontier model handling every workload. “I don’t necessarily think it’s open versus closed. That might be the wrong framing.” → Implication: Open and closed models can coexist because they optimize for different customer needs. “Historically the problem was that that control came with a massive operational tax.” → Implication: Better inference infrastructure can make open models materially more competitive by removing the complexity traditionally associated with running them. “I don’t think open needs to beat the best closed model on every single benchmark. It just basically needs to be good enough for the workload of the given customer while offering a much better combination of control, cost, and deployment flexibility.” → Implication: For production AI, workload-specific economics may matter more than having the absolute smartest model. “Maybe actually the bulk of tokens generated in the future might come from open models.” → Implication: Frontier intelligence could remain dominated by closed labs even while open models capture most production inference volume. “I could really imagine frontier intelligence being really concentrated while most of the production inference becomes super fragmented.” → Implication: AI could consolidate at the intelligence layer while fragmenting heavily at the inference layer across models, GPUs, providers and regions. “I don’t think that necessarily means the margins of open source will be much worse than the ones of closed source.” → Implication: Optimization can potentially make open-model inference highly profitable despite lower pricing. “I think now we’re probably in the middle of phase two... everything feeding the accelerator.” → Implication: The AI trade is broadening beyond GPUs toward networking, packaging, data centers, electrical equipment and power. “It’s no longer about the megawatts, about energized megawatts.” → Implication: Available power on paper matters less than how quickly that power can actually be delivered to operating AI infrastructure. “It’s increasingly about utilisation and conversion now and like how efficiently you convert expensive infra into actual useful AI work.” → Implication: Infrastructure efficiency and utilization become increasingly important as the absolute amount of deployed AI infrastructure grows. “The market tends to really notice demand before it notices what demand breaks.” → Implication: Second-order bottlenecks may offer some of the most interesting opportunities in the next phase of the AI buildout. “The interesting question now is which part of the mine breaks next?” → Implication: Finding the next constraint in the AI supply chain may matter more than simply identifying continued AI demand.

Daniel Koss

49,970 просмотров • 25 дней назад

Baseten’s CEO on why every company will want to own its AI / Tuhin Srivastava on the rise of AI agents, the demand for inference, and building a new kind of hyperscaler. Lately, I’ve been spending a lot of time thinking about the AI inference market and how big it could get. Agents like Muse, Instinct, Town, and Grok Bot aggressively use browsers, run their own computers, and burn through far more tokens than a traditional chatbot. As more people put them to work (Muse is number two in the App Store), the demand for inference, or the computing needed to run these models, should grow enormously. Tuhin and I discuss the rise of agents using browsers and virtual machines to get things done, and Baseten's recent acquisition to help power that shift. We also talk about the data center backlash, why companies are embracing Chinese open models, his plans for Baseten’s new research lab, and why he thinks inference becomes the only market left after AGI. Timestamps: 00:00 What Is AI Inference? 06:05 Competing With the Cloud Giants 09:24 Why Companies Want to Own Their AI 15:30 Building Baseten Before the AI Boom 23:45 DeepSeek and the Race for Open AI Models 29:32 Baseten’s Growth and Expansion 33:05 AI Agents and the Blaxel Acquisition 38:32 Data Centers and the AI Backlash 43:25 What Happens to Inference After AGI? 45:12 When AI Agents Become Customers Thanks to the show's premier sponsors: Atlassian, Granola, and Mercury.

Alex Heath

25,513 просмотров • 10 дней назад

NEW: Premium Inference 101 The Economics & Infrastructure Behind Running Trillion Parameter Models Rodrigo Liang, CEO & Co-Founder of SambaNova "Inference has arrived. 70-80% of those racks are running inference." "[Inference services] are generating lots of revenue, but not enough margin. In order for them to sustain, they've gotta be more profitable." "With SambaNova, that min quantum is down to 1 rack. Where if you have other service providers, [with] say, a DeepSeek model, now 1.5 trillion parameters, to run that, the min for some of the other providers might be 10-20 racks." SambaNova builds full-stack inference infrastructure. 16 chips to a 10kW air-cooled rack that runs trillion parameter models, where a GPU rack pulls 130kW. They just demonstrated the fastest MiniMax M2.7 inference in the world, as benchmarked by Artificial Analysis. The demo paired one NVIDIA H200 rack for prefill with one SambaRack SN50 for decode. Disaggregated inference: GPUs load the context, RDUs generate the tokens. Now serving JPMorgan, SoftBank, Saudi Aramco & DOE national labs, just valued at $11B on a $1B Series F led by General Atlantic. We Cover: › Why inference will need orders of magnitude more chips than training ever did › The 10kW rack vs the 130kW rack, & why air cooling decides geography › Running a 1T parameter model in one rack at full precision, no quantization › The agent latency problem: 20 agents, 2 seconds each, 40 seconds gone › Revenue per rack, & why inference providers have revenue but no margin › JPMorgan, sovereignty, & the move back to on-prem Filmed at the RAISE Summit in Paris. Thank you to Brex, MongoDB & AssemblyAI for helping make this trip & content series happen. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Rodrigo Liang , Co-Founder & CEO at SambaNova Systems (00:59) SambaNova’s Series F: $1B raise at an $11 billion valuation (03:00) The Inference problem nobody saw coming (04:52) SambaNova's chip evolution (07:19) Running a trillion-parameter model on a single rack (11:00) Do $100 billion data centers actually make sense? (14:14) What "premium inference" really means (18:28) Speed is about to become AI's biggest price tag (20:43) Starlink, edge computing, & AI reaching every corner of the planet (24:27) Working alongside NVIDIA & rival chipmakers (27:49) How customers actually measure inference performance (32:07) The biggest bottlenecks in AI's global land grab (35:12) Justifying the billion-dollar AI valuations (37:53) Why SambaNova refuses to build its own cloud (41:03) The "AI sovereignty" debate (43:48) Data privacy fears are driving the return to on-prem AI (47:55) How to actually get ROI out of AI spend (51:16) The one question every business should be asking about AI (56:09) The mentors & lessons behind a 32-year career in chips (58:02) Unveiling SambaNova's newest chip, the SN50

Molly O’Shea

134,960 просмотров • 2 месяцев назад

This is why Nebius will be a trillion dollar hyperscaler (Save this). Nebius is not building another GPU rental shop but rather building a vertically integrated hyperscaler that owns everything from the physical data center, to the server rack hardware it designs in house, to the software stack, to the inference delivery layer. Nearly every other neocloud is essentially a reseller of someone else's infrastructure but Nebius owns the full stack end to end and that distinction is the entire thesis. Here is why vertical integration is the winning architecture for the inference era. AWS and Azure were architected for general purpose computing and every AI workload they run sits on top of infrastructure that was never designed for it, patched, adapted and optimized after the fact. Nebius was built from day one specifically for AI which means every layer of the stack is purpose built and co optimized. The rack design, the networking topology, the cooling systems and the software that orchestrates it all are engineered together as a single system rather than assembled from parts that were never meant to work together. That architectural difference compounds with every passing quarter as AI workloads grow more complex and the performance gap between purpose built and general purpose infrastructure widens. The software layer is where the real competitive moat lives. Most infrastructure companies think of software as a wrapper around hardware while Nebius thinks of software as the product with hardware as the substrate it controls. The company is building an AI native cloud platform where the software layer handles model serving, inference optimization, fine tuning pipelines and developer tooling as first-class primitives. This matters because inference efficiency is almost entirely a software problem. Two companies running identical GPUs can deliver dramatically different performance and cost per token depending on how intelligently the software schedules, batches and routes inference requests across the cluster. Nebius is also building for a fundamental shift in how AI infrastructure gets consumed. Today, enterprise developers navigate massive cloud service catalogs spinning up clusters, managing configurations and building deep expertise in AWS or GCP-specific tooling. The next generation of builders will simply provision agents to interface with infrastructure directly. Nebius is architecting its software layer for that future , one where the interface between the developer and the compute abstraction layer looks nothing like what AWS built in 2006. The entire available capacity has been sold out every quarter. And that is the best possible validation that what Nebius is building is exactly what the market needs and that the market is willing to commit at a scale that makes the current valuation look like the beginning of a much longer story. Long Nebius and make sure to follow me Melvin for more overlooked AI stocks.

Melvin

34,306 просмотров • 2 месяцев назад