Загрузка видео...

Не удалось загрузить видео

На главную

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI Baseten Philip Kiely and ali explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out...

148,203 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 3

Фото профиля John
John1 месяц назад

speculative decoding was the one that surprised me most in practice. the draft model rejection rate changes a lot based on prompt style, so batching similar prompt types together before running speculative decode gave a much bigger throughput lift than just tuning the draft model itself

Фото профиля Aayaan Naqvi
Aayaan Naqvi1 месяц назад

@baseten @philipkiely @waterloo_intern woahhhhh @waterloo_intern sick

Фото профиля Jong Hyun Park
Jong Hyun Park1 месяц назад

@baseten @philipkiely @waterloo_intern Is a durable moat possible in inference tech on a multi year timescale?

Похожие видео

Interview with Nebius Co-Founder Roman Chernin Please like & share this video so that all $NBIS investors on X will see it! :) If you prefer watching on YouTube: Timestamps: 00:00 - Why AI Infrastructure Is So Hard to Understand 00:24 - Market Fragmentation and What Actually Differentiates Providers 01:30 - Consolidation, Segmentation, and the Future AI Cloud Landscape 02:56 - What Analysts and VCs Still Get Wrong About AI Infrastructure 05:34 - Nebius Cloud: Product Readiness and Customer Proof Points 07:42 - Why Inference Workloads Are Exploding 09:11 - Training vs. Inference: How AI Models Actually Reach Production 10:10 - Why Inference Market Share May Concentrate Around a Few Winners 12:36 - Customer Use Cases: Coding, Enterprise AI, and Real-World Adoption 14:01 - Why Integrated Training and Inference Matter Strategically 16:01 - Building Scalable AI Infrastructure With High Utilization 18:24 - Token Factory: Inference as a Managed Service 20:24 - Revolut Case Study: AI-Driven Product Enhancements 22:56 - Token Factory Performance Optimization and Competitive Advantage 25:07 - Scale, Capacity, and Efficiency as Growth Drivers 28:36 - Why Inference Capacity Could Become the Next Major Bottleneck 30:10 - How Nebius Benchmarks Performance Across Providers 33:14 - The Future Size and Shape of the Inference Market 36:38 - Value-Based Pricing: Moving Beyond Cost per GPU Hour 40:55 - How Nebius Wins Deals: Quality, Performance, and Customer Experience 44:53 - Autonomous AI Platforms and the Rise of Agent-Based Models 47:28 - Tavily, Agentic Applications, and the Next Layer of the AI Stack 50:45 - Strategic Trade-Offs: Scaling, Product Roadmap, and Customer Relevance 55:40 - Final Thoughts: Adapting to the Next Shift in AI Workloads Nebius Roman Chernin

Daniel Koss

204,656 просмотров • 4 месяцев назад

Tokenization -- turning text into a sequence of integers -- is a key part of generative AI, and most API providers charge per million tokens. How does tokenization work? Learn the details of tokenization and RAG optimization in Retrieval Optimization: From Tokenization to Vector Quantization, created in collaboration with Qdrant and taught by its Developer Relations Lead, Kacper Łukawski. This course focuses on Retrieval augmented generation (RAG), which has two steps: First, a retriever finds relevant information; then, the generator uses what’s retrieved as context to produce a response. You’ll learn to optimize the first step (the retriever) by understanding how tokenization works and how it impacts the relevance of your search. In addition, you will also learn to measure and improve retrieval quality, speed, and memory. In detail, you’ll: - Learn about the internal workings of the embedding models and how your text turns into vectors. - Understand how several tokenizers, such as Byte-Pair Encoding, WordPiece, Unigram, and SentencePiece work. - Explore common challenges with tokenizers, such as unknown tokens, domain-specific identifiers, and numerical values, that can negatively affect your vector search. - Understand how to measure the quality of your search across relevance, ranking, and score-related metrics. - Understand how the main parameters in "HNSW", a graph-based algorithm, affect the relevance and speed of vector search, and how to tune its parameters. - Experiment with the three major quantization methods – product, scalar, and binary – and learn how they impact memory requirements, search quality, and speed. By the end of this course, you’ll have a solid understanding of how tokenization functions and how to optimize vector search in your RAG systems. Please sign up here!

Andrew Ng

146,313 просмотров • 2 лет назад

The first company I ever joined was Rubrik, Arvind Jain's previous company. It's where I learned how to build systems and engineering teams. Years later, when we started Composio, Glean became one of our first customers, one of the first to believe in what we were building. So sitting down with Arvind felt like closing a loop. In this episode: •How Glean brought transformers to enterprise search before "semantic search" had a name •Glean's partner-first strategy and why they chose Composio for actions •Why Arvind has never worried about competing with OpenAI or Anthropic •GLM 5.2 as the open-source inflection: 90%+ of enterprise AI tasks, majority of inference within 12–18 months •Measuring real ROI: the US telco that cut case-resolution time by 48% •Why agents built on MCP alone act like day-one employees and how context graphs make them tenured CHAPTERS: (00:00) – Rubrik days: how Arvind and Karan met (01:05) – Glean's origin story: transformers before "generative AI" existed (03:04) – What Glean is today: superset of ChatGPT and Claude (05:39) – From embedding search to agents: was it a pivot? (11:39) – The Composio partnership and why actions are hard (14:43) – Why competition doesn't matter yet (20:11) – The open-source inflection point (GLM 5.2) (24:41) – Contracts and model risk: who absorbs the volatility? (26:57) – Enterprise ROI: cost, tokens, and what to measure (39:57) – Measuring value team by team (44:11) – Day-one employees vs. tenured agents (46:09) – Closing thoughts

Karan Vaidya

62,488 просмотров • 1 месяц назад

$AMD | Inference's estimated to be 80%+ in 2027👑🆕 Dr. Su told everyone the world will need a lot more CPUs and Inference will dominate most of compute long term from 2022. Nobody believed her, but I did along with other high conviction investors. 2026 is the first time Inference surpassed training at 65%+ vs 33-35% for training. Early 2026 infrastructure spend: ~55% inference With Agentic AI, 2027 Inference Infrastructure spend is projected to be 80%+ Training still grows in absolute terms (bigger clusters, more experiments). Inference grows faster because usage, agents, and reasoning traces multiply token volume continuously. That’s why hardware and data center design are shifting toward latency, utilization, and cost per token rather than peak training FLOPS. What "token efficient" actually means when AI labs produces better/smarter models? When models get more token efficient, the expensive part of an agent the GPU “thinking” step gets shorter and cheaper. The rest of the loop does not. The agent still has to parse the answer, pick a tool, run code in a sandbox, query a database, open a browser, apply guardrails, and feed the result back. Those steps live on CPUs where AMD has the best CPU in the world. So a more efficient model does not shrink the agent; it shrinks the model’s share of the agent. Wall clock time and cost tilt toward orchestration and sandboxes, which is why you provision more CPU racks even as tokens per decision fall. Cheaper thinking also unlocks more doing. Teams stop designing one shot answers and start adding retries, parallel branches, sub-agents, and24/7 digital workers, Jevons paradox for agents. Each extra loop is another isolated environment, and unlike GPU batching, sandbox demand scales almost linearly with concurrency: fifty candidate patches means fifty containers, not one fatter GPU job. Token saving tricks often push even more work onto that layer. The result is a fleet that thinks less per step and acts more often, so the volume product becomes CPU/sandbox capacity, not just accelerators. Agentic workloads are a big part of why inference is pulling ahead. I’ll pull the latest numbers on token multipliers and how that shows up in 2026 compute mix.Yes. The inference flip is mostly agents + reasoning, not more people chatting. Chat was one prompt in, a few hundred tokens out. Agents turn a single user request into a loop: plan, tool call, read the result, think, retry, hand off to a sub-agent. That is why Gartner’s 2026 range is 5–30× more tokens per task than a chatbot, with coding and research agents often landing higher. On OpenRouter, agentic workloads 14x’d in six months, passed human usage in February 2026, and by early August were ~5× human tokens (~7.3T/day agentic vs ~1.4T human). A typical agentic request there used 15× the tokens of a human query. One important thing to understand, unit cost per token is still falling, but tokens per useful outcome rose faster. An “agentic seat” can burn 50–100× the tokens of a chat subscription. That is impressive growth for inference. It is also why KV cache, speculative decoding, quantization, and inference specific silicon suddenly matter more than another giant training cluster. Not Financial Advice! DYOR!

Mike

16,350 просмотров • 16 дней назад