Загрузка видео...

Не удалось загрузить видео

На главную

Great to be back at another stellar Modular developer conference: #ModCon2026! Thanks to tim for leading a deeply engaging discussion on AI infrastructure, the economics of training and inference, and the looming massive impact of open models. Takeaway: Efficient open models can achieve 90% of frontier performance at 10%...

11,661 просмотров • 23 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

learned a lot from this conversation with Simon Mo and Matt Bornstein. biggest takeaways for me: -there are a lot of reasons why we should like open-weight models. a lot of these arguments stop at handwavy things like "what if the labs stop releasing frontier models to the public" or "it's lower cost." but simon's position as lead maintainer of vLLM and CEO of Inferact give him authority to talk about some of the other, more interesting and concrete reasons to pay attention to open-weight models, namely that they allow end-users to calibrate latency / other performance metrics with way more customizability than what any of the frontier closed-source labs offer (and without the fear that your job might be met with a refusal at some random point where you're deep in a 2 hour job) -re: the above point...for this reason, a lot of US companies (inferact included!) choose to use open-weight models over their closed-source alternatives. this also isn't limited to internal workloads / research - on a recent a16z podcast the team at Decagon spoke about how something like 90% of their customer service ai agents run on open-weight models that they've fine-tuned. -we should really appreciate how many companies/teams came out researchers fascinated by the wave of very small open-weight models that were being distilled from e.g. gpt-3.5 and earlier models in 2022/2023 (prior to the release of chatGPT!). these small models motivated the development of pagedattention, which then led to vlmm/inferact (at other layers of the stack with similar origin stories, you can look at teams like openrouter or ollama). in other words, we have open-weight models to thank for a bunch of the orchestration infra we now rely on. i think yet another, indirect, way we can point to open-source/weight infra pushing the frontier forward. anyway, a lot more in this convo, it was a lot of fun!

Elena

12,917 просмотров • 1 месяц назад

Chamath is making one of the most important business arguments of 2026. Half of large US companies right now cannot generate returns that exceed their cost of capital, which has normalized back to its long run average of 8 to 11%. Another one in seven companies globally is stuck generating persistent returns between 1 and 5% and most businesses don't have room for error and in this environment walks every frontier AI lab saying the same thing, give us your data, your workflows, your processes and our model will make everything better. And companies by the millions said yes. What they didn't fully account for is what happens on the other side of that door. Every time an employee runs a query through a frontier model API, the prompt goes through external servers, workflows, customer data, pricing logic, internal processes, all of it transmitted through a third party. As Alex Karp said companies are spending on tokens while handing over the exact proprietary advantages that make their business worth owning. Microsoft blocked internal use of Anthropic's Claude Fable 5 but over its 30-day data retention policy and the largest software company in the world decided a frontier model's data handling was too risky for its own employees. A US government action revoked access to another frontier model for foreign nationals overnight. Now here's where the cost math becomes impossible to ignore. Deutsche Bank calculated a roughly 65x cost gap between frontier models like Claude Fable 5 at ~$3.25 per task and open-source alternatives at ~$0.05. For 90% of everyday enterprise tasks, performance is comparable. Open-weight models now match closed frontier systems on core agent tasks at roughly one-tenth the cost, a high-volume deployment that costs $250/day on Claude runs at $12/day on an open-source equivalent. Chamath Palihapitiya tested this directly by running a standard enterprise code migration task through an orchestration layer wrapping an open-source model came in 16.4x cheaper than using a frontier model directly.

Milk Road AI

281,771 просмотров • 2 месяцев назад

Quick chat with dylan ツ (Dylan Bristot, GTM @ $NBIS). Also on YouTube (link in first comment) for those who prefer to watch/listen there. Timestamps 00:00 – Dylan's role at Nebius and Nebius Token Factory 01:48 – Dylan's investing philosophy and portfolio approach 05:07 – How working in AI infrastructure influences his investing 08:29 – Training vs. inference and why inference demand could explode 13:22 – Enterprise AI adoption: from POCs to production 18:02 – Open-source vs. closed/frontier models 24:44 – The economics of open vs. closed AI models 29:27 – Where the next AI infrastructure bottlenecks could emerge 31:22 – Dylan's AI Bottlenecks project and approach to stock selection 34:06 – Closing thoughts Key Insights (AI Summary, so you don't have to copy paste and prompt for exactly that ;D) “I seem to like areas where the demand really looks kind of secular, but the supply is genuinely hard to create.” → Implication: The most attractive AI trades may sit in physical bottlenecks where supply cannot quickly respond to demand. “The bottleneck is who has the pricing power and kind of what might get commoditized and where the concentrate might move next.” → Implication: Value capture across the AI stack will keep shifting as individual layers become scarce or commoditized. “Training creates the intelligence and then the inference actually monetizes and distributes.” → Implication: Training and inference are complementary, rather than one ultimately replacing the other. “One user action can become dozens or hundreds of model calls, tools calls, and like verification steps, retries.” → Implication: Agentic AI can drive token consumption far faster than user growth alone would suggest. “The best infra for making any model and the best infra for serving a billion interactions are not necessarily the same.” → Implication: Training and inference could increasingly require different hardware and infrastructure architectures. “The Frontier Labs might be incentivized to run more and more of the inference of these models for internal research instead of providing it to external people.” → Implication: The most capable models and their compute could increasingly be used internally to accelerate frontier research rather than monetized externally. “Enterprise AI adoption is actually much further along than a lot of people kind of think. But probably less mature than the headlines suggest.” → Implication: Enterprise demand is real, but deployment maturity still has significant room to improve. “The POC problem might be solved for a lot of companies, but the production problem isn’t yet.” → Implication: The enterprise bottleneck is shifting from proving AI works to deploying it reliably, securely and economically at scale. “They feel like it’s time for them to actually not only integrate AI, but build some sort of moat out of the AI.” → Implication: Enterprises increasingly want proprietary AI systems built around their own data rather than simply consuming generic models. “The more autonomous the software becomes, the more infra discipline you need underneath it.” → Implication: Agents increase the importance of inference cost, reliability and infrastructure optimization. “Maybe I have fifteen different versions of very different LLMs, fine tuned on fifteen different kinds of tasks that I’m operating across my business, instead of having a one model fits all.” → Implication: Enterprise AI could evolve toward many specialized models rather than one frontier model handling every workload. “I don’t necessarily think it’s open versus closed. That might be the wrong framing.” → Implication: Open and closed models can coexist because they optimize for different customer needs. “Historically the problem was that that control came with a massive operational tax.” → Implication: Better inference infrastructure can make open models materially more competitive by removing the complexity traditionally associated with running them. “I don’t think open needs to beat the best closed model on every single benchmark. It just basically needs to be good enough for the workload of the given customer while offering a much better combination of control, cost, and deployment flexibility.” → Implication: For production AI, workload-specific economics may matter more than having the absolute smartest model. “Maybe actually the bulk of tokens generated in the future might come from open models.” → Implication: Frontier intelligence could remain dominated by closed labs even while open models capture most production inference volume. “I could really imagine frontier intelligence being really concentrated while most of the production inference becomes super fragmented.” → Implication: AI could consolidate at the intelligence layer while fragmenting heavily at the inference layer across models, GPUs, providers and regions. “I don’t think that necessarily means the margins of open source will be much worse than the ones of closed source.” → Implication: Optimization can potentially make open-model inference highly profitable despite lower pricing. “I think now we’re probably in the middle of phase two... everything feeding the accelerator.” → Implication: The AI trade is broadening beyond GPUs toward networking, packaging, data centers, electrical equipment and power. “It’s no longer about the megawatts, about energized megawatts.” → Implication: Available power on paper matters less than how quickly that power can actually be delivered to operating AI infrastructure. “It’s increasingly about utilisation and conversion now and like how efficiently you convert expensive infra into actual useful AI work.” → Implication: Infrastructure efficiency and utilization become increasingly important as the absolute amount of deployed AI infrastructure grows. “The market tends to really notice demand before it notices what demand breaks.” → Implication: Second-order bottlenecks may offer some of the most interesting opportunities in the next phase of the AI buildout. “The interesting question now is which part of the mine breaks next?” → Implication: Finding the next constraint in the AI supply chain may matter more than simply identifying continued AI demand.

Daniel Koss

49,970 просмотров • 11 дней назад