Loading video...

Video Failed to Load

Go Home

Mark Zuckerberg explains the 405B teacher-model flywheel that could make one giant AI the wrong end state "People are gonna wanna do inference directly on the 405 because it's, you know, by our estimates, it's gonna be about 50% cheaper, I think, than GPT-4o to do that directly." "Because...

523,111 views • 28 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,354 views • 29 days ago

Mark Zuckerberg is explaining one of the most misunderstood dynamics in AI and it has direct investment implications (Save this). The concept he's describing is model distillation, and it's one of the most important techniques to emerge in AI over the past year. Here's how it works. You train a massive, enormously expensive model, in Meta's case, Llama 4 Behemoth, a 2 trillion parameter teacher model and then you use that model to teach a much smaller, cheaper model. The smaller model inherits roughly 90 to 95% of the intelligence of the giant while running at 10% of the cost and on a fraction of the compute. Meta already did this with the Llama 4 family and Behemoth serves as the teacher. Llama 4 Scout and Maverick, the publicly released open-source models were distilled from it. Scout runs on a single H100 GPU with a 10 million token context window and outperforms models that cost far more to operate. Maverick, at 17 billion active parameters, rivals DeepSeek V3 in coding at half the parameter count and beats GPT-4o on multimodal benchmarks. Both are completely free for commercial use. What Zuckerberg is pointing at is a structural shift in how AI gets deployed in the real world. Companies aren't taking a frontier model off the shelf and running it as-is but rather taking open-source models, fine-tuning them on their own proprietary data, distilling them into even smaller custom models tailored to their specific use case, and running them on infrastructure they control at a fraction of the cost of a closed frontier API. The investment implication of this is significant and runs in two directions. For Meta specifically, this is a strategic masterstroke. Every company that builds on Llama, fine-tunes it, distills it, or deploys it through their infrastructure is pulling into Meta's orbit while Meta builds the most powerful open teacher model. The ecosystem of companies using it grows and that ecosystem generates commercial activity across Meta's platforms and data services. Meta's AI research benefits from billions of real world deployment signals and it's a flywheel that closed model providers cannot replicate because their strategy requires charging per token, which is now a 65x cost disadvantage against the open-source alternative. For the broader market, distillation changes the economics of inference in a way that has barely been priced in. As intelligence becomes extractable into smaller and cheaper models, the absolute demand for compute doesn't decline but rather it explodes, because now the number of applications that are economically viable expands by orders of magnitude. Every task that was previously too expensive to automate at $3.25 per call becomes viable at $0.05 that means more total token usage, more total GPU utilization, and more demand for the infrastructure companies, the Nebiuses, the GE Vernovas, the Constellation Energies that supply the underlying compute and power.

Milk Road AI

27,908 views • 1 month ago

.Josh Wolfe: Anybody Using DeepSeek App Is 'Absolute Fool' "Anybody using the DeepSeek app is an absolute fool. If you're using DeepSeek on companies like Together Compute, one of Lux's companies, which can get rid of the CCP censorship, then it's probably okay. But remember, the open-source movement is something we deeply believe in. Most great technologists, entrepreneurs, and venture capitalists are on the side of open source. The closed-source models that have consumed tens of billions of dollars are the ones that are really going to be at risk. When you look at Hugging Face, a major repository, or Together Compute, Runway ML, and a lot of Lux's companies, they have been pioneers in open source. Now, why am I not worried about open source, even with the DeepSeek model? As long as you don't have the CCP censorship on it, the models with their open weights allow people to run on their proprietary data. This means companies like pharma or defense companies that have their own siloed, proprietary data—think about Bloomberg with their proprietary longitudinal data, or Meta with their data—are the ones who will have the edge. Even as open source takes hold, these companies will still dominate. I’m not worried about open source being the problem. I’m more concerned about people overfunding closed models with no proprietary source. A lot of capital is going to be burned there, and we’re already seeing that with people worried about OpenAI in some aspects."

Josh Caplan

39,985 views • 1 year ago