Loading video...

Video Failed to Load

Go Home

From Palantir and Two Sigma to turning mechanistic interpretability into a production-grade platform, Mark Bissell (Member of Technical Staff) and Myra Deng (Head of Product) are building Goodfire on the belief that “understanding what’s inside the model” is the path to safer, more controllable AI and Goodfire just raised...

19,389 views • 5 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

New Course: Post-training of LLMs Learn to post-train and customize an LLM in this short course, taught by Banghua Zhu, Assistant Professor at the University of Washington University of Washington, and co-founder of @NexusflowX. Training an LLM to follow instructions or answer questions has two key stages: pre-training and post-training. In pre-training, it learns to predict the next word or token from large amounts of unlabeled text. In post-training, it learns useful behaviors such as following instructions, tool use, and reasoning. Post-training transforms a general-purpose token predictor—trained on trillions of unlabeled text tokens—into an assistant that follows instructions and performs specific tasks. Because it is much cheaper than pre-training, it is practical for many more teams to incorporate post-training methods into their workflows than pre-training. In this course, you’ll learn three common post-training methods—Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Online Reinforcement Learning (RL)—and how to use each one effectively. With SFT, you train the model on pairs of input and ideal output responses. With DPO, you provide both a preferred (chosen) and a less preferred (rejected) response and train the model to favor the preferred output. With RL, the model generates an output, receives a reward score based on human or automated feedback, and updates the model to improve performance. You’ll learn the basic concepts, common use cases, and principles for curating high-quality data for effective training. Through hands-on labs, you’ll download a pre-trained model from Hugging Face and post-train it using SFT, DPO, and RL to see how each technique shapes model behavior. In detail, you’ll: - Understand what post-training is, when to use it, and how it differs from pre-training. - Build an SFT pipeline to turn a base model into an instruct model. - Explore how DPO reshapes behavior by minimizing contrastive loss—penalizing poor responses and reinforcing preferred ones. - Implement a DPO pipeline to change the identity of a chat assistant. - Learn online RL methods such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO), and how to design reward functions. - Train a model with GRPO to improve its math capabilities using a verifiable reward. Post-training is one of the most rapidly developing areas of LLM training. Whether you’re building a high-accuracy context-specific assistant, fine-tuning a model's tone, or improving task-specific accuracy, this course will give you experience with the most important techniques shaping how LLMs are post-trained today. Please sign up here:

Andrew Ng

125,146 views • 1 year ago

Ep. 23 | Free The Money | Only 4 Companies Control 85% of U.S. Beef — Here’s Why That Matters In this episode of Free the Money, I sit down with entrepreneur & investor Camden, to explore how decentralization is reshaping everything from our food system to workforce development and healthcare. Camden breaks down why the U.S. food system is so fragile, highlighting that just four companies control roughly 85% of the nation’s beef supply, and why that level of centralization creates risks for both health and resilience. We discuss how decentralizing agriculture through local ranching, direct-to-consumer models, and even on-chain ownership of livestock can restore transparency, quality, and trust in what we eat. From there, we dive into the health implications of today’s food system, including inflammation, gut health, and simple, accessible ways people can start improving their well-being through diet, sunlight, and daily habits. We also explore Camden’s work with Skillmaker, a company using AI, extended reality, and real-time data to radically transform skilled labor training, cutting training timelines from years down to weeks. This opens the door to solving massive labor shortages across industries like automotive, construction, and infrastructure, while aligning incentives to bring more people back into the trades. On the healthcare side, Camden shares insights into his work with Opiaid, a data-driven platform designed to improve addiction treatment outcomes using wearable technology and real-time patient data that offers a more effective and personalized approach to recovery. Finally, we touch on what’s happening behind the scenes with regulatory reform, the MAHA movement, and the broader shift toward decentralization across systems that have historically been tightly controlled. Sign up for ITrustCapital with this link for $100 funding bonus. See why people are opening a tax-advantaged Crypto, Gold & Silver IRA for their future: Intro Camden’s Entrepreneurial Journey Why Our Food System Is Broken (And How Regenerative Farming Fixes It) The Real Bottleneck in Food Supply: Why Processing Controls Everything On-Chain Agriculture: Owning Cattle Like Crypto Assets Heal Your Gut: Simple Daily Habits That Actually Work Sunscreen, Sunlight & Toxins: What You’re Not Being Told Skillmaker Explained: Turning 2 Years of Training Into 25 Days The Skilled Labor Crisis: Why Incentives Are Completely Broken Using AI to Fight Addiction: Inside the Opiate Crisis Solution Inside MAHA & The FDA: What’s Really Changing Behind the Scenes “Who’s Shaking the Jar?” — A Powerful Lesson from St. Thomas

Bri Teresi

23,206 views • 3 months ago

Perplexity CEO Aravind Srinivas on the biggest threat to the data center industry: It's not competition. It's not regulation. It's decentralisation. "The biggest threat to a data center is if the intelligence can be packed locally on a chip that's running on the device and then there's no need to inference all of it on like one centralized data center." He outlines how this could work in practice. Personalisation doesn't necessarily require on-device model training. Retrieval augmented generation, tool calls, and local data can already tailor AI to individual users. But the real unlock? Test time training. Aravind Srinivas describes a future where AI lives on your device, watches how you work and gradually automates your repetitive tasks. "Imagine we crack test time training where the AI watches tasks you repeatedly do on your local system, adapts to you over time and starts automating a lot of the things you do." The key insight: in this model, the intelligence belongs to you. It's your data, your device, your personalised AI brain. And if that future arrives, the economics of centralised infrastructure start to collapse. "That really disrupts the whole data center industry. It doesn't make sense to spend all this money, 500 billion, 5 trillion, whatever on building all the centralized data centers across the world that do a lot of the intelligence workloads for people." The companies spending trillions on centralised infrastructure may want to rethink where intelligence actually needs to live.

Big Brain AI

90,102 views • 5 months ago

Elon just dropped a MAJOR nugget on how Tesla is going to be training Optimus to do real world tasks. They are building an Optimus Academy, which is a large scale, dedicated real-world training facility to accelerate the development of Optimus. The Academy will deploy thousands of Optimus units, potentially 10,000 to 30,000 robots, in a controlled realistic environment where they perform self-play, experiment with tasks, iterate on behaviors, and continuously generate training data through trial and error. The Tesla bots will also run millions of simulations in Tesla’s high-fidelity physics-accurate engine, allowing Optimus to close the “sim-to-real gap” by using these real-world observations to refine and validate the simulations! “You’re actually highlighting an important limitation and difference from cars. We’ll soon have 10 million cars on the road. It’s hard to duplicate that massive training flywheel. For the robot, what we’re going to need to do is build a lot of robots and put them in kind of an Optimus Academy so they can do self-play in reality. We’re actually building that out. We can have at least 10,000 Optimus robots, maybe 20-30,000, that are doing self-play and testing different tasks. Tesla has quite a good reality generator, a physics-accurate reality generator, that we made for the cars. We’ll do the same thing for the robots. We actually have done that for the robots. So you have a few tens of thousands of humanoid robots doing different tasks. You can do millions of simulated robots in the simulated world. You use the tens of thousands of robots in the real world to close the simulation to reality gap. Close the sim-to-real gap.”

Teslaconomics

42,563 views • 5 months ago

New PNAS paper. Historical GDP per capita data is scarce, but data on the places of birth, death, and occupations of famous individuals is abundant. In this paper we estimate the historical GDP per capita of hundreds of regions in Europe and North America using a machine learning model that leveraged data on about 500k famous biographies. Our estimates more-or-less quadruple the availability of historical GDP per capita estimates for the last 700 years. So why use biographies to augment historical GDP per capita data? Biographical data contains information about people who might have contributed directly to economic growth, like James Watt, or that were attracted to wealthy places looking for patrons, like Michelangelo. So we--mainly Philipp (Philipp Koch)--used this data to construct hundreds of features describing each European region. Then, we trained a machine learning model to find the features that explained most of the variance in a cross-validation test, where we split regions multiple times into a training set and a test set. On average, the model explained about 90% of the variance in GDP per capita of the regions it had not seen during training. But we wanted to go further, and Philipp really went to town by looking at different ways to validate our estimates. We found our estimates correlate positively with historical measures of wellbeing, church building activity, urbanization, and body height. We also used these measures to reproduce the basic Atlantic trade result of Acemoglu, Johnson, and Robison and to explore the economic consequences of the famous Lisbon earthquake of 1755. But what I personally loved most about this project, other than working with Philipp Koch and V, is that it shows that we can use machine learning methods not only to explore the future, but the past. There is a bright and growing future in the use of machine learning for economic history. Hope you enjoy the paper and the data. You can find links to the paper and a data exploration tool in the first comment.

César A. Hidalgo

54,332 views • 1 year ago

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

507,952 views • 3 months ago

Lightspeed's Bucky Moore says the real opportunity in the AI app layer is in large industries far enough afield from where the model providers are today — and where the context engineering to get customer data into the model is extremely nuanced and messy. "I think this is kind of the elephant in the room right now — whether post-training open-source models combined with the unique user feedback you get from being an application provider is defensible enough." "That is going to be an inevitable challenge for any of these industries that hit a maturation point of AI adoption, like legal and software engineering have." "But on the other hand, there are some industries where they're very large, they're far enough afield from where the model providers are today — and probably will continue to be — and the context engineering to actually get the customer data into the model is just so messy. It requires going across different business functions, it requires a lot of hands-on forward-deployed engineering." "Those are the kind of companies that we get really excited about. Because I think being really good at that is not only defensible, but it also allows you to generate a feedback loop with your customers, where you hear a lot of their secrets. And those secrets allow you to feed that back into how you make your product better at the expense of anyone else playing in the space. Because if you're serving the customer, they're only serving you those secrets." "I think Palantir is a good example of this in the pre-AI era, and I think we're going to see many companies ascend in that same way."

TBPN

46,746 views • 4 months ago