正在加载视频...

视频加载失败

Jane Street engineers, in an internal tech talk on training performance: "GPUs are really slow at doing one floating point operation. It's around a microsecond. CPUs are 1,000 times faster than GPUs." Not a podcast. A talk given inside the firm that trades billions a day, to a room...

12,502 次观看 • 18 天前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

September 2009. Jensen Huang walks onto a small stage at the Fairmont hotel in San Jose. About 1,500 people are in the room. He runs a company that makes chips for video games. He spends the next 8 minutes doing math on a whiteboard, explaining why the future of computing won't come from making CPUs faster. He calls it "CEO math" and apologizes in advance to every computer science professor in the audience. Then he lays out an argument that almost nobody took seriously at the time: the way to make computers dramatically faster is to pair a regular CPU with hundreds of tiny parallel processors, the kind that already exist inside graphics cards. One CPU for the sequential stuff. Hundreds of GPU cores for everything else. He calls it "heterogeneous computing." He shows the math. A workload that can be split into many pieces at once gets up to 200x faster on this combined system. A workload that has to run one step at a time loses nothing. "The most important thing in creating a new architecture," he says, "is to make sure it does no harm." This was the first GPU Technology Conference. NVIDIA had launched a software platform called CUDA three years earlier, in 2006, to let developers write programs that run on graphics cards instead of just regular processors. Almost nobody cared. GPUs were for rendering Call of Duty, not for scientific computing. The academic world was polite but skeptical. The enterprise world ignored it entirely. By this point, Huang had been making this argument for years. NVIDIA was a $7 billion company. It competed with AMD and Intel for market share in the graphics market. That was the whole business. Jensen kept saying the GPU wasn't just a gaming chip; it was a computing platform. He kept saying parallel processing would reshape every industry from medicine to finance to physics simulations. People kept nodding, then doing nothing. Then deep learning happened. Around 2012, AI researchers discovered that training a neural network, which means teaching a computer to recognize patterns by running the same calculation millions of times across huge datasets, was exactly the kind of workload Jensen had been describing. GPUs can train AI models 10 to 50 times faster than CPUs. The architecture he outlined in this 2009 talk, with one CPU handling step-by-step tasks while hundreds of GPU cores crunch through massive amounts of parallel data, is now the literal blueprint for every AI data center on earth. ChatGPT runs on NVIDIA GPUs. Claude runs on NVIDIA GPUs. Gemini, Llama, Midjourney, nearly every major AI model you've heard of was trained on NVIDIA hardware using CUDA, the software platform Jensen built for a market that didn't exist yet. NVIDIA was worth about $7 billion when Jensen gave this talk. It is worth over $4.4 trillion today. That's a 600x increase. Jensen Huang, who founded the company at a Denny's in 1993 with two friends, now has a net worth of over $160 billion. He made Forbes' list of the 10 richest people for the first time this year. GTC 2026 is currently ongoing. 17,000 people are packing a hockey arena to watch the same guy explain what comes next. In 2009, 1,500 people showed up at a hotel ballroom, most of them for gaming graphics.

Anish Moonka

413,645 次观看 • 4 个月前

The creator of High Bandwidth Memory (HBM) put a number on the AI build that should stop every infra investor cold. A cluster of a million GPUs runs at roughly 10-20% utilization (Save this). Kim Jung-ho spent thirty years building what feeds the GPU, and his claim is that the GPU is barely working. Here is what is actually happening. Every time a model generates output, the data has to be read out of memory, computed, and written back. The read and the write swallow almost the entire cycle. While that data moves, the GPU does nothing. It sits there, fully powered, fully paid for, waiting. By Kim's estimate the memory is doing only about 30 percent of the work it needs to do. The processor idles the rest. So a million installed GPUs run at 10 to 20 percent. You are not compute constrained. You are memory constrained, and the expensive part is standing around. Adding more GPUs does not fix this. It gives you more processors starving for the same data. Here is the part that decides the next decade. Memory can grow. When a cell cannot shrink any further, you stack it into a high-rise, layer on layer. A GPU cannot be stacked. It runs too hot and needs a cooler bolted to its back, so the one move that rescues memory is closed to the processor. The thing that can keep stacking compounds. The thing that cannot plateaus. The marginal dollar in an AI build now buys more by fixing the memory path than by bolting on another idle GPU. Which is why the companies that control memory bandwidth and supply are not suppliers to the AI trade. They are the AI trade.

Fireside Alpha

38,370 次观看 • 1 个月前

Dario Amodei just told software engineers exactly how long they have. Six to twelve months. Amodei: “I have engineers within Anthropic who say I don’t write any code anymore. I just let the model write the code, I edit it, I do the things around it.” The people building the most powerful AI in history have already stopped writing code. That is not a forecast. That is the current working condition inside the lab closest to the frontier. Amodei: “We might be six to 12 months away from when the model is doing most, maybe all, of what SWEs do end-to-end.” The tech industry spent a decade making software engineers its highest-paid, most protected class. That era has a last day now. When a model can execute an entire software build end-to-end, the ability to write syntax stops being a skill. It becomes a credential for a job that no longer exists. Amodei: “And then it’s a question of how fast does that loop close.” That is the sentence everyone skipped. The code was never the hard part. The hard part was everything around it. The model just learned everything around it. Writing the code is already nearly gone. Testing is next. Deployment is next. When all three collapse into a single autonomous execution loop, the machine no longer needs a human in the chain at all. The corporation or sovereign state that closes that loop first does not gain a competitive advantage. It gains a category of speed that biological engineers cannot match, track, or reverse. That is not disruption. That is replacement at a systems level. Amodei is not describing a future disruption. He is describing the current state of his own building. The loop is already closing. The only question is whether you are inside it or outside it when it seals.

Dustin

318,457 次观看 • 5 个月前

Every Fortune 500 executive is buying AI subscriptions and calling it a strategy. Palantir CEO Alex Karp has a word for that. Karp: “The general approach of just buying models is going to be essentially self-pleasuring for an enterprise at the cost of the enterprise.” Karp: “You buy some large language model, you party with it basically, and the next day you have a hangover.” The entire corporate world is mispricing the AI transition. They are renting intelligence with no foundation to run it on. A raw model floating in a vacuum hallucinates over your unstructured data, generates the illusion of work, and executes nothing. The party ends. The hangover begins. Nothing changed. Karp identified exactly where the value actually goes. Karp: “All the value in the market is going to go to chips and what we call ontology.” Not the models. Not the subscriptions. Not the chatbot interfaces layered on top of them. The ontology. The precise digital architecture of how an organization actually operates. Its security permissions. Its supply chain physics. It’s operational logic. Karp: “The ontology will allow you to take a large language model and use it, refine it, and then impose it on your enterprise in the logic of your enterprise, in the security model of your enterprise.” When you bind a frontier model to the strict underlying logic of a specific enterprise, something fundamental shifts. It stops generating text. It starts generating action. Karp: “We’re using it on the battlefield, we’re using it to compress margins. We’re making engineers better engineers. We’re making people who are not engineers into engineers using our ontology and a large language model.” The traditional engineering bottleneck does not slow down. It disappears. Karp: “We are sitting on the only thing that actually creates quantifiable, transformational value.” The companies renting models are paying for the feeling of transformation. The companies building ontologies are executing the actual thing. One of them will define the next decade. The other will wake up in 2030 wondering where their market share went. Exactly like a hangover.

Dustin

210,533 次观看 • 5 个月前