Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

an infinitely large model trained on infinite data still cannot go below 1.69 GPT-4 burned three months on twenty five thousand GPUs inching toward it plot every model ever trained, error against compute on a log scale, and none of them gets below one straight line. runs ten trillion...

60,149 görüntüleme • 8 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Intelligence was the one thing that never scaled. We scaled everything else. Steel. Energy. Compute. The one resource that built all of it never left the skull. Musk: “People thought defeating Go was either never or 20 years away.” Twelve months later it was over. Musk: “Now that same AlphaGo system can defeat the top 50 players simultaneously with 0% chance of them winning. And that’s one year later.” Fifty lifetimes of mastery against a system that does not know it is playing a game. Zero percent chance. That was not a competition. That was a preview. Musk: “The degrees of freedom to which artificial intelligence is able to apply itself are really increasing by 10 orders of magnitude a year.” Ten billion times. Every twelve months. No brain alive can visualize that number. By design. Every hard problem that ever defeated us did it for the same reason. Not complexity. Scarcity. The only mind capable of solving it was biological and there was never enough of it. Cancer. Fusion. Climate. The physics we cannot even see yet. Not waiting on more data. Waiting on something that can think at a scale biology never allowed. That just arrived. Most people hear this and reduce it to a question about their paycheck. They are watching the single largest expansion of capability in the history of life on this planet and worrying about a job title. For ten thousand years intelligence had one speed. One brain. One lifetime. Every civilization on earth throttled by the same biological ceiling. That ceiling just shattered. We are the only species that ever hit its own limit and built what breaks through it. That is not an ending. That is the point of everything we ever built.

Dustin

24,426 görüntüleme • 2 ay önce

Jensen Huang just told Silicon Valley it’s fighting on the wrong floor. Every boardroom in tech is locked on the same question. Which model wins. OpenAI or xAI. GPT or Claude. Grok or Gemini. Trillions moving on that bet alone. Huang zoomed out and showed them the whole building. Huang: “AI is actually essentially a five-layer cake.” Energy at the bottom. Chips above it. Cloud above that. Models next. Applications on top. Five layers. One war. Everyone crowded onto the fourth floor. Huang: “This is where most people think AI is.” He was pointing at the model layer. Every pitch deck. Every valuation. Every founder story. All packed onto one floor. One floor below the finish line. Three above the foundation. The middle of the building. Huang: “At the bottom is energy.” Not data. Not parameters. Not talent. Power. You cannot out-code the grid. You cannot train a frontier model with a press release. The smartest model on Earth still needs a dumb turbine spinning somewhere. The smartest engineers alive are building on top of someone else’s silicon, inside someone else’s cloud, powered by someone else’s electricity. They own nothing beneath them. Huang: “This layer on top ultimately is where economic benefit will happen.” Healthcare. Finance. Manufacturing. The only floors where AI actually meets money. Every dollar of real value lives at the top. Every physical constraint that decides who gets to play lives at the bottom. The model sits in between. Squeezed from above and below and owning neither end. Silicon Valley is burning hundreds of billions to build plumbing for somebody else’s economy. The basement decides if it runs. The penthouse decides if it pays. The companies building models think they are building the future. Huang just told them they are the middle layer in someone else’s cake.

Dustin

536,261 görüntüleme • 4 ay önce

Dario Amodei just revealed that the AI training bottleneck everyone is worried about doesn’t exist anymore. The industry spent years obsessed with scraping the open web. More data. More text. More human output to feed the models. Amodei: “I don’t think data is quite the most central thing anymore.” The shift is fundamental. Amodei: “Static data is becoming less important. A lot of the data we use today is RL environments that we train on. Dynamic data that the model creates itself.” Not scraped. Not licensed. Not written by humans. Generated by the model through pure trial and error. When you train on complex math or agentic coding, you don’t feed it a textbook. You give it an environment. The model experiments. Fails. Adjusts. Tries again. Amodei: “You’re getting some math problems and the model experiments with trying the math problems.” It generates its own experience. Millions of iterations. Each one building on the last. No human required. This destroys the entire narrative around AI hitting a data wall. You cannot throttle a competitor by locking down copyright. Cannot slow the race by putting up a paywall. When a model learns through its own synthetic experience, the open web becomes irrelevant. The only true bottleneck left is compute. And this is where the geopolitical stakes become impossible to overstate. The nation that wins the compute race doesn’t just build smarter models. It builds models that generate their own intelligence, compounding on themselves, iterating past every limit human knowledge ever imposed. We are no longer training AI on the past. We are letting it simulate the future. The machine has stopped reading the dictionary. It’s doing the math itself now.

Dustin

149,633 görüntüleme • 5 ay önce

Elon Musk's primary residence was a car factory for three years. He slept on the floor during shift change so the people walking in would see him. Musk: "I actually know the people on the line because I worked on the line, and I walked the line, and I slept in the factory, and I worked beside them. So I'm no stranger to them." There are maybe five people at that altitude who could say that sentence without lying. Most executives know their org chart. They know the name of everyone whose approval affects their pay. They do not know the person who installed the battery pack at 3 AM. Musk: "There is no lords and peasants. Everyone eats at the same table, everyone parks in the same parking lot." Any company can put that on a wall. The building is where you find out if they meant it. Musk: "At GM there's a special elevator only for senior executives. We have no such thing at Tesla." A separate elevator. Someone sat in a meeting and decided executives should not have to ride with the people building the product. That elevator is a machine for making human beings invisible. Most people will work their entire lives and never once be seen by the person whose name is on the building. They will give that building forty years. Their back. Their knees. The bedtimes they missed. And they will die as a line item in a spreadsheet nobody reads twice. Nobody calls that a tragedy because nothing dramatic ever happens. It just takes an entire life to finish. Not that people are underpaid. That they are unwitnessed. Every company writes its real values into the floor plan and hopes nobody reads them. The parking spots. The elevators. The floors people are allowed to stand on. People read that long before they read the mission statement. Musk: "We give everyone stock options. Many people who just work on the line who didn't even know what stocks were, we've made them millionaires." Someone took that job because rent was due on the first. No exposure to equity. No path to anything past the next paycheck. Now they own a piece of what they built with their hands. Not a bonus. Not a plaque. Ownership. Their kids will inherit something other than exhaustion. Musk: "I'm incredibly appreciative of those who built cars, and they know it." And they know it. Gratitude that stays in the CEO's head is worthless. Gratitude announced in a press release is worse. The only gratitude that counts is the kind that reaches the floor. You cannot memo it into existence. You cannot hire a consultant to install it. A person will sell you their hours. They will only give you their life if you were standing there when it was hard. None of that is management. It is just what we are. For almost all of human history it was impossible to work unseen. The group was thirty people. Everyone knew who carried weight and who did not. Then we built organizations large enough that a person could pour everything they had into one and vanish. We have had ten thousand years to adjust to that and we never have. We were never built to follow titles. We were built to follow the person who stayed. This is why some companies get everything a person has, and the rest send out engagement surveys wondering where it went. Tesla's people are building something alongside someone who was there when it was ugly, who knows what it cost them, and who never pretended he was above the floor. And almost nobody at the top is willing to pay what it costs.

Dustin

22,212 görüntüleme • 21 gün önce

How good is GPT-4-Vision at extracting text from images? I wanted to find the limit - but I found weirdness instead Most surprising: GPT-4V performance varies depending on the *structure* of text it sees Let me explain A set of images with progressively more text was presented to GPT-4-Vision. GPT-4V was asked what text it saw in the image. The response from the model was compared against the image’s original text and scored for similarity. The model was tested with 4 types of text: essay, random words, random tokens, and random characters. Findings: * Performance degrades - Yes, the models are good at basic OCR, but as you get more text and words then performance drops (this is expected) * Type of context matters - You should expect different recall on your texts based on your context types * Hallucination Errors - I thought that the model would make errors of omission (it wouldn’t return all the words). But instead the model mostly made hallucination errors - it replaced words with made up words. * Evals Matter - This test in isolation doesn’t mean that your data will have the same results, but it should motivate you to create eval tests for your data and anticipate errors which are hard to spot Notes: * Next step would be to add additional image types like tables or PDFs * GPT-4V would routinely get stuck in repeat-token-loops when trying to extract random tokens * GPT-4V would refuse to answer most random character images

Greg Kamradt

49,111 görüntüleme • 2 yıl önce

The creator of High Bandwidth Memory (HBM) put a number on the AI build that should stop every infra investor cold. A cluster of a million GPUs runs at roughly 10-20% utilization (Save this). Kim Jung-ho spent thirty years building what feeds the GPU, and his claim is that the GPU is barely working. Here is what is actually happening. Every time a model generates output, the data has to be read out of memory, computed, and written back. The read and the write swallow almost the entire cycle. While that data moves, the GPU does nothing. It sits there, fully powered, fully paid for, waiting. By Kim's estimate the memory is doing only about 30 percent of the work it needs to do. The processor idles the rest. So a million installed GPUs run at 10 to 20 percent. You are not compute constrained. You are memory constrained, and the expensive part is standing around. Adding more GPUs does not fix this. It gives you more processors starving for the same data. Here is the part that decides the next decade. Memory can grow. When a cell cannot shrink any further, you stack it into a high-rise, layer on layer. A GPU cannot be stacked. It runs too hot and needs a cooler bolted to its back, so the one move that rescues memory is closed to the processor. The thing that can keep stacking compounds. The thing that cannot plateaus. The marginal dollar in an AI build now buys more by fixing the memory path than by bolting on another idle GPU. Which is why the companies that control memory bandwidth and supply are not suppliers to the AI trade. They are the AI trade.

Fireside Alpha

38,370 görüntüleme • 1 ay önce

Demis Hassabis just explained why the real AI bottleneck has nothing to do with training runs. Most people picture the AI arms race as who can build the biggest model. GPT-4 or Gemini Ultra style training runs, a few hundred million in compute, fired once or twice a year. The constraint sits somewhere else. Every time a researcher has a new algorithmic idea, a new architecture, a new training technique, they can't just test it on a laptop. They have to run it at the scale where it would actually be deployed, because ideas that look promising at small scale fall apart completely when you put them into a real system. Every research hypothesis burns significant compute before a single line of production code gets written. At a lab like DeepMind, hundreds of researchers are running hundreds of ideas simultaneously. The demand for experimental compute is continuous. It never stops. Now layer the hardware reality on top. GPU lead times are currently 36 to 52 weeks for data center hardware. Global AI data centers are already drawing 29.6 gigawatts, equivalent to the peak power demand of the entire state of New York, and they still can't meet demand. Companies willing to pay any price can't just buy more compute. They wait in line. The speed of scientific discovery in AI is now gated by hardware availability. The next breakthrough is sitting in a researcher's head right now. Whether it gets validated fast enough to matter depends entirely on whether the compute is there when they need it. The AI race gets won by whoever can run the most experiments per month.

Aakash Gupta

32,150 görüntüleme • 4 ay önce

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,358 görüntüleme • 4 ay önce

Andrej Karpathy just made one of the most interesting arguments about AI model design that most people are completely missing. His take is that frontier AI models are not too big because the technology is complex and too big because the training data is garbage. When you or I think of the internet, we picture Wall Street Journal articles, Wikipedia entries, serious writing. That is not what a pretraining dataset looks like. When researchers at frontier labs look at random documents from the actual training corpus, it is stock ticker symbols, broken HTML, spam, gibberish. One estimate puts Llama 3's information compression at just 0.07 bits per token meaning the model has only a hazy recollection of most of what it trained on. So we build trillion parameter models not because we need a trillion parameter brain but because we need a trillion-parameter compression engine to squeeze some intelligence out of a firehose of noise. Most of those parameters are doing memory work, not cognitive work. Karpathy's prediction is separate the two entirely. Build a cognitive core, a model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization and pair it with external memory that it can query when it needs facts. He thinks a cognitive core trained on high-quality data could hit genuine intelligence at around one billion parameters. For reference, today's flagship models run between 200 billion and 1.8 trillion parameters with most of that weight dedicated to remembering the internet's slop. The trend is already moving his direction. GPT-4o operates at roughly 200 billion parameters and outperforms the original 1.8 trillion-parameter GPT-4. Inference costs for GPT-3.5-level performance dropped 280-fold between 2022 and 2024 driven almost entirely by smaller, cleaner, better-architected models. The real bottleneck in AI right now is not compute but rather data quality.

Milk Road AI

200,517 görüntüleme • 4 ay önce

Elon Musk just scored human civilization on the only scale the universe keeps. Zero. Musk: “We’re practically nowhere on the Kardashev scale. Not registering.” The Kardashev scale measures one thing. How much energy a civilization captures from its star. Type 1 harnesses its planet. Type 2 harnesses its star. Humanity is not Type 1. We are so far below it the math rounds to nothing. A trillion is a million times a million. That is the gap. Every reactor ever built. Every barrel of oil ever burned. Every grid on every continent. Rounding error. Three centuries of industrial progress. Not a flicker on the only scale that tracks whether a civilization is real. Oil is just old sunlight. We killed each other for the crumbs. The sun throws 3.8 × 10²⁶ watts into empty space. Earth intercepts one two-billionth of that. We harness a fraction of even that. The rest hits nothing. Warms nothing. Powers nothing. The offer has stood for four billion years. Musk: “Even one millionth of what the sun outputs. Extremely kickass civilization.” One millionth would make us gods by current standards. We are not at one millionth. We are fighting over condensation on a fire hose. And that is the most hopeful fact about our species. We mistook the floor for the ceiling. Musk: “Launch satellites to orbit Earth and capture solar power. Avoids the need to build massive power plants on Earth.” This is not a policy position. It is a coordinate change. Institutions don’t scale. Physics does. Bureaucracies were designed to manage scarcity. They have no operating manual for abundance. So they fight it. Because abundance makes gatekeepers unemployed. Everything we call history happened in the dark. Every atom in your body was forged inside a star. You are not reaching for something foreign. You are the sun’s output reaching back for the sun. It poured out an empire for four billion years. No one showed up to claim it. Until now.

Dustin

33,665 görüntüleme • 2 ay önce