Загрузка видео...

Не удалось загрузить видео

На главную

$SNDK CEO David Goeckeler says inference is a memory-bound problem and HBF could break through some bottlenecks, so it's an enormous opportunity as inference scales, which they'll talk more at their analyst day next week. "Look, there's a ton of innovation going on right now in inference memory architectures,...

11,234 просмотров • 6 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Etched is deploying two new technologies in chip design: low-voltage inference and cluster-scale memory. CEO Gavin Uberti says they'll make their chips much more power-efficient and way, way faster than today's leading GPUs. He breaks it down: "We looked at a lot of early research directions, and we realized the key things that models need are way more compute and way faster memory." "If you think about inference, there are two key parts: prefill and decode. For prefill, it's a compute-bound problem. You need to have more FLOPS, more operations per second on each of your chips." "On our GPU, the bottleneck's actually thermals. You can't really run a GPU at more than around 50% of what it could theoretically do, or it'll melt." "So we're using a new technology today called low-voltage inference to try to solve this problem. You bring the voltage of the chip down dramatically, which allows us to have way, way better efficiency in terms of how much power is drawn per unit of math, and thus fit way way more flops onto the chip..." "For decode, it's all about bandwidth. Not just bandwidth on a chip, but bandwidth across your cluster. That's why we have this technology we call cluster-scale memory. It reduces the amount of time it takes to communicate from one chip to another dramatically." "As a result we can go use all of our HBM, HBM bandwidth, SRAM, SRAM bandwidth, and our scale-up domain as a single coherent pool. And that means if you're a user, you can go get much faster tokens per second, while still keeping your costs low."

TBPN

20,404 просмотров • 1 месяц назад

Chamath: Two terms you need to pay attention to in AI are Prefill and Decode “There's two terms that I think you're going to hear a ton about over these next few years.” “The first term is prefill, and the next is decode.” “What prefill and decode are, are two very distinct ways of how models think, and how a model goes through the process of answering a question that you ask it.” “And so when you send a prompt to AI, what happens is that the model processes it. This is called the reading phase or prefill.” “It reads your entire prompt all at once. And then it does a bunch of math, calculates all these relationships between all the words, and it stores them in temporary memory.” “The problem is that this is really compute bound. So it requires massive brute force. And Nvidia GPUs crush here.” “And their architecture is designed for massive parallel processing, which makes them really amazing at digesting these long prompts.” “So the problem just gets bigger and bigger, Nvidia just completely dominates.” “But the next phase though, this critical phase, the decode phase, is the writing phase, right?” “So the model starts to generate a response, you ask it a question and its response, one token at a time.” “And then to pick the next token to pick the next word, it has to look back at everything it has said already so that it doesn't hallucinate.” “The problem is that this is incredibly memory bandwidth constrained.” “And in our architecture, a long time ago, we made these design decisions from day one.” “And so what we did was we took a very different architectural approach, we took a very conservative process technology. We weren't pushing the boundaries of physics.” “And we used a lot of what's called SRAM. So memory on the chip so that we could do this decode thing as well or better than everybody else.” “And so now when you put these two things together, I just think it's going to create a huge acceleration in the ability for this entire infrastructure layer to get much cheaper and much more valuable, which I suspect then it'll have a lot more developer pull, you'll get a lot more applications being built, billions and billions of more people using it.”

The All-In Podcast

567,546 просмотров • 7 месяцев назад

ELON MUSK: We believe the AI5 chip will be roughly comparable performance to an NVIDIA Blackwell, and at much less than 10% of the cost Transcription: I'm super hardcore on chips right now as you may be able to tell. I have chips on the brain. I dream about chips, Literally! Because in order to have a functional robot, you have to have a great AI chip. And it needs to be an inexpensive chip and it needs to be very power efficient So we think we believe the AI5 chip will be probably about a third of the power of say something like a Blackwell, an NVIDIA Blackwell, which is a great chip, for roughly comparable performance. And much less than 10% of the cost. This is a chip that is very much optimized for the Tesla AI software stack. So it's not meant to be a general purpose chip, it's meant to be an amazing chip for the Tesla AI software And I mean a couple of things that I think make... like how is Tesla able to achieve such an improvement? I think it is because we are specialized. We're not trying to... you know, NVIDIA has to serve the superset of all past and future customers. So all of their requirements, all of the software that they've written has to work, which is a very difficult problem. Whereas we just need to make it work for our software. And so we're able to simplify the chip dramatically And then we also, I think we're unique in this, but like we have an integer-based system. And integer operations are fundamentally more efficient than floating point operations. So we can do floating point, but the vast majority of our inference is done in integer. Which is, if you're familiar with sort of logic gates, the simplicity of integer... it's integer is much more power efficient, much more silicon efficient, but you have to, you actually have to train for integer inference, which everyone else is training for floating point. That's kind of like a niche technical detail, but it's actually very important. So, yeah, this is going to be a great chip So this chip will be made in basically in four places: TSMC Taiwan, Samsung Korea, TSMC Arizona, and TSMC Texas. And we already know what improvements to make for AI6. So I'm hopeful that we can within less than a year of AI5 starting production, we can actually transition in the same fab to AI6 and double all of the performance metrics

X Freeze

305,109 просмотров • 9 месяцев назад

Building nuclear reactors that are easy to replicate is a SpaceX style problem: "What we're trying to do here is build a reactor which is easy to replicate. It's a little different from mass manufacturing. Different from a Tesla style problem, more like a SpaceX style problem. You have a complicated vehicle that you need to get really good at building and deploying in a repeatable fashion. When I first started the company, we didn't have a size in mind for the first reactor. It was very explicit. We told the team we don't know how big the reactor is going to be. We don't know how powerful it's going to be. We didn't know those numbers until a year to 18 months into the company. We told ourselves we are going to discover the power level through the manufacturing process. Our instinct was as long as that number turns out to be somewhere above 15 megawatts, we should be pretty good for mass production. If you're under 15 megawatts, it's pretty hard to scale the right way. You just end up doing so many different pieces of operations that it becomes more complicated. But our feeling is above the 15 megawatt break point, you have something that can really scale. And we think this reactor ends up somewhere around 25 megawatts. 25 megawatts being the sort of scale factor where if you want a gigawatt, you just build 40 of them. The next challenge for us as we turn this on and turn the next one on is how do we get to the place where we're turning on one reactor every day and then multiple reactors every day. That's how we're going to climb into the gigawatts."

Ti Morse

58,568 просмотров • 1 месяц назад

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 просмотров • 1 месяц назад

Dylan Patel of SemiAnalysis says a worse GPU with better storage and memory now beats the best chip without them, so buying the newest GPU alone no longer wins inference. So, an AMD GPU with more memory can outperform Nvidia in some cases. "So what we have is we have over $80 million of compute, GPUs from Nvidia, AMD, TPUs from Google, Trainium from Amazon, and we run this benchmark constantly on the newest inference engine, newest drivers, newest PyTorch version, whatever it is." "Every day it runs on an automated CI, and we run it on all the latest Chinese models, from GLM, Zhipu, Moonshot, Kimi, Alibaba, all these models we run." "Initially, when we were benchmarking the difference between these chips and different engines, different schemes for parallelism, we were just running it fixed context length." "But now with Agent X, we've analyzed over $5 million worth of Claude Code traces. This is real production traffic that people have donated to us as well as internally generated. Now we know what the actual agent workload looks like." "And then as we implement that and run those benchmarks, it turns out yes, the chip you're using is very important, but now even more important is how are you handling this memory offload?" "And so while an Nvidia GPU is faster than an AMD GPU in most cases, because AMD GPUs have more memory, they actually end up outperforming in some cases." "Or you can have a worse GPU, but a much better storage solution, and now you can outperform what the best GPU can do without those solutions. So just buying the newest and latest GPU alone doesn't get you the best inference economics." "Actually, you need to layer in all these other innovations including storage and memory." [ Who's the top player on your chart? ] "That really is a difficult multivariable problem. And generally that means you need to have, yes, you need to have the best GPU, a GB300, but you also need to have the best storage solutions. And so I won't spoil who's the best right here, but I will say that storage solutions matter a lot and memory solutions matter a lot, as does your front-end networking. That matters a lot."

Fireside Alpha

175,490 просмотров • 16 дней назад

‼️JUST NOW — powerful brief remarks by the leaders of Ukraine, the UK, France, and Germany in London: IN SHORT: Merz: “The coming days could be decisive for all of us. The destiny of this country is the destiny of Europe. Nobody should doubt our support for Ukraine. And this is what we are firmly standing behind. I'm skeptical about some of the details which we are seeing in the documents coming from the U.S. side.” Macron: "I think we have a lot of cards in our hands." Zelenskyy: "Our team came back, and I think that they will brief us about the last talks with Americans and after talks of Americans with Russians." Starmer: "If there's to be a ceasefire, it needs to be a just and lasting ceasefire. Matters about Ukraine are for Ukraine." FULL STATEMENTS: UK Prime Minister Keir Starmer: "But the principles remain, the principles that we've operated on for a very, very long time, which is that we stand with Ukraine, and if there's to be a ceasefire, it needs to be a just and lasting ceasefire. And that's why it's so important, that we repeatedly set out the principle forward that matters about Ukraine are for Ukraine. And we stand here to support you in the conflict and support you in the negotiations and make sure that this is a just and lasting settlement if we can get that far. You're very welcome for our discussions". President of Ukraine Volodymyr Zelenskyy: "Thank you very much, Kir. Thank you for organizing this meeting. Thank you, friends, Emmanuel and Friedrich, and thank you for our joint meeting. I think that it is very important now to organize such a meeting and discuss very sensitive issues regarding these talks that we had, in the United States and between us. Our team came back, and I think that they will brief us about the last talks with Americans and after talks of Americans with Russians. I think there are lot of what we have to discuss and speak about. So there are a lot of things which are very important for today. I think, unity between Europe and Ukraine, and also unity between Europe, Ukraine, and the U.S. There are some things which we can't manage without Americans, things which we can't manage without Europe, and that's why we need to make some important decisions". President of France Emmanuel Macron: "And thank you Volodymyr for being with us. You are always welcome. In France, I think we all support Ukraine, and we all support peace. And peace negotiations to have sustainable, robust peace. And I think we have a lot of cards in our hands. The financing and the quality of equipments and training programs to Ukraine, the fact that Ukraine is resisting in this war, and the fact that the Russian economy is starting to suffer, especially after our latest sanctions and the US sanctions. And now I think the main issue is the convergence between our common positions, Europeans and Ukrainians, and the U.S. to finalize these, peace negotiations and re-engage in a new phase in the best possible conditions for Ukraine, for the Europeans, and for our collective security". Chancellor of Germany Friedrich Merz: "I highly appreciate that we are having the opportunity to see you, Volodymyr, together with Emmanuel, and talking about the upcoming days, because this could be a decisive time for all of us. We are trying to continue our support for Ukraine, as you know. On the other hand, we are seeing these talks and negotiations in Moscow and in the U.S. I'm looking forward to hear from you what the outcome of these talks might be. And we are still and remain strongly behind Ukraine and giving support to your country, because we all know that the destiny of this country is the destiny of Europe. So that's the reason why we are here, trying to figure out what we can do. And nobody should doubt on our support for Ukraine. And this is what we are firmly standing behind. And the outcome is open. I'm skeptical about some of the details which we are seeing in the documents coming from the U.S. side, but we have to talk about that. That's why we are here".

Kateryna Lisunova

373,089 просмотров • 8 месяцев назад

Incredible last few minutes of the Rick Rubin / Adam Neumann interview. I got chills listening to this. --- Adam: We'll connected to this week's portion. We all have something that we're slaves to. We all have an addiction to anger, to ego, to power, whatever it is, idol worship that we have. We all have an Egypt. We all need to leave that Egypt, come out of it. When we come out of that Egypt, we all face a Red Sea. When we see that Red Sea, we all need to have the courage to walk into it so far that we drown. Only when we're willing to drown will the Red Sea be split and we'll go to our next level. We then go to our next level. We see the Red Sea split. Suddenly we're like, oh, everything is okay. And we forget. The moment we forget, we step into the desert. We all have to walk through our version of desert. That desert is between us, us, and God. No one else. The only decision we get to have is how long we're going to journey through that desert. Are we going to journey for 14 days like the children of Israel were supposed to? Or are we going to journey for 40 years because we're not learning our lesson? Our choice is not, are we going to get into the desert or not? It's how long will we be in. And once we go out of that desert, we all have our version of the promised land. And if we have the courage to look inside, see our own correction, our own limitation, look at it clearly now and say, I see you, and I'm going to get over it. I'm going to learn. I'm going to grow. I'm going to understand that life is not about falling down. It's about how I'm going to get up. I'm going to fall 1,000 times so I can get up 1,001 times. And in that last time when I get up and I'm actually ready to go to my next level, I'm going to wish for you, for me, and for everybody who are listening — please, may we have the courage to go to our next level because we know that the promised land is waiting right after it. And may we not be afraid when the whole world is saying this or that about us. Who cares? It's about what you think, what your wife thinks, what your kids think, and what the universe thinks. Rick: Amen. Adam: Amen. Thank you. Thank you, Hashem. Thank you.

Jared Zoneraich

153,441 просмотров • 4 месяцев назад