正在加载视频...

视频加载失败

HBM creator Kim Jung-ho reveals his lab's tHBM flipped stack publicly for the first time saying Samsung's zHBM with memory stacked on a hot GPU simply melts and is likely not the right fundamental direction "So the idea is, let's just put the HBM on top of the GPU....

96,220 次观看 • 5 天前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

The creator of High Bandwidth Memory (HBM) put a number on the AI build that should stop every infra investor cold. A cluster of a million GPUs runs at roughly 10-20% utilization (Save this). Kim Jung-ho spent thirty years building what feeds the GPU, and his claim is that the GPU is barely working. Here is what is actually happening. Every time a model generates output, the data has to be read out of memory, computed, and written back. The read and the write swallow almost the entire cycle. While that data moves, the GPU does nothing. It sits there, fully powered, fully paid for, waiting. By Kim's estimate the memory is doing only about 30 percent of the work it needs to do. The processor idles the rest. So a million installed GPUs run at 10 to 20 percent. You are not compute constrained. You are memory constrained, and the expensive part is standing around. Adding more GPUs does not fix this. It gives you more processors starving for the same data. Here is the part that decides the next decade. Memory can grow. When a cell cannot shrink any further, you stack it into a high-rise, layer on layer. A GPU cannot be stacked. It runs too hot and needs a cooler bolted to its back, so the one move that rescues memory is closed to the processor. The thing that can keep stacking compounds. The thing that cannot plateaus. The marginal dollar in an AI build now buys more by fixing the memory path than by bolting on another idle GPU. Which is why the companies that control memory bandwidth and supply are not suppliers to the AI trade. They are the AI trade.

Fireside Alpha

38,370 次观看 • 1 个月前

Dylan Patel of SemiAnalysis says a worse GPU with better storage and memory now beats the best chip without them, so buying the newest GPU alone no longer wins inference. So, an AMD GPU with more memory can outperform Nvidia in some cases. "So what we have is we have over $80 million of compute, GPUs from Nvidia, AMD, TPUs from Google, Trainium from Amazon, and we run this benchmark constantly on the newest inference engine, newest drivers, newest PyTorch version, whatever it is." "Every day it runs on an automated CI, and we run it on all the latest Chinese models, from GLM, Zhipu, Moonshot, Kimi, Alibaba, all these models we run." "Initially, when we were benchmarking the difference between these chips and different engines, different schemes for parallelism, we were just running it fixed context length." "But now with Agent X, we've analyzed over $5 million worth of Claude Code traces. This is real production traffic that people have donated to us as well as internally generated. Now we know what the actual agent workload looks like." "And then as we implement that and run those benchmarks, it turns out yes, the chip you're using is very important, but now even more important is how are you handling this memory offload?" "And so while an Nvidia GPU is faster than an AMD GPU in most cases, because AMD GPUs have more memory, they actually end up outperforming in some cases." "Or you can have a worse GPU, but a much better storage solution, and now you can outperform what the best GPU can do without those solutions. So just buying the newest and latest GPU alone doesn't get you the best inference economics." "Actually, you need to layer in all these other innovations including storage and memory." [ Who's the top player on your chart? ] "That really is a difficult multivariable problem. And generally that means you need to have, yes, you need to have the best GPU, a GB300, but you also need to have the best storage solutions. And so I won't spoil who's the best right here, but I will say that storage solutions matter a lot and memory solutions matter a lot, as does your front-end networking. That matters a lot."

Fireside Alpha

175,586 次观看 • 27 天前

Gavin Baker, CIO of Atreides Management made one of the most important and nuanced calls on memory stocks in recent months (Save this). His argument is that based on every memory cycle of the last 25 years, the setup today, prices elevated, sentiment high, supply ramping is textbook time to sell but he adds a critical exception. The one cycle in modern memory history where selling was catastrophically wrong was the mid-1990s, which Baker calls the last true capacity cycle in memory. In that cycle, demand was structurally exploding as the internet era required entirely new computing infrastructure to be built from scratch, and memory had to scale with it in a way that had never happened before. His point is that AI may be that same kind of cycle and not a normal boom bust but a once in a generation capacity buildout where the underlying demand is structural, not cyclical. The reason this argument holds weight is the fundamental shift in what memory is in the AI era. Traditional DRAM was a pure commodity, identical specs, interchangeable suppliers, price determined entirely by supply and demand swings. HBM is the opposite because it is custom engineered to fit a specific customer's chip, co-designed between the memory maker and the GPU designer, with SK Hynix's Vice President literally describing it as shifting from a commodity to a customer-tailored custom business. A single Blackwell Ultra GPU now requires up to 288GB of HBM3E, a 3.6x increase over the H100 and major suppliers like SK Hynix and Micron have already sold out their entire HBM production capacity through the end of the year. Because HBM requires advanced packaging processes like CoWoS that can't be spun up overnight, the bottleneck isn't just wafer capacity but rather runs across the entire manufacturing stack. Bank of America projects the global HBM market grows 58% this year alone to $54.6 billion, and Nomura expects the broader memory sector to nearly double to $445 billion. Long Micron!

Milk Road AI

260,701 次观看 • 1 个月前

Dylan Patel on the importance of memory and storage Two key quotes: "An $NVDA GPU is faster than an $AMD GPU in most cases, but because AMD GPUs have more memory, they can outperform Nvidia in certain workloads." “It is a difficult, multivariable problem. Generally, you need the best GPU, such as a GB300, but you also need the best storage solutions. I will not spoil who comes out on top, but storage solutions matter a lot, memory solutions matter a lot, and frontend networking also matters significantly" Full Quote: “We have over $80 million of compute: GPUs from $NVDA and $AMD, TPUs from Google, and Trainium from Amazon. We constantly run this benchmark using the newest inference engines, drivers, PyTorch versions, and other software. It runs every day through automated CI across the latest Chinese models from GLM, Zhipu, Moonshot, Kimi, Alibaba, and others. Initially, when we were benchmarking the differences between these chips, inference engines, and parallelism schemes, we used fixed context lengths. But with Agent X, we have now analyzed more than $5 million worth of Claude Code traces. This is real production traffic that users have donated to us, combined with internally generated data, so we now understand what an actual agent workload looks like. When we implement those workloads and run the benchmarks, it turns out that the chip you are using is very important, but how you handle memory offload can be even more important. An Nvidia GPU is faster than an AMD GPU in most cases, but because AMD GPUs have more memory, they can outperform Nvidia in certain workloads. Similarly, you can use a less powerful GPU with a much better storage solution and outperform the best GPU when it lacks those solutions. Simply buying the newest GPU does not necessarily give you the best inference economics. You need to layer in other innovations, including storage and memory.” Interviewer: “Who is the top player on your chart? Can you tell us?” Dylan Patel: “It is a difficult, multivariable problem. Generally, you need the best GPU, such as a GB300, but you also need the best storage solutions. I will not spoil who comes out on top, but storage solutions matter a lot, memory solutions matter a lot, and frontend networking also matters significantly.”

Daniel Romero

38,220 次观看 • 1 个月前

Important: Who is Opposing the UAPDA? "It's staff! And it's like, who elected you? You were not elected!" ~Burlison "I cannot believe that I am sponsoring the language that Senator Schumer also is sponsoring." ~Burlison ~ Burlison: "We're trying to get UAP Disclosure Acts put on the National Defense Authorization Act. We are facing difficulties, and I mean, it's so frustrating, because it's not members [of Congress]. I don't, at least, if it is members, they're not coming to me and saying, 'Hey, I'm not allowing your amendment.' "It's staff! And it's like, who elected you? You were not elected!' I mean, for you to be able to deny an elected official whose...my job is to answer to the people, and you're gonna deny me the ability to get this amendment, at least even offered, is unacceptable. "And so, right now we're trying to get that made in order, so that when it comes to the floor, at least we get an opportunity to have a vote on the UAP Disclosure Act! And then you'll know who supports it and who doesn't. "But I've gotta get the Intelligence Committee to approve that. And, so any help, André Carson, would be greatly appreciated. If you could lean on our chairman and others, that would be fantastic, but we've gotta get it made in order. And if we get out of the House, I think that we're halfway there. "And then we'll...I think we've got allies in the Senate, whether it'. And I cannot believe that I am sponsoring the language that Senator Schumer also is sponsoring (laughter and applause). But what's right is right, and truth is truth. And we're all Americans, and that's why this is such a common-sense thing. And the only thing that's gonna stop it is these staffers who are killing things behind the scenes in the cover of darkness. And so, that's what we're facing right now, so your help getting vocal on this is probably the only way we're going to get it done."

Joe Murgia

11,580 次观看 • 1 个月前

Micron is one of the most UNDERVALUED stocks in the entire AI trade right now and everyone should be buying at these prices. (Save this). Jensen laid out the situation in one sentence, the supply chain is lined up, the HBM is lined up with the Grace Blackwell GPUs, the only problem is that demand is much greater than the overall capacity of the world. And Michael Dell said it before Jensen even finished that memory is the single biggest supply constraint in the entire AI buildout right now. Every HBM chip that Micron, SK Hynix, and Samsung produce consumes three times the silicon wafer area of standard DRAM. Nvidia's Rubin GPU requires 288GB of HBM per chip, a 260% increase over the H100 in just two generations. Every major hyperscaler has locked up contracts through 2026, and Micron has said publicly it can only fulfill about two-thirds of medium-term demand for some customers. And it's HBM production is sold out entirely for 2026 and HBM4 is also already sold out. The numbers tell the story, DDR4 spot prices surged roughly 15x in eight months. DRAM contract prices rose 90-95% in a single quarter, TrendForce called it "essentially unprecedented" in the history of the memory market. Micron has rallied roughly 68% year to date in 2026, and yet it still trades at a P/E of 37.6x against an industry average of 75.3x. The shortage does not resolve until new fabs come online, Micron's new factories are not producing until 2027 and 2028 at the earliest, and the memory shortage is forecast to run until at least 2027. Milk Road Pro has been covering the HBM memory trade as a core AI infrastructure thesis before it became a consensus Wall Street call and our Pro members are already up massively in $MU. Come join us at the link in bio/below to see our full portfolio and the names we're watching before the rest of the market catches on.

Milk Road AI

146,863 次观看 • 3 个月前

Catherine Austin Fitts on "breakthrough" tech funded by taxpayers' "missing money" "they've built an awful lot of infrastructure... ships in the sky... stuff going on underground... we can run cars from water... [and] if they can get the price of graphene down... what you can do with materials is unbelievable" This clip of Fitts, a former Assistant Secretary of Housing and Urban Development, investment banker, and founder of the Solari Report (The Solari Report | Catherine Austin Fitts), is taken from a discussion with Versan Aljarrah (Versan | Black Swan Capitalist), the founder of Black Swan Capitalist, posted to YouTube on July 2, 2026. ---------------Partial transcription of clip--------------- "I've tried to trace where the money theoretically could go, sort of as a conceptual matter, but I think they've built an awful lot of infrastructure which is not something we don't see when we walk around— "So, much more impressive ships in the sky, much more stuff going on underground... And especially because we're, as a society, we're using 50- and 100- year-old legacy technology. "And we know, I mean, I just know I'm in the Netherlands because my partners here I first met in 2012 because they led the breakthrough energy conferences in, in, in the Netherlands and the US, as well as the secret space program. "Because a lot of the breakthrough energy is associated with the sort of the space programs. And if you just look at the technology we've had for over 100 years to dramatically reduce the price of energy, we know this technology works. We know we can run cars from water. "We know all of this stuff is feasible. And the question is, when is it going to be allowed and integrated into the regular economy? Your guess is as good as mine, but somebody's got it. "So when the head of Lockheed Skunk Works said in 1996, we now have the technology to take ET home— That tech, I, I'm taking him for his word, I think he was telling the truth, but I don't think we've— It was, I don't know if you saw this Lockheed CEO the other day sometime in the last year on a conference call with investors was talking about the incredible technology they, they've developed, they've just developed that he can't talk about.... "I think there's a lot of technology that's going to come out of the lab over the next 10 to 15 years. And, you know, so the other day they, they built literally a modular nuclear plant that they can put in an airplane and fly to a military base. Right? "You know, you have a lot of practical applications coming. If you just look in material science, if they can get the price of graphene down, I mean, what you can do with materials is just unbelievable. "So I think, when, when Trump or these guys talk about a golden age. I don't think they're being ridiculous. I think they understand the possibilities of this technology and they get all excited about it. "But the problem is there's no way to communicate it in a way that people can really understand because the guys who want to make money from it want to, you know, keep it to themselves. So you have a— there's, there are real conflict issues on, on how this technology comes out and who knows about it. But I think the potential for technology to dramatically improve productivity is enormous."

Sense Receptor

16,646 次观看 • 1 个月前

I got to ask Jensen a question today at CES 2026. Question: What advice would you give to a new robotics founder for them to choose the right application space or the right idea so that they can have the most impact and most differentiation? Jensen's answer (short summary): The real strategic choice is between a horizontal play and a vertical one. Horizontal competition comes from every direction, but focusing on a specific vertical allows you to solve the hardest problems for a specific industry. Whether it’s EMS manufacturing or surgical robotics, that deep domain expertise is very beneficial. Jensen's full answer: Well, first of all, let's take a step back. As you know, NVIDIA, here we were just talking about AI factories. And that AI factory—our contribution, our chips, systems, infrastructure, which is software, and model technology. Is that right? That's kind of the NVIDIA stack. And that's an AI factory. In order to build robotic systems, you really need three different computers. You need the training computer, which I just described. And then you need another computer for doing simulation, because the robot needs to learn how to practice and be evaluated inside a virtual world that's physically precise, so that it doesn't have to do crazy stuff in the physical world while it's still learning. And so we create a virtual world that obeys the laws of physics, and I've demonstrated it several times, and that virtual world is called Omniverse. And so that's a second computer. And that computer is much more like one of our gaming computers, and the GPU that we use for that is RTX Pro. Basically an RTX. And then the third computer is the computer that goes into the robot. It's the robot brain. And that robot computer we call Orin today, and then the next generation is called Thor. And it has its own stack. So Thor has a super-fast inference stack. It runs a safety operating system, like the safety operating system we have in the car, because you want the robot to stay in its guardrails and not do things that it's not confident in doing. And so, you have your stack, and then you have your model. And the model could be fine articulation, manipulation, and locomotion, and each one of those, it could be system one and system two thinking. The technology necessary to build a robot is incredible. And so I just described for you, in order to be a robotics company, you have to have three computers. You have to understand all three stacks, and you have to build this robotic system, not to mention all the electronics and the mechanicals necessary to do it. It's incredibly hard. However, as you know, these pieces of technology independently have been coming together. Isn't that right? Which is really what happens to a new industry, is when there's enabling technology necessary for the industry itself, but it rides on the contributions of the technology advances in other industry that it doesn't have to worry about. And so the humanoid industry is riding on the work of the AI factories we're building for other mainstream stuff and other AI stuff. And our Omniverse was designed for other applications and different other digital twin capabilities. And so all this stuff is now coming together. The question is for robotics, ultimately, it comes down to a couple of different questions. Do you want to be a horizontal company, or do you want to be a vertically domain-specific company? The benefit of a horizontal company, of course, is that you less worry about the application, you more worry about the technology, and if you succeed, your scale can be quite large. However, horizontal plays are incredibly hard. Your competition comes from every single direction. Now, on domain, if you want to be domain-specific, then you're going to have to understand the particular application quite deeply. So maybe it's something to do with EMS manufacturing, assembly of these AI supercomputers. Maybe it's related to building cars in factories. Whatever the reasons are, your domain expertise—could be surgical robots—domain expertise could really be a benefit. My preference usually—and you asked, so I'll offer—my preference tends to be to go find verticals, but that's kind of my preference. Some companies, some leaders just would prefer to build horizontal capability, and that's fine, too.

The Humanoid Hub

40,043 次观看 • 7 个月前