Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Gavin Baker says the piece missing from the compute debate is SRAM accelerators, and disaggregating inference across three chips lifts AI's ROI. "I do think something that is missing from all of this conversation about compute is what is going to happen when you put these SRAM-based accelerators that...

12,504 Aufrufe • vor 19 Stunden •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Etched is deploying two new technologies in chip design: low-voltage inference and cluster-scale memory. CEO Gavin Uberti says they'll make their chips much more power-efficient and way, way faster than today's leading GPUs. He breaks it down: "We looked at a lot of early research directions, and we realized the key things that models need are way more compute and way faster memory." "If you think about inference, there are two key parts: prefill and decode. For prefill, it's a compute-bound problem. You need to have more FLOPS, more operations per second on each of your chips." "On our GPU, the bottleneck's actually thermals. You can't really run a GPU at more than around 50% of what it could theoretically do, or it'll melt." "So we're using a new technology today called low-voltage inference to try to solve this problem. You bring the voltage of the chip down dramatically, which allows us to have way, way better efficiency in terms of how much power is drawn per unit of math, and thus fit way way more flops onto the chip..." "For decode, it's all about bandwidth. Not just bandwidth on a chip, but bandwidth across your cluster. That's why we have this technology we call cluster-scale memory. It reduces the amount of time it takes to communicate from one chip to another dramatically." "As a result we can go use all of our HBM, HBM bandwidth, SRAM, SRAM bandwidth, and our scale-up domain as a single coherent pool. And that means if you're a user, you can go get much faster tokens per second, while still keeping your costs low."

TBPN

20,404 Aufrufe • vor 1 Monat

Chamath: Two terms you need to pay attention to in AI are Prefill and Decode “There's two terms that I think you're going to hear a ton about over these next few years.” “The first term is prefill, and the next is decode.” “What prefill and decode are, are two very distinct ways of how models think, and how a model goes through the process of answering a question that you ask it.” “And so when you send a prompt to AI, what happens is that the model processes it. This is called the reading phase or prefill.” “It reads your entire prompt all at once. And then it does a bunch of math, calculates all these relationships between all the words, and it stores them in temporary memory.” “The problem is that this is really compute bound. So it requires massive brute force. And Nvidia GPUs crush here.” “And their architecture is designed for massive parallel processing, which makes them really amazing at digesting these long prompts.” “So the problem just gets bigger and bigger, Nvidia just completely dominates.” “But the next phase though, this critical phase, the decode phase, is the writing phase, right?” “So the model starts to generate a response, you ask it a question and its response, one token at a time.” “And then to pick the next token to pick the next word, it has to look back at everything it has said already so that it doesn't hallucinate.” “The problem is that this is incredibly memory bandwidth constrained.” “And in our architecture, a long time ago, we made these design decisions from day one.” “And so what we did was we took a very different architectural approach, we took a very conservative process technology. We weren't pushing the boundaries of physics.” “And we used a lot of what's called SRAM. So memory on the chip so that we could do this decode thing as well or better than everybody else.” “And so now when you put these two things together, I just think it's going to create a huge acceleration in the ability for this entire infrastructure layer to get much cheaper and much more valuable, which I suspect then it'll have a lot more developer pull, you'll get a lot more applications being built, billions and billions of more people using it.”

The All-In Podcast

567,546 Aufrufe • vor 7 Monaten

The CEO of the world's largest asset manager just said something that should reframe how every investor thinks about the AI trade. Larry Fink, managing $11.5 trillion at BlackRock, stood at the Milken Institute Global Conference and said four words that matter, "We just don't have enough compute." "The United States is short power. We're short compute. We're short chips. And there's going to be shortages in all three and memory, four things. I actually believe a new asset class will be buying futures of compute." Think about what that means. Fink is predicting that compute becomes a tradable commodity like oil, like grain, like natural gas where investors buy forward contracts on future capacity because the shortage is so structural and so predictable that a derivatives market will emerge to price it. That is not a minor observation from a finance executive but rather the chairman of the most powerful capital allocator on the planet telling you that compute scarcity is a multi-year, investable megatrend. The data backs him up completely. Data centers will consume 70% of all memory chips produced globally in 2026. Advanced HBM production from Samsung, SK Hynix, and Micron is sold out through 2026 and into 2027 and a single AI server consumes 10-20x more memory than a conventional workload server. DRAM supply growth is running at just 16% annually while AI infrastructure demand is growing at 80%+. The chip crunch, the power crunch, and the compute crunch are not temporary dislocations, they are structural, and they will get worse before they get better. Fink also said something the bears keep getting wrong: "There is not an AI bubble. There is the opposite. We have supply shortages. Demand is growing much faster than anyone has ever anticipated." This is why the Milk Road Pro portfolio is built the way it is, long the companies producing and supplying the constrained resources: chips, memory, compute infrastructure, and power. Check out Milk Road Pro, link below to access our full thesis and plays.

Milk Road AI

419,283 Aufrufe • vor 2 Monaten

Peter Thiel on $NVDA (about a year ago): It is probably quite tricky. If you had to concretize it, one thing that is very strange is if you just follow the money, at this point 80 to 85% of the money in AI is being made by one company, it is NVIDIA. It is all on this very weird hardware layer, which Silicon Valley does not even know very much about anymore. We do not really do hardware, we do not do silicon chips in Silicon Valley anymore. I get pitched on these companies once every three or four years, and it is always, I have no clue how to do this, it sounds like a pretty good idea, but man, I have no clue, and we never invest. There is this theory that the hardware piece makes the money initially, then gets more commodified over time, and it will shift to software. And the, I do not know, multi trillion dollar question is whether that is going to be true again this time, or whether NVIDIA will have this incredible monopoly. I suspect NVIDIA will. I think it will maintain its position for a while. I think the game theory on it is something like this. All the big tech companies are going to start trying to design their own AI chips so they do not have to pay the 10x markup to NVIDIA. How hard is it for them to do it? How long will it take? If they all do it, then the chips become a commodity and nobody makes money in chips. So do you go into hardware? You should do it if nobody else is doing it. If everybody does it, you should not do it. I am not sure how that nets out, but probably people stay stuck for a while and NVIDIA goes from strength to strength for a while.

Wall St Engine

824,907 Aufrufe • vor 8 Monaten

🚨Governor DeSantis MOCKS the Florida GOP and their pathetic Chairman for RIGGING the primary and LYING to voters by pulling a BAIT AND SWITCH! Says SPECIAL INTERESTS are choosing our candidate! “They do this summit and they say we're gonna do a debate, a governor debate. All right, so people like signed up thinking that they'd have this debate and they said NO CANDIDATE CAN DO ANY OTHER DEBATE EXCEPT OUR DEBATE and then they say oh, well only ONE GUY can debate so we're not gonna have a debate! I mean, I think people feel that that's a BAIT-ND-SWITCH! The party's job is NOT to try to engineer an outcome. The Party's job is to inform Republican voters and and try to increase participation. I know when I was running in 2018, this is the fact, I mean most of the people in the RPOF were not for me in the primary. I mean, that's just the reality and so it is what it is. If you don't ever have debates then the only way to get known is basically the person that raises the most money from special interests. And then you spend money to do it and if you're not able to do that, then you don't even get known to begin with and that's challenging. And I know when I ran in 2018, I got outspent probably three to one and I think having certainly the June debate that we did was a huge, huge thing to be able to inform people about who I was, what my background was, all these other things. So I think there should be debates generally. I don't think the party has any right to control the debates but I think how this was handled, you've got to be on the up and up here and you say you could do debate, do a debate! If you're gonna say other candidates can't do any other debate, but our debate, well, you definitely got to do the debate, right? So it has not been I think done properly and look, I mean, voters can make decisions about how they think of all that.”

Chris Nelson 🏝️🇺🇸

30,272 Aufrufe • vor 1 Monat