Loading video...

Video Failed to Load

Go Home

Storage just stole the show. Between KV cache offload and SSD prices going vertical, storage went from afterthought to headliner. Vik Malyala walks through Supermicro's storage stack. "Storage, whether it be for KV cache offload or just the incredible rise in cost of SSDs, is the primary thing to...

61,919 views • 21 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Dylan Patel of SemiAnalysis says a worse GPU with better storage and memory now beats the best chip without them, so buying the newest GPU alone no longer wins inference. So, an AMD GPU with more memory can outperform Nvidia in some cases. "So what we have is we have over $80 million of compute, GPUs from Nvidia, AMD, TPUs from Google, Trainium from Amazon, and we run this benchmark constantly on the newest inference engine, newest drivers, newest PyTorch version, whatever it is." "Every day it runs on an automated CI, and we run it on all the latest Chinese models, from GLM, Zhipu, Moonshot, Kimi, Alibaba, all these models we run." "Initially, when we were benchmarking the difference between these chips and different engines, different schemes for parallelism, we were just running it fixed context length." "But now with Agent X, we've analyzed over $5 million worth of Claude Code traces. This is real production traffic that people have donated to us as well as internally generated. Now we know what the actual agent workload looks like." "And then as we implement that and run those benchmarks, it turns out yes, the chip you're using is very important, but now even more important is how are you handling this memory offload?" "And so while an Nvidia GPU is faster than an AMD GPU in most cases, because AMD GPUs have more memory, they actually end up outperforming in some cases." "Or you can have a worse GPU, but a much better storage solution, and now you can outperform what the best GPU can do without those solutions. So just buying the newest and latest GPU alone doesn't get you the best inference economics." "Actually, you need to layer in all these other innovations including storage and memory." [ Who's the top player on your chart? ] "That really is a difficult multivariable problem. And generally that means you need to have, yes, you need to have the best GPU, a GB300, but you also need to have the best storage solutions. And so I won't spoil who's the best right here, but I will say that storage solutions matter a lot and memory solutions matter a lot, as does your front-end networking. That matters a lot."

Fireside Alpha

175,586 views • 20 days ago

Dylan Patel on the importance of memory and storage Two key quotes: "An $NVDA GPU is faster than an $AMD GPU in most cases, but because AMD GPUs have more memory, they can outperform Nvidia in certain workloads." “It is a difficult, multivariable problem. Generally, you need the best GPU, such as a GB300, but you also need the best storage solutions. I will not spoil who comes out on top, but storage solutions matter a lot, memory solutions matter a lot, and frontend networking also matters significantly" Full Quote: “We have over $80 million of compute: GPUs from $NVDA and $AMD, TPUs from Google, and Trainium from Amazon. We constantly run this benchmark using the newest inference engines, drivers, PyTorch versions, and other software. It runs every day through automated CI across the latest Chinese models from GLM, Zhipu, Moonshot, Kimi, Alibaba, and others. Initially, when we were benchmarking the differences between these chips, inference engines, and parallelism schemes, we used fixed context lengths. But with Agent X, we have now analyzed more than $5 million worth of Claude Code traces. This is real production traffic that users have donated to us, combined with internally generated data, so we now understand what an actual agent workload looks like. When we implement those workloads and run the benchmarks, it turns out that the chip you are using is very important, but how you handle memory offload can be even more important. An Nvidia GPU is faster than an AMD GPU in most cases, but because AMD GPUs have more memory, they can outperform Nvidia in certain workloads. Similarly, you can use a less powerful GPU with a much better storage solution and outperform the best GPU when it lacks those solutions. Simply buying the newest GPU does not necessarily give you the best inference economics. You need to layer in other innovations, including storage and memory.” Interviewer: “Who is the top player on your chart? Can you tell us?” Dylan Patel: “It is a difficult, multivariable problem. Generally, you need the best GPU, such as a GB300, but you also need the best storage solutions. I will not spoil who comes out on top, but storage solutions matter a lot, memory solutions matter a lot, and frontend networking also matters significantly.”

Daniel Romero

38,220 views • 1 month ago

ABC’s George Stephanopoulos attempted to trip up Secretary of State Marco Rubio by asking the same question about Venezuela three different times. “What is the legal authority for the United States to be running Venezuela?” And three times, Rubio shut him down. RUBIO: “I explained to you what our goals are and how we’re going to use the leverage to make it happen.” “As far as our legal authority on the quarantine, simple. We have court orders. These are sanctioned boats, and we get orders from courts to go after and seize these sanctions. Is a court not a legal authority?” STEPHANOPOULOS: “Is the United States running Venezuela right now?” RUBIO: “Well, I’ve explained once again, I’ll do it one more time.” “What we are running is the direction this is going to move moving forward, and that is we have leverage.” “This leverage we are using and we intend to use. We started using already. You can see where they are running out of storage capacity. In a few weeks, they’re going to have to start pumping oil, unless they make changes.” “That leverage that we have with the armada of boats that are currently positioned, allow us to seize any sanctioned boats coming into or out of Venezuela loaded with oil or on its way in to pick up oil.” “We can pick and choose which ones we go after. We have court orders for each one.” “That will continue to be in place until the people who have control over the levers of power in that country make changes that are not just in the interest of the people of Venezuela but are in the interest of the United States and the things that we care about.” “The legal authority is the court orders that we have.”

Overton

213,086 views • 7 months ago

Most $SUI holders know one thing about the token: total supply is capped at 10 billion. They have never read the mechanic that makes that cap matter. It is called the Storage Fund. And it is the most important thing in the $SUI tokenomics docs that nobody is talking about. Here is exactly how it works: Every time a transaction adds data to the Sui blockchain, the user pays a storage fee. That fee does not go to validators directly. It goes into the Storage Fund, a pool of SUI that never fully depletes. Here is where it gets interesting. The Storage Fund has its own stake in the network. It earns staking rewards the same way every other stakeholder does. Those rewards are then distributed to validators to compensate them for storing historical data. This solves a problem every other blockchain ignores: When a new validator joins Sui, they have to store all the historical data from transactions that happened before they existed. Why would a new validator pay to store someone else's old data? The Storage Fund pays them for it. Past users who created the storage requirements in the first place funded the pool. Future validators get compensated from that pool indefinitely. The fund pays out only the returns on its capital, never the principal. It cannot be drained. It is designed to survive forever. Now here is the part that directly connects to $SUI token value. The Sui docs state this explicitly: Deflation is a feature of Sui, not a bug. Here is why: Total supply is capped at 10 billion SUI. As network activity increases, more transactions are processed. More transactions mean more storage fees flowing into the Storage Fund. As the Storage Fund grows, it holds more SUI. More SUI held in the fund means less SUI in active circulation. Less circulating supply against the same or growing demand means the value of each SUI token increases. Network growth directly reduces circulating supply. That is not speculation. That is the economic model built into the protocol at the architecture level. One more detail worth knowing: If you delete data you stored on chain, you receive a partial refund of your original storage fees. The system charges for storage, rewards deletion, and compounds the fund's stake indefinitely. Most people holding $SUI today are pricing the speed narrative: The parallel transaction processing. The sub-second finality. The Move language safety. They have not started pricing the storage fund deflation mechanic. That gap between what the tokenomics actually does and what the market currently understands is where the long-term thesis lives. The people who read the docs always buy before the people who read the price.

2xnmore

47,562 views • 2 months ago

Free electricity for every American? Chamath explains how it’s possible. Chamath Palihapitiya: “I think the president should try to create a $300B-$500B tax equity fund and help eliminate the electricity costs of 50-100M American households.” “If you actually go to solar and storage, for every 15 million homes, you take a terawatt hour of demand off the grid, that's an extra terawatt hour that you don't need to build.” “So the reason why I like solar and storage is you make these folks so self-reliant.” “Now all of a sudden we're actually catching up to China faster than we would've otherwise.” “We actually get a 2-for-1. The consumer homes don't have that anxiety anymore. They don't have the bill. And because they're generating their own power, they're off the grid, and now what that means is that extra terawatt hour can actually go to the commercial and industrial applications, including these datacenters.” “The president preserved the ability to make those kinds of investments and be tax advantaged for doing it in the One Big Beautiful Bill.” @jason: “And if we want to live in the age of abundance, and we want to bring the bottom half of the country up, and we want affordability, where do people spend their money?” “Food, groceries, their rent, and their electricity, and the utilities.” “So it doesn't come out of Americans' pockets and their taxes, it comes out of corporations, which are printing money.” “If you wanna talk about redistribution of wealth, this is a way for the great American Mag 7 corporations to give something back to Americans.” Chamath: “We will look back and I think that we will want to have seen these big companies who are unbelievably profitable, step up on behalf of American homeowners.”

The All-In Podcast

526,669 views • 6 months ago