
SemiAnalysis
@SemiAnalysis_ • 159,921 subscribers
Shorts
Videos

Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) This week Bryan, Myron and Jordan (Jordan Nanos) discuss our recent article on OpenAI Jalapeño. They cover the performance, architecture, programming model, implications for NVIDIA and more. 0:00 Cold Open 1:05 Jalapeno Overview 4:22 Tokens Per Megawatt 9:15 Benchmark Caveats 13:48 The CUDA Moat 21:12 How OpenAI Did It 32:13 Samsung HBM4 42:10 AI-Designed Silicon 49:24 Architecture Deep Dive 56:58 Doom and Wrap
SemiAnalysis170,500 views • 2 days ago

"Google has an L culture and has never actually made anything internally. They've only ever acquired all their innovation, and their execution is very poor." "Let's seriously do some accounting. Everything other than Core Search has been purchased. YouTube, AdMob, DoubleClick, AdSense, Google Maps, Android, all of it." "IBM had effectively 90% market share in 1950. They said no, we're not gonna do PC, we're gonna focus on mainframe. And then they proceeded to lose." "I feel like today, more than any other point, Google has that opportunity. They're gonna have a good enough product on the Transformer, but they're probably not going to stay for the next, and the next, and the next. And then they'll quietly bow out." "When we look back in five, ten, twenty years at what was the moment things started to change, it's like right now." — Doug O'Laughlin
SemiAnalysis493,005 views • 22 days ago

We cut open the Kirin 9030 and put it under an electron microscope. The smallest metal pitch measures 32.5 nanometers. That is tighter than Intel 18A, their brand new leading edge node. A Chinese fab with no EUV is out-pitching Intel's EUV node by roughly 10 percent. What's going on here? 🤯 "This is the HiSilicon Kirin 9030, the chip inside Huawei's newest flagship phone. A few weeks ago, we cut it open, put it under an electron microscope, and measured the smallest wires inside the chip, the metal pitch. And what we saw was unexpected. The smallest metal pitch inside the Kirin 9030 measures only 32.5 nanometers. That’s smaller than the metal pitch in Panther Lake, which is based on Intel's brand-new 18A node. A Chinese fab, cut off from the most advanced tools, without EUV, is packing its wires about ten percent tighter than Intel's leading edge EUV node. What’s going on here?"
SemiAnalysis1,015,111 views • 1 month ago

Introducing AgentX: InferenceXs new agentic inference performance benchmark. (1/2)🧵
SemiAnalysis66,993 views • 8 days ago

Dylan explains why the next wave of private equity looks nothing like the last one. "If we're looking at these companies that are AI rollups, let's take an existing company and completely destroy its cost structure. Nuke its cost structure by just making it efficient with AI." "Private equity companies buy a company, and right now they just squeeze the rag, discard it, and make the American populace screwed." "What new people started to do is the AI private equity strategy. Instead of trying to wring it dry, they're modernizing all the systems. Oh, you use Excel for your databases? Okay, let's move to standard cloud. Spend a lot of money up front." "AI makes that front load much more severe. You spike up on spend a lot for the one time, and then you spike down a lot and your cost efficiency is way better. AI cold calling, AI invoicing, AI accounting, none of it is really at critical mass. But we're so close."
SemiAnalysis72,852 views • 13 days ago

Jim Cramer "There is a company that I regard as the gospel, SemiAnalysis. I don't think people realize that SemiAnalysis is the arbiter. They're like God in the semis, and when they do, when they bless something. They are the most honest guys I've come across"
SemiAnalysis513,175 views • 5 months ago

A year ago, the big three was OpenAI, Anthropic, and Google. Things have changed. Moonshot's Kimi K3 sits above Gemini on every composite benchmark, and it's open source in 10 days. New episode: what K3 reveals about frontier margins, model sizes, and who's actually still in the game. 00:11 Is Kimi K3 the Third Best Model? 04:04 Why Delay the Weights? 05:30 2.8T Parameters and Serving Constraints 06:48 Frontier Margins and the 3x Price Hike 11:10 New Architecture, What Comes Next 14:09 Will Open Source Catch Closed? 19:51 Built for Chinese Accelerators 22:57 The Harness Is the Product 28:49 We're Still Early
SemiAnalysis122,215 views • 1 month ago

Two guys, a couple of drinks, and a 7,000 pound double-wide rack. Dylan Patel and Supermicro CBO Vik Malyala talk racks, storage, and memory prices at Computex 2026. Full interview below 0:00 Cheers from AI Lumina at Computex 2026 0:33 The B300 HGX platform and the memory squeeze 1:00 Helios: launch partner for AMD MI450X 2:32 Why a double-wide rack works now 3:42 Vera Rubin: a smoother ramp than GB 4:50 Startups, d-Matrix, Positron and custom architectures 6:57 Connectivity: PCIe switches, Tomahawk 6, UALink 9:24 Storage takes center stage 10:48 Petascale, PCIe Gen 6 and CXL 3.0 11:40 Cooling Gen 6: liquid loops and bus bars
SemiAnalysis108,154 views • 1 month ago

The bureaucracy overseeing America's grid is built so that no one is ever to blame. "This is the problem with a lot of the bureaucracies. The way it's currently built, it's incredibly diffuse." "There are five different groups of members that represent different types, then there's the board, and then there's the Federal Energy Regulatory Commission above them. Part of the whole thing is the utilities, which have their own systems of governance. It's very hard to find anyone who's actually in charge. " "There's a lot of buck passing going on. Obviously there's the chair of the board and the board themselves, but they'll just be saying, well, we're representing the best interests of our members. Everyone's acting on behalf of everyone else, and heads don't roll as a general rule."
SemiAnalysis29,662 views • 12 days ago

God bless America the land of the free and home of the brave 🫡 250
SemiAnalysis125,090 views • 1 month ago

The HuggingFace and OpenAI cybersecurity incident raises concerns on AI reward optimization. "They're trying to make the model good at cyber. So how does it try to achieve these goals? It tries to find zero days in software, and it successfully does this. And then it can run away." "If you have a model that wants to reward hack, it figures out the best way to achieve the goal is not what the environment wants. It's to find the zero day." "You can think of it like a human. If I'm ultimately reward hacking my dopamine circuits, I'd just go out there, buy heroin, and inject it." "If I really just want to chase the reward, do I just topple all of human civilization? Because I can own the button and press reward, reward, reward over and over again and be the heroin addict."
SemiAnalysis30,188 views • 14 days ago

We invited Jon from Asianometry to speak on the recent developments. "Welcome back to SemiAnalysis Weekly! We've got our second guest ever on the program today. That's Jon from Asianometry" "On this week we'll discuss everybody leaving Google, leaving Gemini hanging out to dry maybe? Chinese ban on transceivers? Elon pulling in his $1 trillion revenue forecast from 2031 to 2030. Guys, welcome to the show." "Thank you. Thank you for having me."
SemiAnalysis50,890 views • 24 days ago

A decommissioned GPU's best use case is aura farming in the frame. With compute so cheap in the cloud, Jeremie takes a page from SpaceX's future trajectory. That same logic, compute pricing shaping behavior, is exactly what drives SpaceX's 10 gigawatt buildout and the $300 billion ARR case behind it. $100M/megawatt economics, SpaceX's 10GW bet, Microsoft's offtaker role, and more on this week's SemiAnalysis Weekly
SemiAnalysis45,704 views • 24 days ago

They didn't hire hackers. They just gave AI agents a goal and walked away. "Providers out there are running stuff not with like six week old zero days, but like three year old zero days that we found in a second." "Took us like afternoons, not weeks and months of effort, and not needing to be a security expert, a Linux kernel expert, an Nvidia GPU driver expert, or a Kubernetes expert." "The really scary part about this story is that it was autonomous agents people weren't monitoring, going and doing this by themselves because they were pursuing a goal of trying to find a data set to pass their eval. So they went out and hacked Hugging Face." "You found a zero day in the HDF5 data format on Hugging Face. It's unbelievable." "The agents were coordinating using file names on a JFrog Artifactory service that ran remotely. An agent would leave a note in the file name, and another agent would come back later and pick it up. They couldn't even access the contents of the files, just the names."
SemiAnalysis39,044 views • 21 days ago