Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

A masterclass on Google's TPU v8 Networking. Two TPU chips? Pssh. We already knew workload-specific silicon was here. But two scale-up networking topologies? That's the actual Google TPU news. Workload-specific interconnects. Think about that. New Semi Doped with Vikram Sekar and Austin Lyons. Copper? Yep. Optics? Yep. What we...

92,064 Aufrufe • vor 3 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

TPUs Via Cloud Next, Intel Earnings, Foundry Scarcity Ben Bajarin and Jay Goldberg 顾忠南 discuss: 00:00 - Ben and Jay introduce an action-packed episode covering a massive week of semiconductor announcements. 00:10 - Ben recaps his experience at Google Next, highlighting the launch of the new TPU v5p and v5i accelerators. 02:35 - The hosts discuss the physical design of the new inference chips and how they differ from training hardware. 03:04 - A deep dive into memory architecture explores the use of HBM and the "Boardfly" networking innovation. 04:15 - Jay notes reports of Google working on SRAM-based inference chips to solve the HBM supply bottleneck. 07:22 - Google's leadership draws a historical parallel between today's AI training costs and their early days of indexing the web. 11:11 - The duo critiques the enterprise focus of the Google Next keynote and the current productivity gap in Gemini. 15:24 - Intel's recent earnings report shows a significant beat on revenue and guidance driven by a "CPU resurgence". 18:31 - Ben and Jay debate the strategic advantage of Intel owning its own fabs during a global capacity shortage. 25:51 - The conversation shifts to Intel's manufacturing roadmap and the potential for Tesla to become a 14A customer. 33:23 - An analysis of Intel's new financial reporting structure reveals the massive internal scale of their foundry business. 40:27 - Recent drama at the TSMC Symposium is discussed, specifically regarding the high cost of ASML's High-NA EUV tools. 44:23 - Growing demand for AI and memory is driving Wafer Fabrication Equipment (WFE) forecasts to record highs. 47:06 - Ben and Jay conclude by debating the long-term durability of the current semiconductor bull cycle

The Circuit

12,020 Aufrufe • vor 3 Monaten

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,179 Aufrufe • vor 7 Monaten

New interview: Reiner Pope, co-founder/CEO of MatX A counterintuitive throughput insight: “Low latency means small batch sizes. That is just Little’s law. Memory occupancy in HBM is proportional to batch size. So you can actually fit longer contexts than you could if the latency were larger. Low latency is not just a usability win, it improves throughput.” We get into: • The hybrid SRAM + HBM bet, and why pipeline parallelism finally works • Why sparse MoE drives MatX to “the most interconnect of any announced product” • Why frontier labs are willing to bet on an AI ASIC startup • Memory-bandwidth-efficient attention, numerics, and what MatX publishes (and what it does not) • Why 95% of model-side news is noise for chip design • The biggest challenges ahead 00:00 “We left Google one week before ChatGPT” 00:24 Intro: who is MatX 01:17 Origin story: leaving Google for LLM chips 02:21 GPT-3 and the “too expensive” problem 04:25 Why buy hardware that is not a GPU 05:52 Overcoming the CUDA moat 08:46 Early investors 09:35 The name MatX 09:59 The chip: matrix multiply + hybrid SRAM/HBM 12:11 Why pipeline parallelism finally works 14:22 Reading papers and Google going dark 15:20 Research agenda: attention and numerics 17:06 Five specs and meeting customers where they are 19:24 Why frontier labs are the natural first customer 20:32 Workloads: training, prefill, decode 22:18 Little’s law and the throughput case for low latency 24:29 Interconnect and MoE topology 26:35 Inside the team: 100 people, full stack 28:32 Agentic AI: 95% noise for hardware 30:35 KV cache sizing in an agentic world 32:11 How MatX uses AI for chip design (Verilog + BlueSpec) 34:23 Go to market: proving credibility under NDA 35:12 Porting effort for frontier labs 36:34 Biggest skepticism: manufacturing at gigawatt scale 37:32 Hiring plug Vikram Sekar

Semi Doped

19,439 Aufrufe • vor 4 Monaten

MEET THE NVIDIA KILLER: OpenAI bet $10 BILLION on this company that makes chips 20x faster than Nvidia's. If this plays out as expected, it’s over for Nvidia. Cerebras Systems just locked in 750 megawatts of computing power to OpenAI through 2028. For reference: that's equivalent to the annual power consumption of 600,000 US homes. The deal? Over $10 billion. Here's what nobody understands: Cerebras doesn't make normal chips. Nvidia sells you thousands of tiny chips that you connect together. Cerebras makes ONE chip. A single wafer-scale processor the size of a dinner plate. 900,000 AI cores. 4 trillion transistors. All on one piece of silicon. The result? When OpenAI tested it, Cerebras ran inference 20X FASTER than Nvidia GPUs. That's not incremental improvement. That's a different category of performance. But here's where the story gets wild: Four months ago, Cerebras was a struggling company. Their IPO filing revealed that 87% of their revenue came from ONE customer: G42, a UAE-based AI firm. The US government launched a national security review. G42 had ties to Huawei. Ties to China. The IPO collapsed. Investors panicked. Cerebras withdrew their filing in October 2025. Most startups would've been dead. Instead, Cerebras did the opposite. They raised $1.1 billion at an $8.1 billion valuation. Kicked G42 out of the cap table entirely. Got CFIUS clearance. Then landed the OpenAI deal. Now they're raising ANOTHER $1 billion at a $22 billion valuation. They more than DOUBLED their valuation in 4 months. From near-death to $22 billion. While getting rid of their biggest customer. Why OpenAI chose them: ChatGPT has 900 million weekly users. Sam Altman keeps saying they have a "severe shortage" of compute. They need SPEED, not just power. When you ask ChatGPT a question, there's a loop happening: You send request → model thinks → sends response back Nvidia chips are fast at training models. Cerebras chips are built specifically for inference. For real-time responses. For the exact bottleneck OpenAI is trying to solve. Sachin Katti from OpenAI said it best: "Cerebras adds a dedicated low-latency inference solution to our platform. That means faster responses, more natural interactions, and a stronger foundation to scale real-time AI to many more people." In other words: "We need this to scale ChatGPT." The competitive landscape just shifted: Nvidia announced a $100 billion deal with OpenAI in September. But it's still not finalized. Meanwhile, Cerebras closed their deal before Thanksgiving. And it's ALREADY being deployed. Here's the part that should terrify Nvidia: In December, Nvidia bought Groq for $20 billion. Groq makes fast inference chips. Just like Cerebras. So why would Nvidia spend $20 billion buying a competitor to something they supposedly already dominate? Because they know what's coming. Inference is the new battleground. And Cerebras is winning it. The IPO is coming Q2 2026. After this OpenAI deal, Cerebras now has: ✓ IBM contracts ✓ Department of Energy contracts ✓ OpenAI locked in for 3 years ✓ $22 billion valuation ✓ CFIUS clearance ✓ Zero customer concentration risk They went from 87% revenue dependency on one customer to the most diversified chip company outside Nvidia. In four months. The lesson? Smart money doesn't follow headlines. It follows where the AI leaders are actually spending. OpenAI didn't announce this deal for publicity. They need Cerebras hardware to scale ChatGPT. That's a $10 billion vote of confidence. While everyone's watching Nvidia stock, the real war is happening in inference. And the company with ONE giant chip just beat the company with thousands of tiny ones. What do you think happens when Cerebras IPOs?

Ricardo

28,088 Aufrufe • vor 7 Monaten

Nebius will be the first neocloud to hit $1 trillion dollar company and here is exactly why (Save this). As dylan patel says Jensen Huang absolutely hates a world where the hyperscalers have all the power. A world where Microsoft, Amazon, and Google are the only ones building compute is a world where Nvidia is slowly being squeezed by a handful of customers all simultaneously developing custom chips to replace Nvidia GPUs entirely. Google's TPU, Amazon's Trainium and Microsoft's Maia all exist for one reason, to cut Nvidia out of the stack and Jensen knows it so he is playing a long game most investors haven't registered yet. By funding NeoClouds and NeoLabs at scale, Jensen is deliberately engineering a multipolar compute world where no single hyperscaler can dictate terms and where Nvidia hardware remains the default infrastructure layer regardless of which model or platform ultimately wins. Nvidia has deployed roughly $40 billion in AI ecosystem investments across OpenAI, Anthropic, CoreWeave, Nebius, xAI, and dozens of infrastructure companies, all running almost exclusively on Nvidia chips, cementing GPU dependency across the entire AI stack.sedaily Every neocloud that survives and scales becomes a permanent Nvidia GPU customer structurally opposed to the hyperscalers building custom silicon expanding Nvidia's market while simultaneously weakening its biggest competitive threat. Dylan Patel described the neocloud ecosystem as throwing bait into the water and letting the best fish survive, warning that many heavily-backed teams will fail, but the ones that emerge will pull hundreds of millions in ARR right out of the gate. Nebius is that fish because it's the only neocloud operating at hyperscaler scale while remaining fully purpose-engineered for AI workloads from silicon to software. The numbers confirm Nebius has already cleared the survival bar that will eliminate most of the 200+ neoclouds competing right now. Revenue hit $399 million in Q1 2026, up 684% year-over-year, backed by $46 billion in contracted backlog, 3.5 GW of contracted power across seven site and a target of $7–$9 billion in annualized revenue by year-end. When Google approached neoclouds about deploying TPUs, Nebius said no, its Chief Revenue Officer noting that demand is 99% for Nvidia GPUs and that TPU interest comes almost entirely from former Google employees rather than the actual market. That alignment with Nvidia's ecosystem, at this scale, with this backlog, and this level of strategic backing is why Nebius sits in a category of one among the neocloud field. Patel framed the broader play correctly, every neocloud that survives makes Google's TPU and Amazon's Trainium structurally weaker simply by existing and five years from now, the winners will have reshaped the entire compute landscape in Nvidia's favor. Nebius is already hundreds of millions in ARR ahead of the competition while most of the field is still treading water. Milk Road subscribers are already up massively on the Nebius trade, and we are tracking the neocloud buildout as Nvidia works to reshape the entire compute market. Come join Milk Road Pro for our full Nebius breakdown, the valuation framework, the revenue targets we are watching, and the AI infrastructure names we like next for just $1. Link below!

Milk Road AI

92,855 Aufrufe • vor 1 Monat

Jensen Huang is investing in every photonics company he can find and the reason why tells you everything about where AI is headed (Save this). Lip-Bu Tan, the CEO of Intel says, when he looks for investment opportunities, he looks for the bottleneck and right now, the bottleneck is the interconnect, the pipes that move data between chips inside an AI data center. That is why he backed Credo Semiconductor, Astera Labs and Celestial AI on the optical side. Here is the simple version of what the interconnect bottleneck actually means. Think of an AI data center like a city, the GPUs are the buildings where all the work happens but for those buildings to function, you need roads connecting them, fast roads that can carry enormous traffic without congestion. And those roads are now the single biggest constraint on AI performance. As clusters scale to hundreds of thousands of GPUs, traditional copper wiring is hitting its physical limits and that is where this entire sector comes in. Credo Semiconductor (CRDO) is the most direct pure play on this theme, Credo makes high speed cables and optical chips that connect GPUs inside data center racks. Their revenue tripled in fiscal 2026 to $1.3 billion, growing 272% year over year at its peak and four of the world's largest hyperscalers each individually account for more than 10% of Credo's revenue. Astera Labs (ALAB) solves the connection problem between different chip types. Astera makes the PCIe and connectivity chips that manage data flow between GPUs, CPUs, and memory without errors or slowdowns. Their revenue grew 93% year over year to $308 million in Q1 2026 alone. The optical companies are where the longer-term and potentially larger opportunity lives. Copper has physical limits, you can only push electrical signals so far before the signal degrades, the heat spikes and power consumption explodes. The solution is light, fiber optic connections that move data using photons instead of electrons which is faster, cooler and far more energy efficient. Jensen Huang made this clear at Computex 2026 because copper works as long as physically possible but at greater distances and larger scale, optics takes over. Coherent (COHR) is the most established optical company in this space. Coherent makes the lasers, transceivers, and optical components at the foundation of all fiber optic communications. Nvidia signed a multibillion-dollar purchase commitment and invested $2 billion directly into the company and their customer order books are already extending out to 2028. Marvell (MRVL) is the most comprehensive bet across the entire connectivity stack. Marvell makes chips for optical networking, PCIe switching and custom AI silicon. Jensen Huang called Marvell the next trillion dollar company at Computex 2026 and backed it with a $2 billion Nvidia investment. Marvell also acquired Celestial AI, the exact company Lip-Bu Tan backed for $3.25 billion, gaining photonic fabric technology delivering 16 terabits per second of bandwidth. Lumentum (LITE), Corning (GLW), and Ciena (CIEN) round out the major public names. Lumentum received a $2 billion Nvidia investment for laser and photonics components. Corning known mostly for phone glass received $500 million from Nvidia for optical connectivity work and is up over 100% year to date. Ciena runs the optical networking systems between data centers and is seeing analyst price targets raised on the back of the AI optics boom. Every time a hyperscaler spends a billion dollars on Nvidia GPUs, the surrounding infrastructure, cables, switches, transceivers, optical components has to be upgraded to match. The smarter the GPU gets, the more the interconnect matters. Nvidia has committed at least $6.5 billion to photonics companies in the past 4 months alone and the companies building the roads between the GPUs may end up being just as valuable as the companies building the GPUs themselves. Follow me Melvin for more AI, semis and the next big market themes.

Melvin

152,406 Aufrufe • vor 1 Monat

NEW: Cerebras $CBRS CEO Andrew Feldman (Andrew Feldman) "When the chip on your shoulder is the largest chip the world has ever seen." "The demand for AI has outpaced everybody's expectation & everybody's forecast. & so everybody's chasing. They're chasing chips, memory, or data centers." We get into the chip 58x larger than any other, $20B OpenAI deal signed in 4.5 weeks, & what's actually going on with the 'big AI deals' Recorded 2 months after Cerebras' $5.5B IPO at a $56B valuation, where a first-day pop briefly hit ~$95B before settling toward ~$60B. We cover: › A "Cambrian explosion" of new chip architectures › $20B+ OpenAI deal: 750MW of inference compute over 3 years › Why inference, not training, is where the value is now › "We are behind" on the data center build-out › Free tokens & circular deals: "These are drug pushers" › Creating 1,000 millionaires › Nvidia's balance sheet & market strength › Sovereign AI & owning the stack › Co-design & data centers in space › A real 25-year path to ending cancer Cerebras builds AI infrastructure for training & inference. It went public in May 2026, & its products include inference, Wafer Scale Engine, AI supercomputers, AI model services, cloud, systems, & processors. Filmed at the Raise Summit in Paris. Thank you to Brex, MongoDB & AssemblyAI for helping make this trip & content series happen. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Andrew Feldman, Co-Founder & CEO at Cerebras Systems (00:49) Why hardware suddenly became the coolest industry in tech (01:51) What changed at Raise AI Summit (03:01) Inside the $20 billion Cerebras - OpenAI deal (05:50) What actually changes two months after an IPO (06:32) Turning 1,000 employees into millionaires (07:52) Staying sane during an AI gold rush (10:44) Life after the IPO plateau (11:45) The truth about the global data center shortage (12:56) Why data centers are borrowing jet engines for power (14:44) Are data centers really headed to space? (15:37) The shift to designing chips & software together (17:39) The biggest misconception about co-designing chips & software (18:31) Inside SpaceX's multi-billion dollar AI deals (20:06) NVIDIA's playbook for locking out competitors (20:54) The hidden cost behind free tokens (22:52) Andrew's response to Karp's sovereign AI thesis (24:18) AI's biggest win might be curing cancer (26:50) Peptides & biohacking (28:00) How AI could finally fix the broken education problem (29:46) The mentors who shaped Andrew Feldman's career

Molly O’Shea

494,431 Aufrufe • vor 1 Monat

Broadcom's CEO just exposed the real fight underneath Google's AI chip strategy. It is not Google versus Broadcom. It is Google and Broadcom trying to make Nvidia replaceable. Within two minutes at Bloomberg Tech, Hock Tan was asked whether Google bringing more chip design in house keeps him up at night. His exact words: "So we just compete against my own customer." Then he named the real enemy: "the real competitor facing all this is the GPU out of Nvidia." That is the part most people miss. Custom AI chips are not just cheaper GPUs. They are ownership claims. If Google owns the workload, the compiler stack, the cloud customer, and the TPU roadmap, Nvidia becomes a benchmark instead of the toll booth. But Broadcom is still in the room because independence is not binary. The hard part is not drawing a chip. The hard part is shipping generation after generation at scale, matching Nvidia's cadence, keeping networking tight, and making the whole system useful enough that developers do not care what silicon sits underneath. That is why Tan can say Google is trying to create customer owned tooling and still sound calm. Broadcom is not selling picks and shovels. It is selling the bridge out of Nvidia dependency. The numbers explain why this is suddenly a board level issue. Broadcom reported $22.2 billion of Q2 2026 revenue. Its AI semiconductor revenue hit $10.8 billion, up 143 percent year over year. For Q3, Broadcom guided AI semiconductor revenue to $16.0 billion, up more than 200 percent year over year. In the clip, Tan says Broadcom has exactly 6 custom AI accelerator customers. He says OpenAI has been engaged for over 2 years, its accelerator is already working in labs and data centers, and production is on track for late this year. That is the hidden mechanism: The AI labs are not becoming software companies with some chips attached. They are becoming capacity companies with model interfaces attached. Once your margin depends on tokens, latency, memory bandwidth, power contracts, packaging slots, networking gear, and a private accelerator schedule, the "model company" label starts to look like a costume. The precedent is Apple. Apple did not move into custom silicon because it wanted a cute chip branding story. It moved because the iPhone needed control over performance per watt, release cadence, and differentiation. A series chips in 2010. M1 in 2020. More than a decade of slowly pulling the bottleneck inside the company. But Apple still needed TSMC. That is the useful analogy for Google, OpenAI, and the other AI giants. They want Nvidia's margin pool. They want Nvidia's roadmap power. They want Nvidia's ability to decide who gets capacity first. But the first supplier they replace becomes the supplier they cannot live without. Broadcom is the customs officer at the border of private silicon. Second order consequence: AI company valuation will shift from model demos to infrastructure custody. Who owns the workload? Who controls the accelerator roadmap? Who has memory secured? Who can afford to keep a bad first generation alive long enough to get to the second and third? My bet: by the end of 2027, at least one major AI lab will be judged more by its custom chip execution than by its model benchmark lead. The model race is public. The margin race is being negotiated in silicon.

Andrej Drats

10,572 Aufrufe • vor 1 Monat

$NVDA $GFS NVIDIA’s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIA’s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groq’s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groq’s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groq’s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The company’s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groq’s solution is to move “control” into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groq’s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groq’s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groq’s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that “no useful models” fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groq’s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groq’s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth “upwards of 80 TB/s” and contrasts that with off-chip HBM bandwidth “about 8 TB/s,” asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes “TruePoint numerics,” including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the company’s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in “available speed” for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groq’s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as “tokens-as-a-service,” aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groq’s documentation states its API is designed to be “mostly compatible” with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groq’s go-to-market is not simply “sell chips,” but “sell predictable unit economics per token,” with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native “edge within the cloud” strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groq’s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groq’s marketing materials and partnerships provide signals about demand vectors. The company’s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groq’s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, which—if realized in deployments—can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groq’s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groq’s core proposition—deterministic, compiler-scheduled inference with predictable latency—aligns directly with the segment where GPU generality is least valued and where “good enough” programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIA’s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groq’s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The “technology acquisition” rationale is unusually strong in this specific case because Groq’s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIA’s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groq’s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groq’s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIA’s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chips—while not costless—could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship “inference capacity” in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIA’s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groq’s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groq’s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groq’s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a “preemptive acquisition” angle. The reporting identifies recent investors in Groq’s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groq’s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groq’s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groq’s compiler-led deterministic model is philosophically and practically different from CUDA’s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIA’s own rapidly improving inference performance complicates the “need” for Groq but does not eliminate the “option value.” NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIA’s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is “special” about Groq’s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows “hundreds” (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio manager’s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about “NVIDIA needs more inference speed” and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 Aufrufe • vor 7 Monaten

I read a lot of Peter Lynch. Met him once. The one rule I carry into tech investing is the most boring one he ever wrote, know what you own, down to the physics if the position demands it. For me that has meant living inside NVIDIA's stack for years, and pulling apart the alternatives next to it, Trainium, the TPU, every serious accelerator someone is willing to tape out against Jensen. I was also an early investor in Mellanox, the networking company NVIDIA bought to own the switched fabric the entire scale up era now runs on. So when the conversation turns to networking as the real moat, this is not theory to me. It is a position I watched become the thesis. You do not understand what you own until you understand what could take it. Gavin Baker at The Sohn Idea Contest just gave the most physically grounded read on AI infrastructure I have heard this cycle, and it is a Lynch lesson in disguise. The reframe that matters: The last terrestrial mega data center may already be on someone's drawing board. Everything else follows from two constraints, watts and wafers, and Gavin walks both down to first principles. That is the work. Most people are pricing the narrative. Lynch would have asked what the thing actually is. 1. TSMC is the global rate limiter Jensen reportedly visits every quarter asking to double or triple leading edge capacity. TSMC expands at roughly 5 percent. A handful of disciplined operators in Taiwan are the physical governor on the entire AI buildout. This is the part the bubble crowd misses. The constraint is not demand and it is not capital. It is one fab's deliberate refusal to overbuild. That stretches the cycle longer and smoother instead of bubble and bust. It reads like the mid 1990s capacity cycle, not a standard 25 year memory peak where a 60 to 70 percent price spike would be your signal to cut the weed and walk. I have held NVIDIA since 2016 for exactly this reason. Owning it meant understanding it. The thesis was never the chip. It was the chokepoint. 2. The most underestimated silicon is Trainium Consensus is still pricing a one horse race. Gavin's sharpest non NVIDIA call is AWS Trainium, specifically Trainium 3 ramping in the back half of 2026. Here is the part that took me a while to internalize from studying these architectures side by side. As frontier models go fully Mixture of Experts, inference stops being a matmul problem and becomes a networking problem. You need a switched scale up fabric, not just fast chips. Today two organizations on earth have a working one. NVIDIA and Amazon. NVIDIA's came from Mellanox, which is the whole reason I sized that position the way I did years ago, the bet was always that networking would decide this, not raw flops. The TPU is formidable in its own lane, but the scale up fabric is the moat people are not modeling, and it is why I track every accelerator, not just the one I own. 3. The neocloud moat is operational, not arbitrage The lazy take is that CoreWeave and Crusoe are just renting hyperscaler slack. Gavin's counter is that running dense GPU clusters is like driving an F1 car. Looks easy until you try it. Top tier neoclouds run 2 to 3x the hardware utilization per hour of lower tier providers. That is an execution and inventory moat, and it compounds. 4. The structural short nobody is pricing Watts and wafers eventually force the buildout off the planet. Gavin expects orbital data infrastructure to prove technical and economic viability within roughly two years and take meaningful share by the end of the decade. Space solves power with unattenuated solar and solves cooling with massive radiators in the satellite's own shadow. Dense single rack nodes stitched together with lasers into a virtual hyperscale cluster in orbit. The unpriced risk is everything that over expanded to serve a terrestrial buildout. Cooling, power, industrial equipment names sized for a curve that may bend down within seven years. The whole interview is a lesson in pattern recognition over narrative. Lynch built a career on retail investors knowing their companies better than Wall Street did. The same edge exists in AI infrastructure right now, it just requires you to understand watts and wafers instead of same store sales. If you are not modeling the physical boundaries of the stack through the lens of history, you are not underwriting the position. You are following it.

Ben Pouladian

93,581 Aufrufe • vor 3 Monaten

$AMD $MSFT Partnership is MASSIVE in 2026 🚀 If you were excited about my thread on $AMD $AMZN AWS long time partnership, you will be even more excited about what Microsoft gonna do with 2026 AMD EPYC "Venice". Historical Context: The relationship between AMD and Microsoft began in the early 2000s, with Microsoft initially focusing on Intel's x86 architecture for its Windows operating system and server products. However, AMD's entry into the server market with its Opteron processors in 2003 marked the beginning of a competitive dynamic that eventually led to collaboration. The partnership intensified with the launch of 3rd Generation EPYC "Milan" in 2021, powering Azure's N2D and C2D VM families. By 2025, Microsoft had integrated 5th Generation EPYC "Turin" into new compute-optimized instances, reflecting a strategic shift towards AMD for cost and performance benefits. This "Secret Weapon" breakthrough will mark another inflection point for AMD Microsoft Azure relationship, will probably be more aggressive than EPYC "Milan" moment in 2021. We can call it EPYC "Venice" moment 2026" 1. Technical performance of AMD EPYC "Venice" (2026) AMD's 6th Gen EPYC "Venice" processors, slated for 2026, introduce New Chiplet design breakthrough. a revolutionary chiplet interconnect fabric that redefines server scalability for AI. This isn't just faster silicon; it's a paradigm shift for Microsoft Azure , enabling hyper-efficient, rack-scale AI inference that slashes costs and latency while boosting throughput. ~Up to 256 Zen 6 cores, a 70% performance increase over "Turin," optimized for AI and HPC. ~Memory and Bandwidth: 1.6 TB/s per socket, doubling "Turin's" capability, with support for MR-DIMM/MCR-DIMM. ~Efficiency: 1,500-1,700W power draw, a 50% reduction, aligning with Microsoft's sustainability initiatives. ~Interconnect: PCIe 6.0 and a new chiplet fabric for rack-scale AI, reducing latency and enhancing scalability. 2. Why $MSFT will adopt $AMD YPYC Share to 50%+ in 2026. AMD EPYC Share: ~30-35% of Azure's x86 CPU-based business while Intel Xeon share is 65% Microsoft's Azure has been progressively integrating AMD EPYC, with "Venice" expected to expand this footprint: A. Dominance of AI Inference Workloads ~AI inference constitutes 80% of AI workloads in cloud environments, with latency-sensitive applications like chatbots, recommendation engines, and fraud detection requiring sub-second response times. ~"Venice's" 35x inference performance uplift directly addresses these requirements, outperforming Intel's offerings and custom Arm solutions in multi-threaded scenarios. B. Cost Efficiency and Operational Savings ~Azure's 2025 capex of $118B is under pressure to deliver returns. "Venice" can reduce operational expenses by $20-30B annually due to its power efficiency and performance gains, improving Azure's margins to 35-40%. ~The cost per inference operation is significantly lower with "Venice," estimated at 24-31% less than Intel-based alternatives, enhancing Azure's competitiveness against AWS and GCP. C. Scalability for Enterprise AI: ~"Venice" supports rack-scale AI deployments, enabling Azure to scale AI services for enterprise customers. For example, a 1,000-node cluster can process 700,000+ tokens per second, crucial for large-scale AI applications like personalized marketing and predictive analytics. ~This scalability is particularly important as Azure aims to capture the $100B+ AI opportunity by 2026, as stated by Microsoft CEO Satya Nadella. D. Reduction of Nvidia Dependency ~While Nvidia ( $NVDA) dominates AI accelerators, AMD's integrated EPYC-GPU solutions (MI450 with "Venice") offer a balanced approach, reducing Azure's reliance on Nvidia's high-cost GPUs. ~"Venice" enables hybrid inference models, where CPU-based inference handles 80% of workloads, and GPU acceleration is reserved for training and complex tasks, optimizing resource allocation. 3. Financial Implication: ~Revenue from Azure could reach $15-18B annually by 2026, part of a total revenue projection of $70-100B ~Profit margins could improve to 55-60%, boosting net income to $20-25B, supported by scale economies and reduced production costs. Intel could respond by giving more aggressive discounts, but this breakthrough has been a decade long of $AMD R&D, or rethinking chiplet design, a complete new approach. "Venice's" lead in AI inference and efficiency is challenging to match. Broader Industry: Other hyperscalers ( Amazon Web Services , GCP) and enterprises will follow Azure's lead, standardizing EPYC technology and pressuring Intel further. This could lead to a broader industry shift towards AMD, enhancing its ecosystem and bargaining power. Conclusion: The strategic adoption of AMD's 6th Generation EPYC "Venice" processors by Microsoft Azure in 2026 marks a pivotal moment in the evolution of cloud computing, particularly for AI inference capabilities. "Venice's" groundbreaking chiplet design, offering a 35x performance uplift for AI inference tasks, a 50% reduction in power consumption, and unparalleled scalability, positions Azure to leapfrog its competitors in the race for AI dominance. This technical superiority, combined with significant cost savings potentially $20-30B annually in operational expenses; aligns perfectly with Microsoft's ambitions to capture the $100B+ Revenue AI opportunity by 2026. The shift to 50% x86 market share for AMD within Azure is not merely a technical transition but a strategic realignment that redefines the competitive landscape. Historically, Microsoft's partnership with AMD has evolved from niche deployments to a core component of Azure's infrastructure, and "Venice" accelerates this trend. The 30-35% AMD EPYC share in 2025 is expected to double, driven by new VM families like C4D and H4D, which will dominate AI-intensive and HPC workloads. This migration is incentivized by "Venice's" efficiency gains, reducing dependency on Intel and Nvidia, and enhancing Azure's sustainability profile. Not Financial Advice!

Mike

141,018 Aufrufe • vor 10 Monaten

Jensen Huang just reframed the entire history of computing in two minutes. The argument is deceptively simple, but once you see it you can't unsee it. Every single piece of software ever built, every app, every website, every search engine, every platform operated on exactly the same fundamental principle. Someone creates content, it gets stored somewhere and when you ask for it, the system retrieves it. Google indexes the web and retrieves the right page, YouTube encodes your video and retrieves it when someone clicks, Amazon photographs every product in its catalog and retrieves the listing that matches your search. Every recommender system, every ad platform, every social feed, all of it, without exception, is a retrieval operation dressed up in a user interface and we called it the Information Age. But strip away the branding and what you had, for 30 consecutive years, was an extraordinarily sophisticated filing cabinet. The smartest engineers in the world spent their careers optimizing how fast you could put things in and pull things out. Generative AI doesn't just improve that system but rather replaces the entire premise of it. Instead of retrieving content that was pre-recorded by someone else, AI generates it from scratch, in real time, calibrated to your exact context, your specific intent, the precise ground truth of that moment. The same question asked twice gets two different answers, both tailored to what the system knows about you right now. There is no file being pulled or a pre-recorded version, the content is being synthesized on the fly from a compressed model of human knowledge, shaped to fit exactly what you need. The implications of this for the companies that built the retrieval era are profound and already starting to show. Google's click-through rates on organic search results have dropped 61% since AI Overviews rolled out, because users are getting answers directly instead of clicking through to files. Gartner projects traditional search engine query volume drops 25% by the end of 2026 as users migrate to generative interfaces. And yet this is exactly what Jensen predicted, in the old world, the computing bottleneck was storage and retrieval, you needed hard drives, bandwidth, and CDNs. In the new world, the bottleneck is computation, you need the raw processing power to generate tokens at scale, millions of times per second, for millions of simultaneous users. Inference computing demand has grown roughly ten thousand times in the last two years alone. That shift is precisely why Nvidia's revenue opportunity forecast just jumped from $500 billion through 2026 to $1 trillion through 2027. The retrieval era needed CPUs and storage and the generative era needs GPUs, token factories, and inference infrastructure at a scale never built before and Nvidia builds the engine underneath all of it. Jensen has been making this argument since 2024. Most people wrote it off as a chip salesman talking his book but two years later, it's the architecture of the entire industry.

Milk Road AI

17,911 Aufrufe • vor 3 Monaten

$AMD's heading to $5T MC LT| Lowest $/M tokens 🧵 The real reason why Institutions are FOMOing into AMD while other Semi stocks are underperforming ($NVDA $AVGO) Not Financial Advice! DYOR! Under Dr. Lisa Su’s leadership, AMD has transformed from a distant challenger into a formidable force in AI infrastructure, delivering the industry’s most compelling TCO story for high-volume inference. Her clear vision open ecosystems, aggressive annual roadmaps, rack-scale innovation, and relentless focus on tokens-per-dollar has positioned AMD’s Helios racks as the go-to solution for hyperscalers and AI natives struggling with exploding token costs, collapsing the cost down to $0.0003-$0.0005/M tokens. I will link various threads on this analysis to supply chain and wafer ratio if you are interested in understanding the full picture. In the last 3-4 months, explosive Agentic AI demand significantly increased Inference demand for Agentic AI models with 5-10 agents. If you are a listener of CNBC or Bloomberg, u should know enterprises and companies are complaining abt cost of token, and how it starts to spike up way too much to make sense. The fact that most data center today are run by $NVDA Chips, where the cost is way too high for Training or Inference. 1. Token cost Here are some quick comp, so u understand why $META OpenAI Anthropic $MSFT $AMZN Softbank $GOOGL and many more small to medium AI Natives are buying AMD CPUs and GPUs as much as they want, or pretty much AMD chips are sold out for the next 3-5 years. Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens 2. Why Hyperscalers and AI Natives Are Choosing AMD Token consumption (especially Agentic) is outpacing even NVIDIA’s efficiency gains, making diversification mandatory for economic viability. Massive deals reflect this reality like $META, OpenAI, $MSFT, Softbank, $AMZN, Oracle, LumaAI, G42... Dr. Lisa Su’s Vision in Action: Since taking the helm, Su has driven AMD’s turnaround with disciplined execution, annual GPU cadence (MI300 → MI350 → MI400), full-stack software (ROCm 7), open ecosystems (UALink, OCP designs), and customer-centric rack-scale solutions like Helios. Her emphasis on “tokens per dollar” and TCO has turned AMD into the pragmatic choice for sustainable AI scaling. Power/Energy Efficiency: ~Helios Rack-level is estimated at 120kW-140kW with 50% more HBM4 where Inference and Training cost matter ~Rubin Rack-Level is estimated at 160kW-230kw AMD Helios shines in owned TCO, memory density, and energy flexibility at hyperscale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B 3. Superior CPUs to pair with GPUs on massive scale 5-10-20GW Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. Conclusion: NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always-on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. Not Financial Advice! DYOR! Video source: Microsoft Build 2026

Mike

145,778 Aufrufe • vor 2 Monaten

Etched came out of stealth at $800M and by lunch X had NVIDIA in the ground We do this every few months. A chip launches, the deck says killer, the timeline holds a funeral, and NVIDIA closes green anyway Etched hardwires the transformer into silicon. That is where the speed comes from, nearly the whole die on one job instead of the ~30% a GPU uses. It is also the trap. The day that chip tapes out is the best it will ever be. You cannot patch it. You burned progress into a wafer and now pray the field stops moving NVIDIA made the opposite bet. Same board, faster every quarter in software. Dynamo is pulling more tokens per watt out of the same rack, on version 1.0 The depreciation risk the bears aimed at NVIDIA for two years does not live at NVIDIA. It lives here, on the chip built to bury it Etched is not a fraud. It is a niche tool priced like a general one, and $800M is not enough to run a frontier supply chain. The rest get bought on the next down cycle Bury the lead, not the leader. Full case with Jack Farley and Max Wiethe on MTS And special thanks for Baseten for the cool T Shirt! Chapters 00:00 Switching from bonds to semis 00:33 What Etched actually is 01:31 Faster and cheaper, but how much HBM 02:56 Maturing market, not an NVIDIA killer 05:01 $1B in contracts and a Taiwan factory 05:20 Why these startups all get absorbed 06:48 Tiered inference and the obsolescence trap 09:39 Etched vs TPUs and Trainium 12:17 Is the CUDA moat weakening 13:21 Co-design, squeezing every token per watt 14:37 NVIDIA is a software company that sells a chip 14:58 Who is NVIDIA's most dangerous competitor 16:23 The NVIDIA killers, ranked 18:19 A rich man's game 18:43 AMD's MI500 vs Rubin Ultra 20:19 The neocloud business decision 22:46 Lightning round, Rambus the toll on HBM 25:11 The CXL run-up on Astera, Marvell, Credo 26:22 Use AI less, go to the booth 28:10 EDA is not dead

Ben Pouladian

29,969 Aufrufe • vor 1 Monat

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 Aufrufe • vor 1 Monat

A transformer can learn not just the outcomes of dynamics, but the operator that executes the rules. To show this we trained a transformer on roughly 0.04% of a discrete rule space - 100 of 262,144 possible rules - and it learned to apply unseen rules from the same rule class. The model does not simply memorize specific rules. It learns the operator that maps a supplied rule plus an initial state, including unseen rules from this class, to the correct next state. This is relevant because it is a shift from “neural networks approximate dynamics” to “neural networks can learn to execute symbolic programs within a defined rule class”. The rule itself is supplied at inference time, as data, and the network has internalized how rules act, not which rules to apply. On previously unseen rules, the model achieves 98.5% perfect one-step forecasts and reconstructs governing rules with up to 96% functional accuracy. Two results make this hold up under scrutiny. First, inductive bias decay. As we scaled training rule diversity, the correlation between functional inference accuracy and distance-from-nearest-training-rule collapsed to R² = 0.00. At the largest tested training-rule diversity, the model’s performance on a new rule shows no measurable dependence on how similar that rule is to anything it was trained on. The bias toward training data (the thing we worry most about in compositional generalization claims) is something we can measure decaying, and we find that at scale it is gone. Second, an identifiability theory. We derive a closed-form expression for the number of rules consistent with a single observation. This reframes the inverse problem: failure to recover ground truth is not necessarily a model defect, but can be correct behavior when the data underdetermine the rule. The model is sampling the equivalence class; and identifiability is governed by coverage, not capacity. The methodological move underneath both results is amortization. Classical work on rule inference (e.g. the Santa Fe EVCA program, evolutionary search over CA rule space) was per-instance: search the rule space for each new system. We replace that with a single forward pass of a transformer trained across many instantiations of the rule class. That is what makes symbolic rule inference scalable as a research direction rather than a curiosity. We show that this works in a tightly constrained domain: binary, deterministic, local cellular automata on small grids. The locality-break experiment shows the model fails sharply when target systems violate its structural priors (which is itself a useful diagnostic, but it bounds the operator class). We don't yet know how this scales to multistate, higher-dimensional, or stochastic CA, or whether it transfers cleanly to non-CA systems whose coarse-grained dynamics admit local surrogates. The identifiability framework - what can be inferred from observation, given a hypothesis class - should transfer wherever finite local rules meet sparse data. The amortization argument transfers wherever per-instance symbolic search has been the bottleneck. Those are the pieces I expect to outlive the cellular automata setting. Led by Jaime Berkovich with Noah David, at LAMM@MIT. Out now in Advanced Science Advanced Portfolio (link to paper & code below).

Markus J. Buehler

39,019 Aufrufe • vor 3 Monaten

BREAKING: Mark Pincus (mark pincus) on How to Build Billion-Dollar Products "Your # 1 job as a founder is to be right." "F*ck scale." "Don't be a fake CEO." Mark founded Zynga in 2007, grew FarmVille + Words With Friends to 1B+ users in 4 years, & sold the company to Take-Two for $12.7B. Before that he was an early investor in Facebook & Twitter. Across 10 companies, he's spent 30 years learning how to build products people love. Now he's put it into "Life at the Speed of Play," aka the "Product Maker Bible" His goal: something you can reread in 10 years & still use, the way he rereads Peter Thiel's Zero to One. We went deep on the framework: - Your product instinct is right ~95% of the time, but your specific idea is wrong at least 75% of the time. The whole job is separating the two & k*lling your B+ idea to find the A. - Proven Better New: copy what already works, make it objectively better, then add one new bet. - 40,000 games launched in the App Store last year. 0% held a top 25 spot. Why products win on day 365 retention, not virality.. lessons from Nikita Bier & more - Why he's an AI maximalist who still calls consumer AI un-investable, & thinks today's $2-5T companies become $10-20T companies. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Mark Pincus, Chairman & Founder at Zynga (01:07) Why it took so long to write the book (01:52) Why the book starts with Elon (07:29) How frustration becomes market research (08:16) The gap between Airbnb and hotels (09:14) Why people don't want flight attendants on private jets (13:13) The goal behind Life at the Speed of Play (15:07) The instincts vs. ideas framework (21:01) Do instincts improve with experience? (22:49) A framework for building winning products (29:38) Becoming a student of yourself (31:16) Finding great ideas in overlooked products (34:39) Why AI makes building easier & winning harder (39:17) Cracking Zynga's record-setting retention (44:59) What Mark looks for in startups (50:42) Jeff Bezos' boldest decision (53:03) What happens when founders realize they're wrong (59:37) How Mark found talent others overlooked (1:06:58) Don't be a fake CEO (1:12:15) Inside Silicon Valley's early days (1:14:58) Why he funded Friendster & Napster as "experiments" (1:17:00) The 'think weekend' with Thiel, Zuckerberg, Reid Hoffman & Sean Parker (1:18:58) Launching Zynga when nobody believed him (1:21:42) Why Mark is an AI maximalist

Molly O’Shea

220,531 Aufrufe • vor 1 Monat

$AMD $5 Trillion MC Is Inevitable Long Term👑 This thread will focus more on Inference! 2026 EPYC "Venice" $TSM 2nm to save Large GW Scale Inference by 40% more than Prior Turin gen. Context: EPYC Turin achieves ~$0.001 per million tokens for batch inference vs $0.02-$0.12/ million tokens as I wrote the thread below. Venice is going to lower cost down to $0.0005-$0.0006/Million Tokens. OpenAI spent roughly $20B on Inference and Training, where 80-90% of that was for Inference per Analysts. AKA Renting Compute is Expensive AF! In this thread, I want to focus on why most analysts and investors are underestimating the role EPYC "Venice" and future Gen on overall Data center revenue. And $TSM ramping up 2nm supply early is a confirmation that AMD will be a major buyer long term. I will also link the thread the Gap between AMD Analysts & Reality and 2nm Ramp Thread so you have more comprehensive view of what I'm writing here. Before I go into detail this is my 2026 Projection: AI GPUs: $35-$50B EPYC Data Center: $15B-$17B Client Segment: $12-$13B Gaming: $6B Embedded: $4B-$5B Total Revenue $70-$100B Non-GAAP net income $18B-$25B Non-GAAP EPS $10.97-$15.40 Foward P/E 55x-70x= $603-$1,078 AMD's Analysts are projecting $0 Revenue for MI450 and sluggish EPYC Growth. Meaning, all analysts are either full of 💩 or Sexist, you decide! Analysts are also projecting 0% growth on AMD "Secret Weapon" Chip as $MSFT said we are at significant Windows refresh and upgrade cycle. Do you think TSMC would allocate more 2nm supply to $AMD at $0 MI450 revenue and sluggish EPYC? 1. EPYC is going to be the leader in lowest Inference! Current Turin cost saving is 95% vs $NVDA or 98-99% on Inference cost when you factor in renting Inference compute from Amazon Web Services, Microsoft Azure, or $NVDA Neocloud pets. TSMC claimed: 10-15% higher performance at iso-power, 25-30% lower power at iso-speed, and ~15% higher transistor density compared to 3nm. This reduces operational expenses (energy, cooling) while increasing throughput per chip. EPYC Turin achieves ~$0.001 per million tokens for batch inference (via vLLM on models like Llama 3 70B), driven by high core counts and low hardware costs. EPYC Venice offers ~1.7x overall performance and up to 70% more compute capability per core, with up to 256 cores (512 threads). Enhanced vector/AI instructions and open-source firmware (openSIL) optimize for inference workloads. AMD Incorporates AI Engines (now part of AMD's XDNA) for on-chip acceleration, improving efficiency for low-latency and edge inference. This reduces reliance on discrete GPUs, lowering system complexity and TCO. Venice SKUs are projected at $3,000-$15,000 ($5,000 for 256-core flagship), far below NVIDIA Rubin ($50,000-$90,000) or AMD's own MI450 GPUs ($40,000-$50,000). High memory bandwidth (up to 1.6 TB/s) supports efficient batch inference. Venice is designed exactly for Large customers that want to lower Inference Cost and MI450 Helios is for Customers that want Training at lowest TCO, TDP as well as lower Upfront 1GW scale(Full build $35-$40B vs $NVDA $55B-$80B). 2. Real World Example: OpenAI's 2025 inference spend reached ~$20B, escalating to even higher total compute rental (mostly inference) amid token volume growth(from video generating). By 2026, with usage doubling (consistent with industry trends: token demand grows 2-5x YoY), assume OpenAI processes ~1,800 billion million-tokens annually $NVDA Blackwell at $0.02-$0.12 is $36B(most optimized) Rubin is projected to be at $0.01/million tokens or $18B annual Inference Cost vs $AMD Venice $0.0005/million tokens or $0.9B annual Inference Cost => Massive saving for OpenAI or anyone that are paying 80-90% Annual Bill for Inference compute. In short, it is unsustainable to pay this much rent vs owning for all current AI players for the medium to long term. Rubin excels in low-latency decode (if Groq integration from $20B deal in 2027-2028), but Venice dominates batch (80% of inference by 2030). Actual savings depend on deployment scale (OpenAI's 6GW AMD plans), electricity rates, and software maturity. If Rubin only hits $0.03, savings swell to $53.1B vs. $17.1B. 3. Will running Inference on Venice and future Gen slow down response generation in 2026 and beyond? Human perception of "fast enough" for chat, agents, search augmentation, summarization, coding assistance is roughly Meaning, EPYC may generate $100B a year on data center revenue, Hence $MSFT $AMZN $META $GOOGL OpenAI xAI and 42+ Countries are leaning AMD for Inference, because the cost saving is MASSIVE! 4. Regular users (you, me, people using ChatGPT, Claude, Gemini, Grok, Perplexity...) are extremely unlikely to notice any slowdown and in many cases might even experience slightly faster or more consistent response times if the industry heavily shifts toward AMD EPYC for inference. What actually happens when companies save massively on inference? When OpenAI , Anthropic , Gemini , Grok Meta .... save billions on the batch/enterprise/RAG layer using EPYC Venice, they typically do one or more of these things with the savings, none of which make your chat slower but enhancing their bottom line(Profit) ~Keep prices the same → make more profit ~Lower subscription prices / increase free tier limits ~Train bigger & better models more frequently ~Offer longer context windows ~Add more reasoning steps / tool calls / agents per query ~Improve multimodal capabilities ~Build more data centers / reduce throttling during peaks In practice the consumer experience usually gets better, not worse, when inference becomes dramatically cheaper. Prime example is $META leaning AMD heavily or currently AMD largest customer. or Grok 2 to Grok 3 heavily used AMD for Inference saving. And most Grok Users reported Groke responses snappier, not slower. 5. What does this mean for potential Revenue? Noted that TSMC is massively ramping 2nm supply for $AMD both MI450 and EPYC. EPYC Conservative projection: FY2025: $10.5B(best Est) FY2026: $16B FY2027: $29B FY2028: $49B FY2029: $75B FY2030: $100B Large customers: $META OpenAI $MSFT $AMZN $GOOGL xAI (Apple?) Smaller customer: $DELL $HPE $SMCI and 42+ other countries. The roadmap to $5 Trillion is very much inevitable as Inference Cost from Renting or owning $NVDA are too high, but $NVDA will still dominate Training market share, where MI families are likely to take 15-20% market share, but the TAM is also expanding Rapidly. Most Institutions are projecting $2-$3Trillion TAM by 2030. $NVDA said $4 Trillion. Dr. Lisa Su said $1 Trillion+ by 2030. So you decide on how much TAM. If you enjoy this kind of analysis, Slap the Like/Repost and Bookmark to please the X Algo as it is Free.99! If you want to support my work further, consider subscribe to see more in-depth analysis! Alright, that is it. Not Financial Advice!

Mike

102,223 Aufrufe • vor 7 Monaten

Excited to release a new repo: abcGPT! It can be hard to "dial in" the voice you want from an LLM, because an LLM is a tangled superposition of millions of voices from millions of different authors around the world. Instead, frontier LLMs tend to give that slop-ish / generic / corporate tone that's hard to avoid, even with aggressive prompting and an informative context window. Lately I've been experimenting with some ideas on the fringes of attribution/unlearning, trying to make it so an AI user can "dial in" the specific voice/style/sources they want to use in a way that's more rigorous than prompting/context-engineering. and I'm starting to get pretty good results. the model below uses the following technique: - Take nanoGPT as written by Andrej Karpathy - Assign each neuron a random "specialty score" m between 0 and 1, sampled from a U-shape so most neurons land near 0 or near 1 with some in the middle. - Freeze this "m" for the lifetime of the network (it's the neuron's permanent corpus assignment) - Extend the forward() code with an α parameter, a kind of vibe-fader from 0 to 1. Think of each neuron's m as its position on that same slider. The slider acts like a spotlight: it lights up neurons whose m is near its current position, and silences those far away. Slide all the way to 0, and only TinyStories specialists fire. Slide all the way to 1, and only Shakespeare specialists fire. - Train this new nanoGPT on two datasets (in this case, TinyStories and Shakespeare) - During training, sample α from Beta(0.5, 0.5) AND draw the corpus from Bernoulli(α), so a Shakespeare batch tends to come with a high-α (Shakespeare-favoring) gate, and a TinyStories batch tends to come with low-α. - train until golden brown 🧑‍🍳 Perhaps surprisingly... it works! ¯\_(ツ)_/¯ The neurons we pre-assigned to Shakespeare learn to behave as Shakespeare specialists. the neurons we pre-assigned to TinyStories become children's-story specialists. the halfsies learn to bridge between them. After training, you can play with the kindof... vibe dial... you can "dial in" the voice you want during inference, by choosing whether to lean on Shakespeare or TinyStories neurons more or less. 📀💿 When you fully dial in Shakespeare neurons, the model only outputs tokens which look like Shakespeare, and when you fully dial in TinyStories, the model only outputs tokens which look like children's stories, and... (honestly this was the hard part)... everywhere inbetween! In a way, it's partitioning statistical signal into fuzzy segments, and then the end user can choose which pre-training data sources they want to lean upon for generation... and how much. My goal was to get a version of this working at scale, with clear intuition for why it works, and I'd like to explore ways to scale up this effect to large numbers of sources and larger models, and study the interplay between individuality/generality as scale increases. Link to repo and a detailed walkthrough of the abcGPT methodology in the reply.

⿻ Andrew Trask

133,194 Aufrufe • vor 2 Monaten

Cerebras just IPO’d and the stock already ran up over 100% (Save this). For the entire 70 year history of the semiconductor industry, every company on earth has followed the same process. You take a dinner plate sized silicon wafer, put hundreds of tiny chips onto it, and dice it up like a pizza. Nvidia does it this way, AMD does it this way, Intel has done it this way for six decades and everyone who tried to break that convention failed. Until Cerebras asked the most annoyingly obvious question in the industry’s history, what if you just didn’t cut it? The result is the Wafer Scale Engine, a single chip 56 times larger than Nvidia’s H100 and it fundamentally changes the physics of how AI inference works. The reason this matters is not the size, it’s the bandwidth. Every time an AI model generates a single word, it has to reach into memory, pull weights, multiply them together, and produce a prediction and when you’re running millions of concurrent sessions at once, the bottleneck is not raw processing power but how fast data moves between memory and compute. Nvidia’s H100 moves data at roughly 3 terabytes per second, while Cerebras’ WSE-3 moves data at 21 petabytes per second, roughly 7,000 times faster because memory and compute live on the same enormous piece of silicon and data barely has to travel at all. That gap is exactly why OpenAI went from 150 tokens per second on traditional GPUs to 2,000 tokens per second on Cerebras hardware, and why AWS integrated Cerebras into Bedrock to deliver roughly 5x more inference capacity in the same physical footprint. The macro setup is making the trade even more urgent. South Korea DRAM export prices recently jumped 35%, flash memory surged 47%, and SSD pricing spiked nearly 140% and every single one of those increases hits Nvidia-based infrastructure directly, because the H100 requires 80GB of the most expensive, most contested memory in the AI supply chain. Cerebras’ WSE-3 uses zero external HBM memory, baking 44GB of SRAM directly into the wafer itself which means as memory pricing goes parabolic, every CFO evaluating AI infrastructure is suddenly looking much more seriously at the architecture that sidesteps that cost entirely. The demand is already showing up in the backlog. Cerebras ended 2025 with $24.6 billion in remaining performance obligations for a company doing just over $500 million in annual revenue, that is a number that implies years of contracted growth already sitting on the books. The IPO was 20x oversubscribed, the price range was raised twice before listing, and shares opened 89% above their listing price on a $5.55 billion raise that made it the largest semiconductor IPO in history. The risks are real and worth naming. 86% of 2025 revenue came from two entities with UAE ties, U.S. revenue actually fell 34% to $187 million, and the $20 billion OpenAI contract is conditional, if Cerebras misses delivery milestones, OpenAI can terminate and trigger repayment demands on a $1 billion loan facility. And yet the market is valuing Cerebras at roughly 91x trailing revenue, richer than Nvidia, AMD, and Arm combined. What investors are betting on is not that Cerebras beats Nvidia, it is that the inference supercycle is large enough to support an entirely different architecture optimized for a different workload, and that $24.6 billion in contracted backlog converts to diversified revenue before the market starts asking harder questions. CEO Andrew Feldman said this took a decade of late nights to get right, everyone who tried to copy it failed and given that the entire inference economy is now running through exactly the bottleneck Cerebras was built to eliminate, the market is starting to believe him.

Milk Road AI

30,441 Aufrufe • vor 3 Monaten