Загрузка видео...

Не удалось загрузить видео

На главную

Vik and Val Bercovici (WEKA) map where AI inference memory is headed. Every 100x cut in KV cache gets swallowed by ~10,000x more usage, so demand climbs. - NVLink beats the board: 128 lanes vs 32 PCIe - WEKA serves NAND-backed storage faster than DRAM over network - DeepSeek's...

59,992 просмотров • 18 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

The Cost of Intelligence is Heading to Zero | Hyperspace P2P Distributed Cache We present to you our breakthrough cross-domain work across AI, distributed systems, cryptography, game theory to solve the primary structural inefficiency at the heart of AI infrastructure: most inference is redundant. Google has reported that only 15% of daily searches are truly novel. The rest are repeats or close variants. LLM inference inherits this same power-law distribution. Enterprise chatbots see 70-80% of queries fall into a handful of intent categories. System prompts are identical across 100% of requests within an application. The KV attention state for "You are a helpful assistant" has been computed billions of times, on millions of GPUs, identically. And yet every AI lab, every startup, every self-hosted deployment - computes and caches these results independently. There is no shared layer. No global memory. Every provider pays the full compute cost for every query, even when the answer already exists somewhere in the network. This is the problem Hyperspace solves where distributed cache operates at three levels, each catching a different class of redundancy: 1. Response cache Same prompt, same model, same parameters - instant cached response from any node in the network. SHA-256 hash lookup via DHT, with cryptographic cache proofs linking every response to its original inference execution. No trust required. Fetchers re-announce as providers, so popular responses replicate naturally across more nodes. 2. KV prefix cache Same system prompt tokens - skip the most expensive part of inference entirely. Prefill (computing Key-Value attention states) is deterministic: same model plus same tokens always produces identical KV state. The network caches these states using erasure coding and distributes them via the routing network. New questions that share a common prefix resume generation from cached state instead of recomputing from scratch. 3. Routing to cached nodes Instead of transferring KV state across the network for every request, Hyperspace routes the request to the node that already has the state loaded in VRAM. The request goes to the cache, not the cache to the request. Together, these three layers mean that 70-90% of inference requests at network scale never require full GPU computation. This work doesn't exist in isolation. It builds on research from across the industry: SGLang's RadixAttention demonstrated that automatic prefix sharing can yield up to 5x speedup on structured LLM workloads. Moonshot AI's Mooncake built an entire KV-cache-centric disaggregated architecture for production serving at Kimi. Anthropic, OpenAI, and Google all launched prompt caching products in 2024 - priced at 50-90% discounts - because system prompt reuse is so pervasive that it changes the economics of inference. What all of these systems share is a common limitation: they operate within a single organization's infrastructure. SGLang caches prefixes within one server. Mooncake disaggregates KV cache within one datacenter. Anthropic's prompt caching works within one API provider's fleet. None of them can share cached state across organizational boundaries. Hyperspace removes this boundary. The cache is global. A response computed by a node in Tokyo is immediately available to a node in Berlin. A KV prefix state generated for Qwen-32B on one machine is verifiable and reusable by any other machine running the same model. The routing network provides the delivery guarantees, the erasure coding provides the redundancy, and the cache proofs provide the trust. What this means for the cost of intelligence Big AI labs scale linearly: twice the users means twice the GPU spend. Every query is a cost center. Their internal caching helps, but it's siloed - Lab A's cache can't serve Lab B's users, and neither can serve a self-hosted Llama deployment. Hyperspace scales sub-linearly. Every new node that joins the network adds to the global cache. Every inference result enriches the cache for all future requests. The cache hit rate rises with network size because query distributions follow a power law - the most common questions are asked exponentially more often than rare ones. The implication is simple: as the network grows, the effective cost per inference drops. Not linearly. Logarithmically. At 10 million nodes, we estimate 75-90% of all inference requests can be served from cache, eliminating 400,000+ MWh of energy consumption per year and avoiding over 200,000 tons of CO2 emissions. The first person to ask a question pays the compute cost. Everyone after them gets the answer for free, with cryptographic proof that it's authentic. Training is competitive. Inference is shared Open-weight models are converging on quality with closed models. Labs will continue to differentiate on training - data curation, architecture innovation, RLHF tuning. That's where the real intellectual property lives. But inference is a commodity. Two copies of Qwen-32B running the same prompt produce the same KV state and the same response, byte for byte, regardless of whose GPU runs the matrix multiplication. There is no moat in multiplying matrices. The moat is in training the weights. A global distributed cache makes this separation explicit. It doesn't matter who trained the model. Once the weights are open, the inference cost approaches zero at scale - because the network remembers every answer and can prove it's correct. No lab, no matter how well-funded, can match this. They cannot share caches across competitors. They scale linearly. The network scales logarithmically. The marginal cost of intelligence approaches zero. That's the endgame.

Varun

37,362 просмотров • 4 месяцев назад

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,170 просмотров • 7 месяцев назад

I asked Dan Martell to walk me through every level of making money with AI. He gave me the most simple, practical advice I've ever heard on this subject. Level 1 - Making $0 - $100k Level 2 - Making $1m - $10m Level 3 - Building a $10m++ enterprise. 0:00 Only 5% of the World Has Ever Paid for AI 0:46 The Easiest Thing to Sell With AI Right Now 1:56 The Marcus and Sophie Framework 4:24 Theory of Constraints (Right Problem to Solve) 5:33 What Is the Number One Business Constraint 7:13 How to Leave Your Job and Go All In 8:27 Business Is Simple Find a Problem and Solve It 9:08 Stop Getting Ready to Get Ready 9:33 The Sarah Story One Text and $10K 9:53 Pull Up Your Phone and Message Your Contacts 11:05 Dan's Son Gets His First Client at $800/Month 12:41 Best Employee vs. Best Employer 13:59 What Other Services Can You Sell With AI 14:44 Sales Is Not Talking It's Asking 17:01 What to Do When You Hate Your Business 18:40 Pain and Pleasure Are the Only Two Motivators 19:13 They Haven't Made It a Must Yet 20:29 Make It a Must Not a Nice to Have 21:06 The Jen Story and the Gasping Moment 22:17 How to Find Your First 10 to 15 Clients 28:38 The Personal Brand Play 33:06 Vision Is What AI Cannot Do 34:55 Hard for Computers Easy for Humans 36:13 Level 2 Making Your First Million With AI 37:18 The Replacement Ladder Framework 37:39 Admin First Then Delivery Then Marketing 39:09 Why Marketing Is the Biggest AI Category 39:32 Why You Should Keep Sales for Yourself 40:00 Level 5 Leadership and AI Agents 41:41 What a Fully AI Systems Business Looks Like 43:13 The Gym Owner With Three Locations 46:16 Shutting Down the Company for Two Days 46:37 Teaching the Whole Team to Code in Claude 49:28 Wayne the 62 Year Old Who Made $12K a Month 52:38 I Only Share What Actually Works 53:21 Whisper Flow and Talking to Your AI 56:41 Claude Chat Claude Coworker and Claude Code 57:57 The Claude Browser Extension 58:49 Claude Code Is Not Just for Developers 1:00:06 How to Migrate Your AI Memory Across Tools 1:01:08 Level 3 $1M to $10M and the Brand Play 1:02:05 Nobody Buys AI They Buy Trust 1:03:25 Brand Is Association and Association Is Trust 1:05:12 A Million Followers Is $10M in Activated Revenue 1:07:03 How to Keep AI From Becoming Slop 1:07:42 Human in the Loop 1:08:16 The 10 80 10 Rule and Why AI Is Now the 80 1:10:01 The Team FIRED Themselves 1:11:45 Dan's Free AI Curriculum for Your Team

Grant

166,363 просмотров • 1 месяц назад

🔥Web3 Beyond the hype! In this exclusive conversation, hardik, Founder & CEO of The Crypto Times, sits down with Max Rebol, Founding Partner at Harbour Industrial Capital, to explore the future of Polkadot, why focused VC strategies win, and how DePIN and real-world use cases will power the next phase of blockchain adoption. From scaling to a billion users to data monetization and sustainable crypto projects, this one’s packed with insights for Web3 builders and investors. ------------------------------------------------ 00:00 – Introduction to Max Rebol & Harbour Industrial Capital 00:28 – Max Rebol’s Background & Launching a Polkadot-Focused VC Fund 01:47 – Web3 Foundation's Role in Fund Two 02:05 – Fund Size, Track Record, and Investment Strategy 03:05 – Top Investment: Peaq Network & Depin Use Case 04:14 – Web3 for Real-World Users vs. Crypto Natives 04:40 – Polkadot in Gaming: Mythical Games & FIFA Rivals 05:46 – Blockchain Gameplay: Trading Virtual Players 06:21 – Polkadot vs. Ethereum as Backend Infrastructure 07:32 – Cost of Deploying Parachains & Entry Barriers 08:29 – Agile Core Time: Polkadot’s New Scaling Model 10:01 – Scaling from 10 to a Billion Users with Agile Core Time 10:59 – Ethereum’s L2 Challenges and Polkadot's Efficiency 12:30 – L2s as Free Riders? Ethereum’s Sustainability Problem 13:31 – Protocol Revenue: The Crypto Chain Survival Metric 15:06 – Stablecoin Transfers & Nova Wallet Fee Benefits 16:34 – Why Are People Still Using Tron? 17:27 – Polkadot Fee Flexibility: Use USDC/USDT to Pay Fees 18:04 – Protocol Revenue Models for Developers 19:29 – What Gives DOT Token Real Value 20:34 – Polkadot Compared to AWS Credits 22:26 – Polkadot’s TPS Benchmark: 143,000 TPS 24:22 – Polkadot JAM and the Future of Multicore Blockchain 25:46 – Solana Centralization & Hardware Limitations 27:15 – Meme Coin Culture vs. Blockchain Ethics 28:29 – Why Polkadot Won’t Compete in the Meme Coin Space 30:41 – Crypto Narratives from 2017 to 2024 32:56 – Real World Use Cases as the Next Big Narrative 34:23 – Tokenization of Every Asset: The Long-Term Vision 35:42 – Identity & Privacy Use Cases on Polkadot 36:44 – Moving Beyond KYC with Privacy-Preserving Tech 38:14 – Why Traditional Investors Would Use Tokenized Platforms 40:00 – Fractional Real Estate & Global Financial Inclusion 41:48 – Interoperability & Modular Compliance on Polkadot 43:13 – Smart Contracts Enforcing Local Regulations 44:55 – Permissionless Execution vs. Centralized Chains 47:08 – Polkadot NPoS vs. PoS: More Efficient Validator Use 49:09 – Governance on Polkadot: The World’s Largest DAO 50:52 – How Polkadot Treasury Proposals Work (Subsquareio) 51:53 – What Is Polkadot JAM? (Join-Accumulate Machine) 53:47 – Multicore Architecture: Like a Cloud Server for Web3 55:28 – Coherence Across Cores: JAM’s Secret Weapon 56:17 – Doom on the Blockchain: JAM Demonstration 57:53 – Why Not All Games Should Run Fully On-Chain 1:00:02 – Fun Gameplay First, Crypto Utility Second 1:01:55 – Owning & Trading Game Assets on Blockchain 1:03:22 – Mythical Games: Call of Duty Devs Build for Web3 1:05:06 – Raising Capital for Fund 2: Focus on Institutional Investors 1:06:34 – Family Offices’ Strategic Web3 Bets: Why They Choose Polkadot 1:08:32 – Why Focused VC Funds Beat Diversified Strategies 1:10:25 – Capital Deployment Strategy: Revenue-Driven Projects 1:11:45 – Depin & Data Monetization: The Next Web3 Goldmine 1:13:17 – Example: Cars Selling Anonymous Driving Data 1:14:49 – Exchanges, Bridges & Revenue-Focused Infrastructure 1:15:56 – Investing in HyperBridge: Security & Bridge Revenue 1:16:34 – The Future of VC: Sustainable Projects Over Hype 1:16:52 – Closing Words & Thanks from the Interviewer

The Crypto Times

21,911 просмотров • 1 год назад

Research suggests that up to 40% of cancer cases could be prevented through lifestyle changes. The evidence is now overwhelming: exercise is not just supportive—it’s a therapeutic intervention that recalibrates tumor biology, enhances treatment tolerance, and improves survival outcomes. Today’s interview features Dr. Kerry Courneya. With over 600 peer-reviewed studies, he is one of the most influential figures in exercise oncology. Even if you aren't someone who has personally experienced cancer in one form or another, you need to watch this episode. Episode 99 is Available now on X, YouTube, Spotify, and Apple Podcasts. Chapters: 0:00 - Introduction 1:47 - Why exercise should be effortful 2:33 - How to meaningfully reduce risk of cancer 6:22 - What type of exercise is best? 7:59 - How exercise reduces risk—even for smokers and the obese 10:48 - Weekend-only exercise 13:49 - 150 vs. 300 minutes per week (more is better—up to a point) 16:03 - Why pre-diagnosis exercise matters 19:09 - Why resilience to cancer treatment starts with exercise 21:01 - Why low muscle mass drives cancer death 23:58 - Why BMI fails to measure true obesity 27:51 - Why daily activity isn't enough (structured exercise matters) 29:34 - Breaking up sedentary time—do ‘exercise snacks’ help? 31:50 - Supplements vs. exercise 32:32 - Where exercise fits with chemo and immunotherapy 35:30 - Why rest is not the best medicine 41:20 - Aerobic vs. resistance 42:13 - How weight training improves 'chemo completion' 44:41 - Why exercise creates vulnerability in cancer cells (limitations do apply) 47:09 - Why exercise might be crucial for tumor elimination 53:03 - Why cardio may be better at clearing tumor cells 56:18 - When cancer spreads quickly—and when it doesn't 57:43 - Why liquid biopsies may prevent over-treatment 1:02:56 - Exercise-sensitive vs. exercise-resistant cancers 1:06:06 - Prostate cancer therapy—why strength training matters 1:08:10 - When exercise is the only therapy—does it work? 1:09:26 - Why HIIT reduces PSA in prostate cancer 1:11:40 - Avoiding overtreatment—can exercise buy you time? 1:12:00 - Why high-intensity exercise boosts anti-cancer biology 1:13:11 - Turning a diagnosis into a wake-up call 1:16:11 - Why oncologists are rethinking exercise 1:18:50 - Why exercise eases anxiety about cancer—proven psychological benefits 1:25:00 - Before, during, and after treatment 1:27:02 - Why exercise is unique among cancer therapies 1:28:16 - Why cancer patients stop exercising—the risky mistake almost everyone makes 1:30:41 - How to get sedentary cancer patients exercising (realistically) 1:33:15 - The $1 million case for including exercise 1:34:56 - Why recurrence trials haven't convinced doctors—yet 1:37:36 - The bottom-line message 1:37:55 - The myth of a cancer panacea (exercise included) 1:44:07 - What's the best $50 investment for staying active? 1:44:40 - Only 15 minutes per day—what’s the best anti-cancer exercise?

Dr. Rhonda Patrick

219,784 просмотров • 1 год назад

$AMD $5 Trillion is Inevitable LT| Agentic AI🧵 Agentic AI is the new $5 Trillion TAM 🚨🚨🚨 This thead will do Comp with $INTC and how to quantify this massive Agentic AI demand spike, and forcing Jensen to rush a CPU design. Global Agentic AI Market size is estimated to be $3-$5Trillion TAM by 2030(McKinsey) Quantifying the demand from agentic AI for AMD involves assessing the broader market growth for agentic systems, their unique computational requirements (particularly for CPUs in orchestration and reasoning tasks), and AMD's positioning very well through products like EPYC processors and partnerships. AMD EPYC Venice is the most superior choice in 2026-2027 for most Agentic AI workloads Agentic AI refers to autonomous AI agents that perform multi-step tasks, involving sequential logic, tool integration, and decision-making workloads that heavily rely on CPUs for handling orchestration, memory management, and context switching, rather than just GPU-parallelized training or batch inference. Agentic AI is often cited as 40-100x more "hungry" than traditional AI due to its continuous, 24/7 operation and complex workflows. This stems from factors like chain-of-thought reasoning (multiple LLM calls per query), API/tool interactions, memory management, and orchestration loops, which can generate 10-100x more tokens and require real-time responsiveness. For example, a single agentic query might trigger 5-20 model inferences, making it 10-20x more compute-intensive than simple chatbots, and the always-on nature compounds this to 40-100x overall. Nvidia's CEO has highlighted this as driving "easily 100x more computation" for inference in agentic/reasoning setups. AMD's EPYC Venice (6th Gen EPYC, codenamed "Venice") and Intel's Xeon 7 Diamond Rapids represent the pinnacle of server CPU technology in 2026, both targeting high-performance data center workloads like AI inference, agentic AI orchestration, cloud computing, and HPC. Venice builds on AMD's Zen 6 architecture, emphasizing core density and efficiency, while Diamond Rapids leverages Intel's Panther Cove P-cores for balanced performance. Both chips adopt similar advancements like 16-channel DDR5 memory and PCIe Gen 6, but differ in core counts, process nodes, and overall design philosophy. Intel has faced acute supply constraints across its Xeon lineup, including legacy nodes (Intel 7/3) and the ramping 18A process for next-gen parts. Intel shortage is expected with lead times up to 6 months or longer. 1. AMD EPYC Venice vs Intel Xeon 7 Diamond Rapids Architecture AMD: Zen 6 chiplet design with 8 CCDs and dual IODs Intel: Panther Cove P-cores; multi-die architecture with 4 compute tiles Core/Thread Count AMD: Up to 256 cores / 512 threads (Zen 6c variant) Intel: Up to 192 cores / 192 threads Process Node AMD: TSMC N2 (2nm) Intel: Intel 18A (1.8nm-class); in-house fab Memory Support AMD: 16-channel DDR5; up to 1.6 TB/s bandwidth. Intel: 16-channel DDR5 ; up to 1.6 TB/s bandwidth I/O and Connectivity AMD: PCIe Gen 6 (up to 128 lanes); twice the CPU-to-GPU bandwidth Intel: PCIe Gen 6 (up to 128 lanes); LGA 9324 socket Power (TDP) AMD: Starting 400-500W, potentially lower due to efficiency gains from TSMC 2nm Intel: Starting 400-500W, as it targets competitive efficiency Performance Projections AMD: Up to 70% uplift vs. 5th Gen Turin (1.7x in multi-threaded/AI tasks) Intel: ~40% faster than Granite Rapids (Xeon 6, 128-core). Lags AMD in per-core perf and 40-50% behind Venice core-for-core comp Target Workloads AMD: AI inference/orchestration, HPC, cloud virtualization. Partnerships Intel: Hyperscale AI, general enterprise. Custom silicon Pricing: AMD: estimated $10k-$20k for top SKUs Intel: estimated $8-$18k Availability: AMD: Significant Ramp H2 2026 due to higher allocation from TSMC Intel: H1-H2 2026 delayed, but trying to catch up Overall: ~Venice's 256 cores provide a 33% edge over Diamond Rapids' 192, making it superior for massively parallel tasks like AI training/inference or virtualization ~TSMC's N2 vs. Intel 18A debates rage on which is "better," but AMD's mature chiplet approach yields better density ( 32 cores/CCD vs. Intel's 48/tile). Venice's redesign reduces latency, aiding agentic AI where CPUs handle orchestration ~ Early projections show Venice widening AMD's lead matching or exceeding Diamond Rapids' perf with fewer watts in multi-threaded benchmarks. Intel's no-SMT design (to prioritize AI) handicaps it vs. AMD's 512 threads, though Clearwater Forest (E-core) could compete in density-focused niches. ~Power & Cooling: Both push above 400-500W, demanding liquid cooling. ~AMD been taking market share now above 40%. AMD EPYC Venice emerges as the superior choice in 2026 for most server workloads. Its higher core/thread count (256/512 vs. 192/192), stronger per-core performance, and architecture optimized for AI-driven tasks (agentic orchestration with GPU integration) provide decisive advantages in throughput, scalability, and efficiency. Projections indicate Venice delivering 1.7x the performance of prior gens while widening the gap over Intel ( 40-70% leads in multi-threaded benchmarks). AMD's fabless model with TSMC ensures reliable scaling, and its ecosystem ( open ROCm) appeals to AI adopters. Intel's Diamond Rapids is competitive in single-threaded enterprise apps and custom hyperscale ( NVLink), with potential fab advantages for supply/security. However, without SMT and lower density, it falls short in core-for-core battles—exposing Intel to another generation of AMD dominance unless 18A yields surprise efficiency gains. For data centers prioritizing raw compute ( AI, HPC), Venice wins; for Intel-centric ecosystems or specialized I/O, Diamond Rapids holds ground. Real benchmarks post-launch will confirm, but logic points to AMD pulling ahead. 2. Market size , Potential Revenue and Supply Global Agentic AI market size is projected to be $3-$5 Trillion by 2030 according to McKinsey, where consensus points to 40-50% CAGR driven by small to large enterprise demand. I also wrote a full thread on how and why Agentic AI is so explosive that AMD will blow all anlaysts estimate for subscribers. Link below if you are interested. AMD's data center segment hit a record $5.4B in Q4 2025 (up 39% YoY), with EPYC shipments ramping due to agentic demand. With 2GW of deployment in H2 2026, AMD AI data center revenue has $40-$50B+ at the lowest or most conservative projection; or Total Revenue in the $77-$94B For FY2026. However, Agentic AI massive demand spike could send EPYC revenue 3x to 4x in the next few years, potentially surpassing MI series GPU demand as enterprises prioritize CPU-dense Rack setups. This is pushing $NVDA Jensen to rush a CPU design and acquired Groq, a new CPU player due to this massive TAM. Noted that this is just popping just in weeks, highlighting we are just so early in this AI Supercycle and the pace of adoption is insane, and clearly productivity will skyrocket. Why? Because Agentic AI is 24/7 Smart AI agent working for you or your businesses is a mad compelling, and it is estimated to be 40-100x more Inference Hugnry! Many experts already said it is impossible to project this kind of Inference Demand. AI CapEx is expected to ramp up even more in 2027-2028-2029 and 2030 as Global Agentic AI is going to scale to $3-$5 Trillion TAM by 2030. The nature of Agentic is driving higher CPU/GPU ratio, with CPUs handling 50-90% of Agentic workflows. For example, The current Helios Rack: 18 compute trays per rack with 72 GPUs + 18 CPUs. The beauty of this $META and $AMD long term partnership is, that it is absolutely flexible to adjust racks to higher CPU rato or equal to service different needs. Helios rack can be easily swap to 2 GPUs 2CPUs or even CPUs only trays for dedicated orchestration/head nodes. You see, the beauty of this open rack-scale is flexibility and evolvability. If Agentic AI demand pushes much higher, AMD should be able to adjust variant trays without abandoning Heilos Rack. We can't talk just about massive Agentic AI demand without talking about the Supply side or TSMC. TSMC, AMD's primary foundry for advanced nodes ( Zen 6/Venice on N2/2nm), is addressing AI-driven shortages through massive expansions. TSMC accelerates fab construction with up to 10 facilities targeted for 2026. TSMC is accelerating its domestic manufacturing expansion, with industry sources indicating that as many as ten fabs could be under construction or preparing to begin operations across Taiwan’s major science parks. TSMC Capex: $52-56B in 2026 (up 37% YoY), with $45B already approved for new/upgraded capacities. 70-80% for advanced processes (2nm/A16), 10-20% for packaging (CoWoS quadrupling to 120-140K wafers/month by late 2026). In addition, Taiwanese companies (led by TSMC) commit to at least $250B in direct investments in US-based advanced semiconductor, AI, and energy production/innovation capacity.Taiwan provides $250B in government credit guarantees to facilitate additional investments and build a full US semiconductor ecosystem (including industrial parks). TSMC completed a second land purchase in Arizona (January 2026) for gigafab scaling, with an additional $100B+ (potentially four more modules) to further expand and qualify for tariff exemptions. AMD with secured 12GW from OpenAI and $META and massive Agentic AI will mean higher priority acess to 20-30% more wafers on TSMC advanced nodes, as TSMC has multi-year agreements with AMD for AI chips. Dr. C. C. Wei, CEO of TSMC quote: "I spend a lot of time in the last three or four months talking to my customer and then customers. Customer. I want to make sure that my customers demand are real. I talk to those cloud service providers, all of them. Their answer is. I'm quite satisfied with their answer. Actually they show me the evidence that the AI really help their business. So they grow their business successfully and he or she in their financial return. So I also double check their financial status. They are very rich." Amid shortages, the US buildout ensures AMD can ramp production of Instinct GPUs and EPYC CPUs without the constraints hitting competitors like Intel. By diversifying away from Taiwan (85% of advanced nodes today), the agreement mitigates supply disruptions, ensuring stable flows for AMD's chips. Scaling production and securing supply will matter for AMD the most in the next 5-10 years growth. The growth could be 80-100% YoY or higher; or it could be in the 60%. The aggressive TSMC supply ramp is reassuring the higher growth point. Conclusion: AMD stands at a pivotal inflection point in 2026, where the explosive rise of agentic AI demanding 40-100x more inference compute through its 24/7, multi-step orchestration positions the company to potentially triple its EPYC CPU revenue to $45-60B+ by 2028 while scaling Instinct GPUs to tens of billions annually by 2027. Agentic AI demand could push AI CapEx closer to $1 Trillion in 2027, far higher than most estimates. Dr. Lisa Su, AMD's visionary CEO, is masterfully securing supply to harness this massive demand by prioritizing operational execution and deep TSMC collaboration, ensuring readiness for the second-half 2026 AI ramp. Dr. Su has explicitly called out surging EPYC demand for agentic tasks where CPUs power head nodes and traditional workloads alongside GPUs while guiding for data center dominance through proactive capacity planning and partnerships like Nutanix ($150M investment for open agentic platforms) or providing tens of millions CPUs for OpenAI, $META, $ORCL, $AMZN, $MSFT, $GOOGL and others. Her strategy includes multi-year TSMC agreements for advanced nodes (N2 for Venice CPUs and future Instincts), diversifying beyond Taiwan to mitigate risks, and unveiling innovations like the MI455X GPU at CES 2026, which she touted as enabling "the next trillion-dollar market opportunity" in physical AI. Dr. Su's forward-looking vision predicting AI reaching 5 billion users emphasizes "AI everywhere," backed by hardware like Ryzen AI chips, all while declaring demand "going through the roof" and committing to scale without bottlenecks. TSMC's aggressive ramp-up, fueled by $52-56B in 2026 capex (up 37% YoY) and 10+ new fabs across Taiwan, the US (Arizona cluster expanding to 6+ modules with $165B+ investment), Japan, and Europe, provides profound reassurance for AMD's supply stability. The January 2026 US-Taiwan agreement committing $250B in investments and credit guarantees for US reshoring accelerates this, granting tariff relief (15% rates with 1.5-2.5x exemptions) tied to capacity buildouts, enabling TSMC to potentially double output over the decade to meet AI wafer hunger. This translates to 20-30% higher wafer allocations on key nodes, sidestepping Intel-like shortages and empowering Dr. Su's team to deliver on hyperscaler demands without disruption. Ultimately, this synergy cements AMD's leadership in the agentic era, promising sustained growth, $5T+ valuations at scale, and a resilient path forward as AI reshapes the world. This is NOT Financial Advice! Video source: AMD CES 2026

Mike

44,460 просмотров • 5 месяцев назад

New episode on the optimal exercise intensity, duration, and frequency to prevent and reverse heart aging! In this podcast episode, Dr. Benjamin Levine discusses his groundbreaking research, which reveals how three weeks of bed rest can have a more detrimental impact on fitness than 30 years of aging. Dr. Levine details his research findings that show how a structured exercise regimen can reverse up to 20 years of heart aging by improving both shrinkage and compliance, as well as enhancing aspects of vascular age by 15 years. He also discusses how resistance training and aerobic training have profound differences on the heart, what risks are linked to high-intensity exercise, why recovery is key for the heart, how exercise duration and intensity affect coronary calcium levels, what exercise dose increases Afib risk, and so much more. This episode is a must! Available on YouTube, Spotify, X, and everywhere else. Links in comment. Timestamps: 0:00 - Introduction 1:31 - Bed rest vs. 30 years of aging 5:18 - Recovering from bed rest 6:49 - Does exercise protect against long COVID? 11:27 - Bed rest as a model for space flight 12:24 - How bed rest affects heart size 13:52 - Why a brand-new rubber band mimics a lifetime of endurance training 17:23 - The exercise dose that preserves youthful cardiovascular structure 19:32 - Reversing 20 years of heart aging 23:14 - Reversing vascular age by 15 years 28:38 - Why start an exercise regimen in your 70s? 34:26 - High-intensity exercise risks 37:51 - Balancing high- & moderate-intensity training 42:49 - Training for health vs. training for performance 43:57 - Why muscle mass & cardiorespiratory fitness are like retirement funds 45:12 - Make exercise part of your personal hygiene 46:16 - Why VO2 max correlates with longevity 53:43 - Cardiorespiratory fitness & mortality 59:21 - How does change in fitness over time affect mortality? 1:01:34 - Exercise non-responders 1:05:23 - Limiting factors for VO2 max improvements 1:08:20 - How marathon training affects heart size 1:12:34 - Heart adaptations in purely strength-trained vs. endurance athletes 1:18:23 - Why pure strength-trainers should incorporate endurance training 1:22:07 - How strength training affects blood pressure 1:26:41 - How exercise influences cardiac output 1:28:39 - Does CrossFit count as endurance training? 1:31:04 - Exercise for improving blood pressure 1:36:11 - Lifestyle strategies for treating hypertension 1:38:40 - Why recovery is key 1:42:36 - The best indicator of being overtrained 1:43:36 - Estimating training zones 2-5 1:50:00 - Why HRV is a poor recovery indicator 1:55:16 - Why men are faster runners than women 1:58:49 - Can women achieve similar aerobic exercise benefits doing 2x less? 2:00:21 - Possible cardiovascular benefits of HRT in women 2:02:12 - Defining “extreme exercise” 2:04:00 - How exercise volume affects coronary plaque calcification 2:10:50 - How exercise duration & intensity affect coronary calcium levels 2:14:03 - Why high exercise duration & intensity increases Afib risk 2:16:33 - What exercise dose increases Afib risk? 2:17:59 - Managing stroke risk in athletes prone to Afib 2:21:14 - Why you shouldn’t become an endurance athlete to “live longer”

Dr. Rhonda Patrick

266,327 просмотров • 2 лет назад

$AMD $5 Trillion MC Is Inevitable Long Term👑 This thread will focus more on Inference! 2026 EPYC "Venice" $TSM 2nm to save Large GW Scale Inference by 40% more than Prior Turin gen. Context: EPYC Turin achieves ~$0.001 per million tokens for batch inference vs $0.02-$0.12/ million tokens as I wrote the thread below. Venice is going to lower cost down to $0.0005-$0.0006/Million Tokens. OpenAI spent roughly $20B on Inference and Training, where 80-90% of that was for Inference per Analysts. AKA Renting Compute is Expensive AF! In this thread, I want to focus on why most analysts and investors are underestimating the role EPYC "Venice" and future Gen on overall Data center revenue. And $TSM ramping up 2nm supply early is a confirmation that AMD will be a major buyer long term. I will also link the thread the Gap between AMD Analysts & Reality and 2nm Ramp Thread so you have more comprehensive view of what I'm writing here. Before I go into detail this is my 2026 Projection: AI GPUs: $35-$50B EPYC Data Center: $15B-$17B Client Segment: $12-$13B Gaming: $6B Embedded: $4B-$5B Total Revenue $70-$100B Non-GAAP net income $18B-$25B Non-GAAP EPS $10.97-$15.40 Foward P/E 55x-70x= $603-$1,078 AMD's Analysts are projecting $0 Revenue for MI450 and sluggish EPYC Growth. Meaning, all analysts are either full of 💩 or Sexist, you decide! Analysts are also projecting 0% growth on AMD "Secret Weapon" Chip as $MSFT said we are at significant Windows refresh and upgrade cycle. Do you think TSMC would allocate more 2nm supply to $AMD at $0 MI450 revenue and sluggish EPYC? 1. EPYC is going to be the leader in lowest Inference! Current Turin cost saving is 95% vs $NVDA or 98-99% on Inference cost when you factor in renting Inference compute from Amazon Web Services, Microsoft Azure, or $NVDA Neocloud pets. TSMC claimed: 10-15% higher performance at iso-power, 25-30% lower power at iso-speed, and ~15% higher transistor density compared to 3nm. This reduces operational expenses (energy, cooling) while increasing throughput per chip. EPYC Turin achieves ~$0.001 per million tokens for batch inference (via vLLM on models like Llama 3 70B), driven by high core counts and low hardware costs. EPYC Venice offers ~1.7x overall performance and up to 70% more compute capability per core, with up to 256 cores (512 threads). Enhanced vector/AI instructions and open-source firmware (openSIL) optimize for inference workloads. AMD Incorporates AI Engines (now part of AMD's XDNA) for on-chip acceleration, improving efficiency for low-latency and edge inference. This reduces reliance on discrete GPUs, lowering system complexity and TCO. Venice SKUs are projected at $3,000-$15,000 ($5,000 for 256-core flagship), far below NVIDIA Rubin ($50,000-$90,000) or AMD's own MI450 GPUs ($40,000-$50,000). High memory bandwidth (up to 1.6 TB/s) supports efficient batch inference. Venice is designed exactly for Large customers that want to lower Inference Cost and MI450 Helios is for Customers that want Training at lowest TCO, TDP as well as lower Upfront 1GW scale(Full build $35-$40B vs $NVDA $55B-$80B). 2. Real World Example: OpenAI's 2025 inference spend reached ~$20B, escalating to even higher total compute rental (mostly inference) amid token volume growth(from video generating). By 2026, with usage doubling (consistent with industry trends: token demand grows 2-5x YoY), assume OpenAI processes ~1,800 billion million-tokens annually $NVDA Blackwell at $0.02-$0.12 is $36B(most optimized) Rubin is projected to be at $0.01/million tokens or $18B annual Inference Cost vs $AMD Venice $0.0005/million tokens or $0.9B annual Inference Cost => Massive saving for OpenAI or anyone that are paying 80-90% Annual Bill for Inference compute. In short, it is unsustainable to pay this much rent vs owning for all current AI players for the medium to long term. Rubin excels in low-latency decode (if Groq integration from $20B deal in 2027-2028), but Venice dominates batch (80% of inference by 2030). Actual savings depend on deployment scale (OpenAI's 6GW AMD plans), electricity rates, and software maturity. If Rubin only hits $0.03, savings swell to $53.1B vs. $17.1B. 3. Will running Inference on Venice and future Gen slow down response generation in 2026 and beyond? Human perception of "fast enough" for chat, agents, search augmentation, summarization, coding assistance is roughly Meaning, EPYC may generate $100B a year on data center revenue, Hence $MSFT $AMZN $META $GOOGL OpenAI xAI and 42+ Countries are leaning AMD for Inference, because the cost saving is MASSIVE! 4. Regular users (you, me, people using ChatGPT, Claude, Gemini, Grok, Perplexity...) are extremely unlikely to notice any slowdown and in many cases might even experience slightly faster or more consistent response times if the industry heavily shifts toward AMD EPYC for inference. What actually happens when companies save massively on inference? When OpenAI , Anthropic , Gemini , Grok Meta .... save billions on the batch/enterprise/RAG layer using EPYC Venice, they typically do one or more of these things with the savings, none of which make your chat slower but enhancing their bottom line(Profit) ~Keep prices the same → make more profit ~Lower subscription prices / increase free tier limits ~Train bigger & better models more frequently ~Offer longer context windows ~Add more reasoning steps / tool calls / agents per query ~Improve multimodal capabilities ~Build more data centers / reduce throttling during peaks In practice the consumer experience usually gets better, not worse, when inference becomes dramatically cheaper. Prime example is $META leaning AMD heavily or currently AMD largest customer. or Grok 2 to Grok 3 heavily used AMD for Inference saving. And most Grok Users reported Groke responses snappier, not slower. 5. What does this mean for potential Revenue? Noted that TSMC is massively ramping 2nm supply for $AMD both MI450 and EPYC. EPYC Conservative projection: FY2025: $10.5B(best Est) FY2026: $16B FY2027: $29B FY2028: $49B FY2029: $75B FY2030: $100B Large customers: $META OpenAI $MSFT $AMZN $GOOGL xAI (Apple?) Smaller customer: $DELL $HPE $SMCI and 42+ other countries. The roadmap to $5 Trillion is very much inevitable as Inference Cost from Renting or owning $NVDA are too high, but $NVDA will still dominate Training market share, where MI families are likely to take 15-20% market share, but the TAM is also expanding Rapidly. Most Institutions are projecting $2-$3Trillion TAM by 2030. $NVDA said $4 Trillion. Dr. Lisa Su said $1 Trillion+ by 2030. So you decide on how much TAM. If you enjoy this kind of analysis, Slap the Like/Repost and Bookmark to please the X Algo as it is Free.99! If you want to support my work further, consider subscribe to see more in-depth analysis! Alright, that is it. Not Financial Advice!

Mike

102,223 просмотров • 6 месяцев назад

Just in $AMD Anush "Speed is the moat"|ROCm🎙️ In the race to define the future of AI, what's the one advantage that truly lasts? It's not proprietary tech, argues Anush Elangovan Elangovan, VP of AI Software at AMD , but the sustainable speed of innovation. He explains why AMD is rejecting the "walled garden" model for its open source ROCm stack, betting that an open community flywheel is the key to victory. Listen to understand how this open strategy is designed to out-innovate closed systems by empowering developers to solve everything from frontier-model challenges to the mundane, everyday problems that define the "last mile" of AI. AMD ROCm Software: Part 1 Transcript [00:00:00] Andrew Zigler: Joining me is Anush Elangovan, VP of AI software at AMD. And when people talk about AI compute, the conversation often stops at hardware specs, but it's more than just physical chips that win the game. It's also the software ecosystems supporting them. [00:00:18] Andrew Zigler: The prevailing strategy in the industry has been to build something like a walled garden. You know, something closed, proprietary locks, developers in. But AMD is betting on an entirely different play, open source acceleration, and with rock, their open source AI software stack. AMD is building not just hardware parity, but an innovation flywheel that's powered by the community with interoperability and the freedom to scale without all of that pesky lockin. [00:00:48] Andrew Zigler: And in this world, speed is your moat and how fast you can innovate while your platform remains open, flexible, and standardize across all of its applications. That's what we're gonna explore [00:01:00] today. So Anush, I'm really excited to have you here. Welcome to Dev Interrupted. [00:01:04] Anush Elangovan: Thanks for having me. Uh, super excited to chat about it. [00:01:07] Andrew Zigler: Amazing. Well, let's go ahead and dive right in with kind of what I laid it out with in the beginning, the idea of the moat and it being about speed. I wanna unpack that a bit because that came from you when you and I first spoke. And I, and I want to know, you know, how do you define speed inside of AMD beyond just things like hardware, benchmarks. [00:01:27] Anush Elangovan: Yeah, that's a very good question. So when we typically talk about speed, everyone's like, Hey, hardware benchmark specs, right? Like, uh, memory bandwidth or, or flops. And that is one important part of it, uh, AMD does very well. With that, we do have, a, a very good history of executing on that axis. [00:01:47] Anush Elangovan: But when I say speed is the moat, it is about, uh, how we prepare, how we build the muscle to run the race for a long time and run it fast. And it is [00:02:00] not about a single point in time that you've, you've beat some you know, benchmark and, and you declare victory. It's about building the ability to consistently develop and deliver. [00:02:13] Anush Elangovan: Both hardware and software innovation at scale and do it fast, right? Like, you know, we we're increasingly getting to a point where models come out and they're, uh, you know, a year or two ago it was like, Hey, they work on AMD on day zero, which is great, but now they are performing on AMD the day it releases, right? [00:02:32] Anush Elangovan: So, what does it take to Prefetch where the industry is going? Be prepared to intercept. At that point is what you know, I, I refer to as you know, the, the speed factor in, in creating this mode, right? And the mode is just shed all things that hold you back and run as fast as you can. [00:02:53] Anush Elangovan: Uh, because the pace of innovation that is, uh, being seen in, in AI [00:03:00] industries is just. Amazing. Right? And it's like, it's transformational at at how you generate electricity. It's transformational as at how you build data centers. It's transformational at how you deploy compute, networking. It's transformational at what kind of use cases you, you know, uh, use AI for. [00:03:17] Anush Elangovan: Uh, and for that, you need to be prepared to, see what comes tomorrow and be prepared to run the race tomorrow. [00:03:23] Andrew Zigler: Yeah, it's a really great perspective because it highlights that it's not just like a checkpoint that you run through. I like how you called out, like it's not just hitting that benchmark or being the best in class at that moment, in that snapshot, it's about having a. The throughput and about having that dedication to the idea and continuing to deliver on it. [00:03:43] Andrew Zigler: It's not just crossing the threshold, but it's also being the engine. And that's what, that's what protects a business. That is the moat, because the moat is that innovation layer, the faster and more, uh, future forward. That you can work and think, [00:04:00] you know, the better. Uh, we, we talk a lot about like future forward work styles. [00:04:04] Andrew Zigler: Like what are the things I could be doing right now today that are gonna be like, way more useful tomorrow? Let, let's abandon those, workflows that are older and that kind of like, that translates into. An advantage when you work that way. You know, what kind of things have you learned working with, uh, like across all spectrums of people who would use ROCm, right? [00:04:23] Andrew Zigler: You have like the developers, but then you also have the enterprises and you have this large span of adoptees, right? So what is the, what does that look like that you learn? [00:04:32] Anush Elangovan: Yeah, so, so the way I look at it is there are gonna be pockets of different, uh, you know, cadences, right? Like, so people who are deploying in enterprises, for example, right? The validation and how long it takes for them to deploy an LLM that's secure. It's, with guardrails, et cetera, maybe longer. [00:04:52] Anush Elangovan: but you still have to go through the process and you have to be prepared to like, walk that walk to deploy an enterprises. That doesn't mean it's [00:05:00] not fast, that's as fast as you can do for that industry, right? And if you are deploying AI in healthcare, right, it's, it's got its own, uh, cycle. [00:05:07] Anush Elangovan: but in each one of these, you want to see how, like, go down to the essence of what is it that you actually have to do. And, you know, I, I, I like how you framed it. It's like it's, you shed your prior assumptions of how things are done, right. And, and you kind of build up from a, uh, first principles, uh, approach to say, this is how I could use AI to unlock, whatever I'm doing. [00:05:33] Anush Elangovan: And, and, some of it, you know, it's good to really step back and look at. Just question every part of it, right? Like right now you're getting chat GPT and, Gemini competing for like, math, olympiads and, and, uh, college, uh, reasoning, uh, tests. Right? And, and those are like that, that is amazing and increasingly like complex tasks that they're trying to do. [00:05:58] Anush Elangovan: But there may also be like. [00:06:00] More mundane things that AI could, could get applied to. Right? And, and so when we think about shedding old ways, you wanna shed it not just in like the tip of the spear. It's like, you know, I'm gonna see what's the frontier model. It's also, it could be something as simple as. [00:06:18] Anush Elangovan: How do you choose a, a movie, uh, you know, like a recommendation system, right? Or, or, uh, an automated, uh, flight, uh, rebooking system. So the moment, you know, your flight is late, uh, right now it's a notification, right? It's like, oh, you got a text message saying your flight's late. And I got that like three times this week. [00:06:38] Anush Elangovan: But anyway, uh, and, and, and, and, I was just like, okay, so if I were to rethink this. All this MCPs that we have that should be hooked up into an MCP that says, your flight's delayed. Here are your options. If you want, you know, these are the paid options. Yeah. Here are the free options. This will get you back into your you know, Toronto airport [00:07:00] tonight. [00:07:00] Anush Elangovan: Or if you stay, here's a hotel plus this, plus this, plus. It's just like, go ahead is all I should say. Versus now I'm like, okay, can someone, you know, can I call a travel agent? Can I do this? Can I go online and log into And you know, so we gotta fundamentally rethink even those like small, nuances of, things that we do that can be automated out and AI is really, really good at doing something like this, right? Maybe I just explained an AI startup idea right now. Somebody should just start that. [00:07:29] Andrew Zigler: I think you did. Yeah, you definitely did. Someone, one of our listeners is definitely going to lift that off of you. I, I, I, you know, I hate being on the receiving end of those. You feel a little helpless and then you have to like, follow the whole flow. So I know what you mean. Like I, I like how you called out that the build and this like. [00:07:45] Andrew Zigler: Where speed is your moat and the innovation layer is protecting you, is what makes you better than your competitors. How you scale that and you bring that to market. So by understanding the problems that you're solving, uh, throwing away those older assumptions, but also [00:08:00] recognizing that like. We're building every single day, new things and new ways of using stuff that we're still figuring out the implications of. [00:08:08] Andrew Zigler: And so when you have a lot of velocity and you're introducing a lot of new ideas, and maybe you have that workflow now that automatically rebook your flight off of your late flight text message, and uh, I know I would certainly use it, but you know, what kind of philosophies guide the way that y'all think about building this ecosystem to manage that stability while letting folks. [00:08:29] Andrew Zigler: Play with the speed and the assumptions and the airplane re bookings. [00:08:34] Anush Elangovan: so, so I think, you know, we need to peel one layer down, right? and the philosophy is, Hey, we, we just discovered electricity, right? And you know what we're gonna do? We are gonna make motors, uh, or dynamos, right? Like engines. Uh, sure. We don't know if it's gonna be a Ferrari that you're gonna make, or it's a a a a dump truck. [00:08:57] Anush Elangovan: That's good for doing this. But let's [00:09:00] let, which is also required, right? You need a dump truck. You need a garbage truck. And, [00:09:04] Andrew Zigler: Yeah. You need the [00:09:04] Anush Elangovan: course you need, uh, a Ferrari for a midlife crisis, right? So, [00:09:09] Andrew Zigler: precisely. [00:09:10] Anush Elangovan: But, but my, uh, point is what do we build next? And, uh, and this is what I meant by like, okay, let's, let's take those baby steps to build the. [00:09:20] Anush Elangovan: Infrastructure that's required that we know we'll have to use, right? So, so if I just discovered electricity, okay, great. Now one, how do I save this electricity and how do I use it? So there's battery technology, so you need to do something like that, right? Like so. But then you also want to make it into an actionable thing. [00:09:37] Anush Elangovan: You want to make it for like automobiles, or you wanna use it for, you know, powering, uh, entire cities. So it is that transformational. So, uh, AI is that transformational. So, if you distill down, it'll, it'll come down to how do we think about, what we can do with this this fundamental technology that, We may not be aware of what it [00:10:00] is gonna unlock next, but at least you know the next step is clear, right? It's like a dense fog, you know, it's gonna be like, it, it's the right path. You see the light, but it's kind of like out there and, and the steps you're taking are concrete and you're like, okay, this is good. [00:10:16] Anush Elangovan: I, this is better than where I was or where we were. So we are moving forward. So you can build with the. Intuition from what you see in the short term and a tactical view, but towards what you think the future is gonna be. [00:10:28] Andrew Zigler: Right. You almost like we're all in this like fog of war, right? And like you said, you're reaching out and you're trying to step through it. You could think of it too, as like you're in the dark and your hands are up in front of you and you know that. You're, you're not gonna run your face into a wall because your hands are out in front of you, but you're not gonna maybe do much better than that. [00:10:45] Andrew Zigler: So that's kind of like, I think the eco, the, the industry, the world that we find ourselves in, uh, and we all have to, then this becomes the power of an ecosystem, of a group of people working together to create that layer of, [00:11:00] uh, of establishing the [00:11:01] Anush Elangovan: exactly. And I, I, I just, instead of, you know, saying fog of war I describe it as like, you're in this. Beautiful valley with like a morning, uh, fog that's in. You can smell the flowers. You, you hear the birds. You are like, okay, it's, we are in like, uh, utopian paradise and yes, I just need to like, continue the walk, right? [00:11:24] Anush Elangovan: and then move forward with that, conviction that you're in the right spot. [00:11:27] Andrew Zigler: Yeah. So let's talk about that ecosystem world. This nice, I love how you describe it, this grassy side of a hill in the morning that's covered in some mist and maybe we can't see 30 feet in one direction, but it sure is a beautiful hill and it smells nice. And so we're all here. And why is, in that world, why is. [00:11:44] Andrew Zigler: You know, open source, their strategic advantage that y'all are going for in the AI hardware market. And, and then how does like ROCm turn that into wins for people within that ecosystem? [00:11:56] Anush Elangovan: you know, the, the way we look at it is this, is kind of like how I view [00:12:00] AI and the ecosystem, right? But, but it is for everyone to enjoy. Uh, and so we do want to make sure that. You know, it is, uh, beneficial for everyone. [00:12:09] Anush Elangovan: The ecosystem can come in and, and innovate. It's an open innovation engine. and uh, it is very different from, you know, having a walled garden with, Hey, only I know how to do this and I'm gonna do it and throw it over the fence and you can use it or keep walking, right? So we'd like to be good citizens that way, but also. [00:12:30] Anush Elangovan: Uh, it is self-fulfilling in a way, right? Like it, the, the pace at which we innovate with open source is unmatched. Like, you know, our serving engines are like VLLM and, and sg l. Those things, uh, those frameworks are like super, super aggressive in terms of how fast they come out with features and how fast they can you know, get performant models out. [00:12:52] Anush Elangovan: And that compared with what, uh, you'd get from, you know, the likes of like T-R-T-L-L-M or something is always lagging, right? Because you [00:13:00] just can't keep up with you know, 200 commits a week just on one particular model to get that model really performant [00:13:06] Andrew Zigler: And, and, and in that world where, you know, everyone can enjoy the winds of this, what kind of customer stories or innovation stories have really stood out to you and excite you about building and creating this place for developers? [00:13:19] Anush Elangovan: Yeah. So I think the parts that are super exciting for me are when when we get to see a customer that is first skeptical. Then they start a little like, okay, fine, we'll give you a chance. Uh, we do a simple, uh, POC and then they're like, huh, this seems to work. Yeah, we told you it works. [00:13:42] Anush Elangovan: You don't have to change one line of code. Really? Yes, no need to change one line of code. Okay, let's try a production workload. So then they try it. Oh, you're more performant than the competition. Yes. We're more performant than, than the competition. So how much does it cost? And we're like, oh, it's your TCO is better with, uh, [00:14:00] AMD. [00:14:00] Anush Elangovan: So again, they're like, wow, okay, good. So now how do we deploy at scale? And then we go deploy it at scale. And when they give a thumbs up on that and they say, this is good, right? That's when you know, you, you see it go full circle from like, oh, we, we've never heard about AMD to like actually deploy to tens of thousands of GPUs In the order of a few months, right? It, it, it really is fascinating to see and very exciting and invigorating to [00:14:28] Andrew Zigler: Yeah. At like a great exposure to a lot of interesting problems. And, and then people using the infrastructure, the, the technology available to solve those problems. Really specific problems by the way, that's often why they're bringing their data and AI to it, uh, is because it is really specific and important for them. [00:14:45] Andrew Zigler: And there's a, a lot I think that other engineering orgs can learn and even emulate from AMD's success and, and having this open source ecosystem and it causing this acceleration within. You [00:15:00] know, uh, customers and enterprises that use and adopt the tools and, and, and that creates an advantage. And that goes back to why we're talking and like the real thesis of our conversation today. [00:15:10] Andrew Zigler: So how do you think engineering leaders that are listening to this and obviously tapping into this great success AMD has from an open source flywheel, how do you think other, other folks building in the same space can foster that open, first, that open source oriented culture in order to, you know, accelerate their innovation goals? [00:15:29] Anush Elangovan: Yeah, that's a very good question. So the startup that um, was acquired by AMD we, we built, I mean, we started off doing iot stuff and you know, smart ring and all that, right? But in the, the end of like, uh, and not the end, the last six years of the company was building ML compilers. [00:15:47] Anush Elangovan: And ml, ML compilers are like super, uh, complicated, sophisticated, advanced algorithms, dah, dah, dah. but it was all open source, right? So our VCs were like, wait, what do you mean your core [00:16:00] IP is open source? And um, the speed is the moat applied even then, right? It was just like, yes, if you have an idea that. [00:16:08] Anush Elangovan: Because someone saw this idea that you are, they're gonna be able to catch up, then you probably have the wrong idea anyway. But if they are, you know, you execute and they're gonna catch up, that you should assume they're gonna catch up. Right? So you gotta move forward. So keeping it open source is super important. [00:16:25] Anush Elangovan: But also to your question on like, you know, the learnings from an AMD standpoint, right? If there are, hard problems, I'd say dig in and work through it, right? Like there's no way but through it, right? That should be the simple mentality. And more, uh, frequently than not. you'll see that you'll just make it through in a, in, in good form. [00:16:52] Anush Elangovan: But if you doubt it and you're like, oh, I don't know if I should commit, if I'm, I, you know, what should just commit to do the right thing [00:17:00] every step, right? Every step, and just keep taking one step in front of the other. And in no time you'll see that you'll be running. Right. And, and yes, the first few steps will be like, yeah, everyone's complaining about your software quality. [00:17:15] Anush Elangovan: Everyone's complaining about this and that, and it doesn't work. And, and a few steps in, you know, you get, you get the hang of all the complaints that are coming in. You get the feedback loop. You're like, okay, what, what are you prioritizing again? One step in front of the other, right? You just keep knocking that out and then you get to a point where you're, it just becomes second nature, right? To do the, to do the right thing. And, and then yes, if someone gives you two options, you'll be like, fine. This is, uh, you know, there's always the resource trade off. There's always a human capital trade off, but what's the right thing to do? of course, I, I'm pragmatic about what we choose, but, but if the right thing for your long-term success is dig in, go first, principles, make it [00:18:00] happen. [00:18:00] Anush Elangovan: Well. Then just go for that. There's, there is no shortcut to [00:18:04] Andrew Zigler: acknowledging, you know, how it aligns with your mission, your core company goals, and what you're looking to achieve. And, and I, I love how you rightfully called out that in the open source world and you know, you have your technology that you've built, what you think is your moat upon, right? [00:18:22] Andrew Zigler: It's your code and, and to open source that, or to just make it where anyone could peer in is, you know. Scary in one regard, but two, it just kind of feels like you're handing away your throne room in some kind of sense, a very direct feeling sense. But the ultimately, you were really right to call out, and this is something I think about all the time, that the real power there is still the speed This the speed. [00:18:42] Andrew Zigler: That was the moat at the beginning of our conversation. It's the speed in combination with your. Very specific domain understanding of what you're building and what you're creating, and your new role as the steward of that world and how people plug into it, which [00:19:00] has frankly, a lot more influence and power than lording over a closed. [00:19:04] Andrew Zigler: You know, repository or an ecosystem, and like you said, like throwing things over the wall. Sure. There, there might be people always on the other side of that wall, but you're not gonna have a great connection with them. You're not gonna be able to really clearly understand them. I, I like your metaphor of the side of the field of the mountain a lot more. [00:19:23] Andrew Zigler: But, but in the, in this world, you know, where. That speed is, is the power and, and open source is just one way that you can harness that speed to get really far ahead and to innovate. , There's other parts of this equation that you can be experimenting with too, and I'd love to pick your brain about them as a software leader and, and, and one of them is about looking forward and kind of understanding that future that we're all building towards and beyond today's models and hardware. [00:19:48] Andrew Zigler: You know, what do you see as the next major bottleneck or opportunity in the AI compute space? As, as you know, enterprises and folks start to get a little more mature about what's available to [00:20:00] them. [00:20:00] Anush Elangovan: Yeah, I think, the bottleneck and opportunity is, uh, what I'd call, call walking the last mile of ai. Right. Uh, and like I I, I gave you an example, uh, previously, but, but it's similar to that. It's like there are cases where Humans have so many, uh, things to do in your day. You know, like the, if we sit down and actually had a customer focus like, okay, these customers lives, I'm gonna save four hours of this customer's life. And if you actually sit down and look at all of that, it'll be. Easily automatable, easily you know, uh, applicable, uh, for ai, right? [00:20:39] Anush Elangovan: Like, but then making it happen is gonna take a little bit, right? It's like maybe it's, uh, paying your utility bill, right? Or something like that, right? Or, or, your healthcare explanation of benefits. Uh, like, I'm sure you get an explanation of benefits, and I'm like, I, I don't even know what that thing is. [00:20:55] Anush Elangovan: It's just like EOB and like. [00:20:57] Andrew Zigler: it's a big, a big old PDF. Yeah, [00:21:00] exactly. [00:21:01] Anush Elangovan: Like, like, I'm like great straight to the, uh, shredder, right? And but that could be, you know, automated with the ai, right? It, it, it'd be like, Hey, the summary of this thing is you went and visited this day. Everything is okay. Everything is paid for, so don't worry, it's not a bill. [00:21:17] Anush Elangovan: That again, the same, uh, thing, but the sense of what that information overload is could be. Digested by ai, uh, accumulated over time and retrieved when you need it. Like, I don't, I actually don't even need to know this EOB right now, unless of course, whenever I need to know it, that maybe, you know, like for some benefits I need to figure out what do, what did I do over the past year and how do I apply it? Source:

Mike

14,195 просмотров • 8 месяцев назад

$AMD| The FOMO to buy AMD Chips is NOW 🧵 Not Financial Advice! DYOR! Research Purpose Only! The Inference Queen is the biggest winner in Agentic AI where all other CPUs are struggling to compete with a 2yr old EPYC Turin and EPYC Venice is in mass production phase. AMD stresses deployability today on standard x86 platforms (no proprietary architectures required), full software compatibility, and open standards. This positions Venice + Helios as a practical, high-density alternative to competing solutions while underscoring that agentic AI shifts the balance toward CPU-rich racks alongside GPUs, and most importantly, lowering the cost of token to accelerate adoption and innovation. Context: The Wall Street Journal yesterday came out with an article that OpenAI is condiering drasstically lowering the token prices to win more customers from Anthropic. The narrative "they" are trying to exacerbate the current AI selloff won't last long. This is a fundamental misunderstanding of what is going on, or what I already discussed for months and years. Followers and Subscribers already knew this for years, that this day would come, where token cost will bcome the central discussion among enterprises as there is no such thing as unlimited budget or Tokenmaxxing when they use $NVDA chips or In-house Hyperscalers chips. I will link various threads if you are interested in understanding the full picture from supply chain to recent TSMC Rapid 2nm expansion up to 12 Fabs total by 2027/2028. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. The OpenAI-AMD 1GW Helios deployment (starting H2 2026) represents a pivotal vertical integration move that directly supercharges the inference economics. This isn't incremental; it's a structural shift toward ownership of massive, optimized rack-scale capacity, enabling the lowest token costs and triggering the enterprise adoption flywheel. We need to be honest, $AMD is the only company that made a big bet on Inference since the day Chatgpt became sensational where $NVDA and others were betting big on Training. At the end of the day, Token bill from Anthropic has to obey economics. Meaning the bills rise, companies have to get more out of it to justify the cost. It cannot be an unlimited inference budget, and it has to show up on efficiency, profitability and operating leverage. 1. Tokenomics After you understand this, you will understand why Citi cited Anthropic is likely to sign a deal with $AMD along with Hyperscalers, AI Labs, Sovereign AI like Softbank 5GW in France and many other countries. However, OpenAI and $META are now wanting faster deployment, and they are AMD shareholders now, they have prioritized allocation. Anthropic and Hyperscalers just cannot compete when Helios Rack lower token cost to$0.0003–$0.0005 per million tokens at GW scale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens Now, OpenAI, META and Hyperscalers can lower Inference cost even further with $AMD EPYC Venice "dense rack" or Agentic AI Rack. AMD published a detailed technical blog emphasizing that the future of agentic AI autonomous, multi-step AI systems requiring heavy orchestration, databases, caching, APIs, and control planes demands massive CPU-dense rack-scale infrastructure, not just GPUs. The catalyst prominently positions their upcoming 6th Gen EPYC "Venice" processors as the key enabler for next-generation dense racks, delivering leadership throughput under real-world power, cooling, and density constraints. ~EPYC Venice (Zen 6 architecture, up to 256 cores / 512 threads per socket) is projected to deliver exceptional rack-level performance. In AMD’s modeled 100 kW rack comparisons, Venice-powered systems are expected to achieve ~3.30x the throughput of NVIDIA’s Vera (88-core Olympus) baseline across a broad mix of agentic-supporting workloads. ~This builds on current-generation 5th Gen EPYC "Turin" (up to 192 cores), which already delivers ~2.37x rack throughput vs. Vera and ~1.6x vs. Intel’s Xeon 6980P (128 cores). ~ Liquid-cooled Turin deployments already support >27,000 CPU cores per rack today. Venice is architected to push this beyond 36,000 cores in the same rack class, dramatically increasing concurrent agent capacity and overall infrastructure efficiency. 2. Ownership vs renting compute from Hyperscalers matter to OpenAI and only owning $AMD chips can meaningfully lower token cost for enterprises. ~Eliminates cloud overhead: No provider margins, utilization buffers, or egress fees. Direct control over power contracts, cooling, scheduling, and orchestration at dedicated facilities. ~Helios optimizations at GW scale: Rack-level density (1.4+ exaFLOPS FP8 per rack), high HBM4 bandwidth, EPYC orchestration for agentic workloads, and superior TCO/TDP. AMD's long-standing focus on tokens per dollar/watt shines here 20-40%+ efficiency edges in inference-heavy scenarios. ~At 1GW+ optimized deployment, inference hits $0.0003–$0.0005 per million tokens (community/analyst models tied to Helios metrics). This is dramatically lower than typical rented/cloud equivalents, especially for high-volume output tokens in agentic flows. High token bills today, enterprises running heavy agentic/coding/analysis workloads can face $50-100M+/month at current API rates (flagship models $5-30+/M output, scaled to massive volumes). Post-Helios compression, same volume will drop to $10-15M/month (or better) via lower underlying costs passed through as pricing flexibility, volume tiers, caching, or batch discounts. ROI thresholds collapse. More companies greenlight pilots → production → massive scaling. Agentic AI (autonomous workflows) multiplies token demand exponentially, but affordability removes the friction. OpenAI gains flexibility, Unlike more cloud-dependent rivals (Anthropic), they can lower effective pricing, offer aggressive enterprise bundles, or absorb volume without margin destruction directly tackling "high token bill" complaints while maintaining profitability as usage explodes. 3. Agentic AI Models shifted CPU:GPU Ratio to 1:1 toward 3-5:1 with Explosively Token-Hungry Workloads Agentic AI (autonomous, multi-step agents with planning, tool use, iteration, and self-correction) is fundamentally more compute and token intensive than conversational or single-turn generative AI. Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. ~Agents often generate 10–100x+ more tokens per task due to iterative reasoning chains, multiple tool calls, verification loops, and long-context orchestration. ~Goldman Sachs forecasts token consumption multiplying 24x by 2030 (to 120 quadrillion tokens/month) largely driven by agentic adoption in consumer and enterprise. ~Enterprise data shows agent-pattern workloads growing at 680% annualized rates, projected to surpass conversational AI in token volume by Q3 2026. ~Daily enterprise agent token consumption is already in the billions, with complex workflows (coding, workflows, analysis) amplifying this dramatically. 4. Competitive Edge: Winning Customers from Anthropic Anthropic’s Claude models (especially Opus/Sonnet) excel in complex reasoning and agentic coding, commanding premium positioning. However, their higher underlying costs (heavier reliance on third-party cloud with margins) limit pricing flexibility compared to OpenAI’s owned Helios capacity. Anthropic is on track to generate $10.9 billion in Q2 revenue. The company expects to achieve its first-ever quarterly adjusted operating profit of $559 million. However, sustaining full-year profitability remains challenging due to immense computing and model training costs The truth is, Anthropic has no choice but to buy as much $AMD chips as possible if they want to compete with OpenAI or get investors attention. This 5% adjusted operating profit to revenue ratio is just pathetic. Current pricing dynamics (2026): OpenAI already undercuts on many tiers ( flagship output tokens significantly cheaper than equivalent Claude Opus). Nano/mini models offer 5–10x advantages for volume work. Anthropic holds edges in long-context flat pricing and certain reasoning quality. OpenAI after Helios Rack Ownership, At $0.0003–$0.0005/M effective costs, OpenAI gains massive headroom to: ~Aggressively discount high-volume agentic tiers or bundles. ~Offer “unlimited” enterprise plans or usage-based models that Anthropic struggles to match without margin erosion. ~Target cost-sensitive, high-throughput agent deployments (dev tools, automation platforms) where token bills explode. Enterprises facing $ millions in monthly agentic bills will migrate to the provider delivering better economics at scale. OpenAI’s combination of strong models (o-series reasoning) + lowest TCO positions it to erode Anthropic’s enterprise share, especially as agentic becomes the dominant token consumer. Cheaper tokens expand the total addressable market dramatically. This feeds the data/model improvement loop, justifying further capex. AMD benefits from proven scale pulling in more customers (Meta, Oracle, Microsfot, Amazon, Softbank, TensorWave, LumaAI ... already aligned on Helios). Conclusion: Dr. Lisa Su has been laser focused on inference economics since at least 2022–2023, repeatedly emphasizing that the real battleground for AI scalability would be TCO, power efficiency (TDP), and ultimately tokens per dollar and per watt not just raw training FLOPS. While many viewed inference as a secondary, commoditized workload, Dr. Su architected AMD’s roadmap around rack-scale systems optimized for high-volume, sustained inference that would dominate as models matured and usage exploded. Helios represents the culmination of that multi-year bet: a fully integrated, open platform designed precisely for the economics of massive token throughput. This deep, strategic partnership with OpenAI starting with the 1GW Helios deployment in H2 2026 and scaling to 6GW, is the embodiment of that shared vision. Both companies foresaw a future where agentic AI models evolve to become extraordinarily token-hungry: autonomous agents executing complex, iterative workflows with planning, tool use, verification loops, and long-context reasoning. These workloads can consume 100x+ more tokens per task than traditional chat or single-turn generation, driving exponential demand as capabilities improve and enterprises deploy them at scale. By owning and optimizing this massive Helios capacity at GW scale, OpenAI achieves inference costs as low as $0.0003–$0.0005 per million tokens. This structural cost advantage allows OpenAI to absorb the coming token explosion profitably, dramatically lower effective pricing for enterprises, and win high-volume agentic workloads from higher-cost competitors like Anthropic. What was once a prohibitive monthly token bill becomes an affordable accelerator for productivity and innovation. The OpenAI-AMD alliance validates Dr. Su’s prescient strategy and turns the Agentic flywheel into reality: Collapsing inference costs → explosive token consumption → richer data and better models → accelerate greater demand. This partnership doesn’t just address today’s economics, it positions both leaders at the center of the infrastructure buildout that will power AI’s next decade. By delivering the lowest inference economics at scale, OpenAI not only solves enterprise bill pain but gains a decisive weapon to win share from higher-cost rivals like Anthropic. And that is why OpenAI and $META will deploy EPYC Dense Rack Not Financial Advice! DYOR! Research Purpose Only!

Mike

84,951 просмотров • 1 месяц назад

Alex Hormozi’s business advice has taken the world by storm. He’s published two books and has more than 9,000,000 followers on the Internet. This is the first interview he’s ever done all about his writing process. Some highlights: 1. Be loyal to the truth, not your own ideas. 2. Silence the world when you write: Alex closes the windows, wears earplugs, and uses noise-cancelling headphones to create an environment where the outside world ceases to exist. 3. His book, $100M Leads went through 19 drafts (so expect to rewrite, a lot). 4. There’s a big difference between becoming known and becoming respected. Don’t let an algorithm convince you otherwise. 5. The pain is the pitch: The more vividly you describe someone’s problem, the less you need to sell the solution. 6. People buy things from people who can describe their pain better than they can. They assume: “If you understand my problem that well, you must have the solution.” 7. Sell at the point of greatest deprivation, not satisfaction. Offer the steak when someone is starving, not when they’re already full. 8. When writing ads, capture moments of pain as specifically as possible. Don’t say “I was overweight” when you can say “I wore a cover-up at the beach and avoided photos.” 9. Structure your writing time around long, uninterrupted blocks of time. Alex shoots for six hour blocks, even if that means waking up at 5am. 10. Memory is unreliable, so capture your best stories. Alex has an Excel spreadsheet with more than 600 stories from his life. 11. The #1 creator mistake: They keep building new products for their audience rather than building more audience for their product. 12. Alex’s best ideas come from reconciling contradictions. He says: “I look for places where two things seem true, but seem to conflict.” 13. The back cover should function as a mini sales letter. Each bullet point should tease a lesson in a way that makes people curious​. 14. If an idea can’t be operationalized, it’s useless – “What does this change about what someone actually does?” If the answer is nothing, cut it. I’ve shared the full interview with Alex Hormozi below. And if audio is your thing, I’ve linked to Apple, Spotify, YouTube in the reply tweets. — — Timestamps Below 00:00:00 Intro 00:00:20 Hormozi’s cave 00:05:35 Discovery process 00:12:32 Fix problems using MECE 00:13:43 Marketing “$100M Leads” 00:20:55 Writing for YouTube vs Books 00:27:13 The pain is the pitch 00:31:37 Break down the world into frameworks 00:34:09 10x more effort = 1000x results 00:38:24 Understanding the table of contents 00:49:27 The only app I use to read articles (Readwise Reader) 00:50:53 How to think about your audience 00:54:22 The #1 creator mistake 00:57:58 The business of Hormozi’s books 01:02:04 The love for writing as a kid 01:04:34 Hormozi’s ideal university writing class

David Perell

33,676 просмотров • 1 год назад

A beanie that reads your thoughts and turns them into text — no surgery required?@jason grills the co-founders on their noninvasive brain-computer interface, backed by Vinod Khosla, and calls cap on the whole thing (until he doesn't). This episode of This Week in Startups covers a lot of ground: Jason's tactical tip of the day on making everyone the CEO of their domain, a deep dive into Sabi's thought-to-text beanie, a live demo of AI-powered podcast sidebars built by the TWiST audience, the announcement of a new $5K bounty for an annotation tool, and Jason's big five wellness framework. 0:00 Intro & tactical tip: Make everyone the CEO of their domain 1:49 Matt Coffin's "CEO of X" management philosophy 3:04 Community ownership: Deputizing Ricky, Lawn, Bianca, Maddie, Kabir 4:02 Building the Noti Gang: X group chats and community flywheels 5:08 Founder takeaways: Activate your top 1%, make someone the CEO of it 6:18 The streamer trick, parasocial dynamics, and creator ethics 7:59 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at and use code TWIST for 10% off! 9:34 Guest intro: Rahul Chhabra of Sabi 10:01 What is Sabi? Noninvasive BCI in a beanie with 100,000 sensors 10:02 LinkedIn Jobs - Hire right, the first time. Post your first job and get $100 off towards your job post at 11:02 How it works: From fMRI to EEG, from hospital to hat 13:46 The brain foundation model and thought-to-text decoding 15:55 Vetting the founders: BITS Pilani, Stanford, athlete fatigue AI 17:22 Vinod Khosla's investment thesis: BCI must be noninvasive 19:30 Jason's challenge: Say "Calacanis" or it doesn't count 19:50 Northwest Registered Agent: Get more when you start your business with Northwest. In 10 clicks and 10 minutes, you can form your company and walk away with a real business identity — Learn more at 21:30 Reserve your device, release by end of 2026 22:26 Privacy concerns: Does the beanie read everything? 23:47 Jason's Big Five wellness framework: sleep, nutrition, exercise, meditation, socialization 25:41 Bounty #1: AI live sidebar contest — demos from the TWiST audience 26:02 What the bounty asked for: AI personas watching the show in real time 27:04 Oliver's breakdown: What's easy vs. hard about live AI commentary 28:40 Demo #1: Armchair (by Mark Colebrook) — fact-checker + troll personas, live 30:26 Render: Find out why 5 million developers are already using the all-in-one cloud platform, Render. Go to and apply for the Render Startup Program to get $500-$100,000 in free credits, depending on your stage and backers. 35:30 Live political violence test: The sidebar in real time on the WHCD shooting 38:45 Demo #2: Pod Commentators / SideCast — browser-based, Gemini-powered 40:35 Jason's revised Bounty #1 spec: Fact-checker + cynic, public stream or Zoom 42:41 Timeline: check-ins May 1, May 8; final winner May 15 44:51 Demo #3: BMD Pat (by Patrick Hughes) — instant URL, all-snarky personas 46:52 Bounty #2 announced: — a fair-use multimedia annotation tool 49:31 Annotated: the of media commentary 51:11 Contest rules: Jason owns the domain, winner gets $5K + potential ongoing work 53:13 Wrap-up: Bounty 1 = AI sidebar, Bounty 2 = Annotated 🎥 Watch the full episode here 👇

This Week in Startups

16,332 просмотров • 3 месяцев назад

The enemy of sleep is effort. In this episode, Dr. Michael Grandner (Dr. Michael Grandner) explains how insomnia fundamentally stems from a learned pattern of wakefulness and why, for many, the bed itself becomes a source of stress and vigilance rather than rest. We explore practical strategies to retrain your brain to effortlessly anticipate sleep rather than forcing it. Dr. Grandner also highlights why sleep apnea is shockingly prevalent yet frequently undetected, and provides effective treatment alternatives beyond CPAP. Additionally, we discuss how supplements and other substances influence sleep architecture, how to optimize sleep tracking devices, and how to strategically use sleep to enhance cognitive and physical performance. This is an essential podcast for anyone looking to improve their sleep. Links in the comments. Timestamps: 0:00 - Introduction 1:33 - Insomnia vs. poor sleep 3:59 - Does stressing worsen insomnia? 10:28 - Why CBT-I targets wakefulness 12:59 - Reserving bed strictly for sleep 17:11 - Can trying too hard backfire? 18:26 - A fix for nighttime scrolling 21:47 - Can't fall back asleep? Get up 24:39 - Why effort keeps you awake 26:18 - Sleep restriction therapy 28:57 - How to fall asleep faster 31:39 - Avoiding bedtime cliffhangers 33:20 - Sedatives vs. CBT-I 37:33 - Insomnia by the numbers 38:54 - Why sleep apnea is so common 42:31 - Nighttime waking & sleep apnea 48:37 - At-home apnea tests 50:10 - What causes sleep apnea? 52:52 - Sleep stages 1:01:20 - Why do we dream? 1:05:38 - Apnea destroys sleep architecture 1:07:07 - Sleep apnea & Alzheimer’s risk 1:10:07 - Poor sleep, attention & memory 1:13:23 - CPAP alternatives 1:17:26 - Mouth taping 1:19:30 - Measuring apnea treatment success 1:21:33 - Sleep hygiene for chaotic schedules 1:25:00 - Do blue-blockers enhance sleep? 1:25:46 - Why morning light matters 1:30:33 - Should you delay morning coffee? 1:34:30 - Consistent mornings vs. bedtimes 1:38:01 - Revenge bedtime procrastination 1:42:49 - Is 5 mg melatonin too much? 1:50:25 - Are melatonin labels misleading? 1:53:19 - Melatonin & immune system boost 1:54:14 - Melatonin safety myths 1:58:35 - Magnesium, glycine, & L-theanine 2:01:37 - Glutamine & B12 disrupt sleep 2:03:08 - THC & REM sleep 2:09:36 - Does CBD help? 2:12:09 - Alcohol as a sleep aid 2:15:06 - How late is too late for caffeine? 2:19:19 - Staying up late & unhealthy eating 2:24:09 - Is shift work worse than smoking? 2:27:51 - Ideal power nap length 2:29:38 - Napping advice for shift workers 2:32:18 - How to beat jet lag 2:40:22 - Do trackers detect wakefulness? 2:43:56 - Sleep stage tracking accuracy 2:48:24 - Wearable sleep scores 2:57:56 - Evening habits elevating heart rate 2:59:59 - Insufficient REM & deep sleep 3:02:55 - Orthosomnia from sleep trackers 3:07:12 - Better sleep & cognitive resilience 3:09:42 - School start times vs. teen biology 3:14:26 - A sleep hack for athletes 3:16:36 - Sleep banking before competition 3:19:03 - Poor sleep & injury risk 3:24:00 - Why caffeine can't fix poor sleep 3:25:38 - Eye masks & earplugs 3:27:15 - Proven ways to fall asleep faster 3:29:11 - Reading before bed 3:30:02 - Algorithm for falling back asleep 3:31:04 - Proven strategy for deeper sleep 3:32:27 - Reducing nighttime urination 3:34:11 - Does sharing a bed disrupt sleep? 3:35:50 - Are you sleeping enough? 3:37:28 - Do you really need 8 hours? 3:38:42 - Adjusting to your chronotype

Dr. Rhonda Patrick

192,656 просмотров • 9 месяцев назад

$NVDA $GFS NVIDIA’s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIA’s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groq’s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groq’s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groq’s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The company’s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groq’s solution is to move “control” into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groq’s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groq’s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groq’s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that “no useful models” fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groq’s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groq’s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth “upwards of 80 TB/s” and contrasts that with off-chip HBM bandwidth “about 8 TB/s,” asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes “TruePoint numerics,” including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the company’s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in “available speed” for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groq’s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as “tokens-as-a-service,” aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groq’s documentation states its API is designed to be “mostly compatible” with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groq’s go-to-market is not simply “sell chips,” but “sell predictable unit economics per token,” with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native “edge within the cloud” strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groq’s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groq’s marketing materials and partnerships provide signals about demand vectors. The company’s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groq’s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, which—if realized in deployments—can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groq’s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groq’s core proposition—deterministic, compiler-scheduled inference with predictable latency—aligns directly with the segment where GPU generality is least valued and where “good enough” programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIA’s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groq’s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The “technology acquisition” rationale is unusually strong in this specific case because Groq’s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIA’s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groq’s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groq’s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIA’s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chips—while not costless—could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship “inference capacity” in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIA’s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groq’s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groq’s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groq’s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a “preemptive acquisition” angle. The reporting identifies recent investors in Groq’s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groq’s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groq’s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groq’s compiler-led deterministic model is philosophically and practically different from CUDA’s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIA’s own rapidly improving inference performance complicates the “need” for Groq but does not eliminate the “option value.” NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIA’s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is “special” about Groq’s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows “hundreds” (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio manager’s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about “NVIDIA needs more inference speed” and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 просмотров • 7 месяцев назад

$AMD is easily a $1,200 stock IMO| CPUs TAM 🧵 Not Financial Advice! DYOR! In this thread, I want to discuss the actual TAM for CPUs data center for just 2026, where many are giving different ranges, where I don't agree with. I will explain in detail why I disagree with these research firms and financial analysts using Math. And this thread should not be treated as Financial Advice. I'm just explaining my research and thought process so we can have a discussion. In 2024/2025, I gave out $620 PT for FY2026 was too conservative for AMD potential. At the time, It was early and many were just laughing, that PT was unrealistic and the AI world is run on GPUs only. Today, most of these folks are laughing with me. That is ok, I dont offer financial advice, and I do not need everyone to agree with me. I respect other opinions. If you enjoy this kind of thread, slap the like/repost/bookmark. If you want to support my work further and gain more in-depth analysis, consider subscribe! In early 2026, hyperscalers, enterprises, and OEMs are scrambling as Intel and AMD server CPUs are largely sold out for the year, with prices jumping 10–20% and lead times stretching from weeks to months (or longer for certain SKUs). What was once a GPU dominated story has flipped: the shift to explosive Agentic AI with its multi-step reasoning loops, tool calling, multi-agent orchestration, real-time data movement, and reinforcement learning, is dramatically tightening CPU:GPU ratios from the old training-era 1:4–8 all the way to 1:1 to 5:1 or even CPU-heavy configurations. CEOs across NVIDIA, AMD, Intel, Google, Meta, Microsoft, and public companies have been sounding the alarm on CNBC, Bloomberg, and earnings calls. CPUs are “cool again,” and in many agentic deployments they are becoming the new bottleneck alongside (or even ahead of) GPUs and custom ASICs. In 2025, roughly 12-15m AI GPUs + AI ASICs GPUs shipped, and is expect to be 15-20m units by 2026, where it suggesting Training demand is not going away. The actual TAM is structural, multiplicative demand that has already forced AMD to double its long-term server CPU TAM forecast to >$120 billion by 2030 (>35% CAGR), with Dr. Lisa Su noting Q2 2026 server CPU sales expected to surge 70%+ year-over-year and demand “far exceeding expectations.” At the same time, AMD’s secured 30–40% share of TSMC’s initial 2nm capacity (behind only Apple’s >50%) positions it to ramp Zen 6-based EPYC Venice exactly when this agentic wave hits hardest but even that aggressive five-fab 2nm expansion (with plans scaling toward 11 total advanced facilities) cannot instantly close the gap in the near-term. Supply constraints on wafers, advanced packaging, and power are compounding the squeeze, just as hyperscalers forward-buy and lock in long-term deals. 1. The actual potential TAM Various sources and institutions are giving $50-$160-$200B CPUs TAM toward 2030, and i disagree, where supply is severely behind vs Demand by at least 2-3 years or even longer by some estimates. The actual TAM will probably be 15-20m for FY2026. The typical average selling price from low to high end is $5,000 to $15,000, but due to rising memory, and different inflationary pressures on Semi, it would be more logical to think between $7,000-17,000. A. CPU:GPU Ratio at 1:1 A basic calucation at mid range =12,000 x 15-20m CPUs= $180-$240B TAM B. CPU:GPU Ratio at 5:1 = $12,000 x 75m-100m CPUs= $900B-$1.2T TAM Of course TSMC cannot even supply 20% of this massive inflection TAM in 2026. But do we think of Demand for TAM or Supply for TAM? Hence we are seeing massive 2nm Ramp from TSMC for $AMD. IMO, conservatively, I would take down 15-20% on 1:1 or $135-$192B TAM for just 2026. Im not even talking about 2030. We are just months into this, it is impossible to estimate Cagr atm, but this is 1-5 agents running tasks, I wrote a thread on 24/7 autonomous agents thread, where companies could use 50-250 agents to run tasks for them 24/7. It would require a different structural CPU:GPU to bring down the cost of token as well as handling the Orchestration bottleneck. GPUs would be useless and sit idle waiting for CPU due to highly CPU-intensive nature. The cost per Million tokens must come down more rapidly for this 50-250 autonomous agents to work, otherwise the token cost would be too enormous. Helios Rack is estimated to bring inference cost down to $0.0003-$0.0005/M tokens with 18 EPYC Venices along with 72 MI455x and other chips+ Components. A heavier or CPUs dense rack would bring down inference cost further. EPYC Verano(2027 gen 7 AI-optimized) is expected to drive inference costs meaningfully lower than the Venice baseline likely to the $0.00002–$0.00025 per million tokens range (or even sub-$0.00015 in highly optimized agentic/batch workloads). Verano have higher core counts than Venice, LPDDR5X SOCAMM2 memory support, more AI optimized and Next-Gen rack density & efficiency. 2. $AMD secured at least 30-40% of TSMC 2nm capacity and Memory from Samsung through 2028-2030. 2 2nm fabs are entering ramping phase toward 60-65k wafers per months and 5 dedicated 2nm fabs entering mass production/ramp in 2026. Will link sub threads below if you are interest for full detail. Apple is reported to secure 50%+ 2nm capacity for Iphone 18 and Mac chips and AMD secured at least 30-40% capacity while $NVDA $AVGO $ARM $AMZN $GOOGL and others are on 3nm. This broader aggressive ramp from TSMC to target up to 11 fabs is to address $AMD massive growth ahead. Where $ARM is facing massive CPUs supply constraints as they have to compete with other Mega Cap players on 3nm allocation. And $INTC is also facing supply constraints for data center CPUs and PC per management with lead times extrended to longer than 12 weeks. Dr. Su is aiming for higher than 50%+ Market share, and I believe it is achievable in 2026 or 2027 as AMD has the strongest CPUs offerings. Dr. Su did not want to take advantage of the shortage and she said during the Q1 earning call, AMD is prioritizing Units shipped while guiding margin to be inching 60%. If Jensen were in charge, I'm sure margin would be 70-75% in this kind of severe CPUs shortage condition. But that is not how Dr. Su operates for more than a decade. She wants most market share. So we will see it in revenue growth, but as TSMC ramps faster and faster, AMD Operating and FCF margin will massively improve vs prior decade. A significantly higher margin profile than before. 3. How I came up with $1,200 withint 12-18 months? At $1,200/ share, that would be around $2 Trillion MC. I expect FY2027 revenue to be $124-$144B where data center revenue dominates overall revenue. AI GPUs: I will stick to the lowest end so show u that I'm conservative at $18B for each GW vs $NVDA Rubin is $30B+ (most likely Helios Rack in the $20B+ due to memory price rising). We know deals with OpenAI and Meta are around 12GW and additional multi-customers at multi-GW scale were hinted and will be revealed as we get to July 22-23 2026 Advancing AI event. For now I will conservatively add a bit more to this model. (3-6GW Helios Rack Range) EPYC Venice is reported to be in $15,000-$20,000. However large customers will likely to enjoy $10-$12k discount. I expect AMD to be able to ramp 7m EPYC Venice for entire 2026 and 3-4m of EPYC Verano(higher price than Venice). If we take an average selling price of $10,000 to be on the conservative side. Take down another 30% to be even more conservative on projection. I like to be conservative. That would be ~ 7m EPYC CPUs(Venice + Verano) for FY2027 or 583,000 units per month or 15,000 additional 2nm wafers per month which is completely reasonable for current TSMC Ramp, and I may be too conservative here. EPYC Verano and MI500 series will also be on 2nm. AI GPUs: 3GW x $18B= $54B EPYC CPUs: $10k x 7m CPUs= $70B = Data center revenue alone is $124B Other segments= probably in the $20-$25B FY 2027. FY2027 revenue = $124-$149B At 7m EPYC CPUs for entire 2027, that would be more than 50% market share when we comp it to availability from supply side, not from total Demand. It is possible that TSMC could significantly ramp even more capacity in 2027, so we will see. Metric Q1 2026 FY2027 Gross Margin 55-56% 60-62% Operating Margin 25-26% 32-35% Net Income Margin ~22% 26-30% FCF Margin 25% 28-30% At $124-$149B Revenue FY 2027 Net Income would be $32-$44B EPS would be $20-$27 (GAAP) Non-GAAP would be $25-$31 At $1,200 a share or $2T valuation that would be: 13.4-16x Price to Sales (P/S) 38-48 P/E At this kind of growth of AI SuperCycle, I think it is very reasonable valuation. If we use today at $406/share or $661B MC: 2027 P/S = 4.4x-5.3x 2027 P/E = 13x-16x Is AMD today expensive or cheap to you? Above is already a very conservative where I trimmed 20-30% of doable units. Meaning, there could be upside if TSMC is able to ramp meaningfully like they are planning. Conclusion: A $1,200 per share valuation IMO for AMD in FY2027 is not expensive at all; it is, in fact, conservative when viewed against the structural explosion in agentic AI demand we have mapped out. With server CPU TAM potentially scaling into the $100–$200B+ range in just CPU:GPU 1:1 Ratio for just 2026. AMD positioned to capture 50%+ share thanks to its 2nm TSMC allocation advantage and full-stack leadership, the company could realistically deliver $124–149B in total revenue and $25–$31+ non-GAAP EPS. At those levels, $1,200 implies a 2027 P/E = 13x-16x. Entirely reasonable for a company that will have become the clear Inference Queen (and in many workloads the preferred) AI infrastructure provider, with operating margins expanding above 30% and tens of billions in high-margin rack-scale AI revenue. Dr. Lisa Su was right presciently so about the Agentic AI inflection all the way back to her early 2022–2023 commentary on the coming shift from pure training to inference and orchestration-heavy workloads. While the broader market only fully woke up to this in 2026 when she doubled AMD’s long-term server CPU TAM forecast to >$120B by 2030 (with >35% CAGR), Dr. Su and her team have consistently positioned the company at the center of the CPU renaissance. The explosive demand we are seeing today, sold-out lines, rising ASPs, and hyperscalers forward-buying entire gigawatts of Helios-class systems is exactly the outcome she forecasted years ago. Not Financial Advice! DYOR!

Mike

301,322 просмотров • 2 месяцев назад

E166: Gracy Chen @Bitget - Single Mom. 2200 Employees. Top-5 Exchange Gracy Chen is CEO of Bitget - a top-5 global crypto exchange with 25 million users and $20B in daily trading volume. She went from TV journalist to MIT MBA to co-founding a unicorn startup to running one of the world's largest exchanges. She's also one of the few female CEOs in crypto who's been told directly by investors they won't back women who are married but don't have kids yet. So she built her own path! Timestamps: 0:00 Introduction 1:34 Please Subscribe 2:00 The Only Bad Experience With When Shift Happens 4:37 How Do We Motivate More Women To Join The Crypto Space 6:13 Bitget’s Overall Work Environment Explained 7:06 How Gracy Manages 2,200 Remote Workers 10:09 Partnerships: Jupiter KAST 10:50 Who Is Gracy Chen? 13:04 Where Does Gracy’s Fire In The Belly Come From? 15:01 Why Is It Necessary To Go To The Best Schools 17:00 How Does Happiness Connect With Being The Best? 19:38 How Did Gracy End Up As A TV Host 23:45 The Hardest Part About Building A Business 26:10 Want-trepreneurs vs Entrepreneurs 28:52 Partnerships: Ethena Sumsub 29:52 Just Go and Do It 32:57 Why Investors Didn’t Invest In Gracy’s 2017 Project 34:55 Why Should People Care About Crypto Currencies & Blockchain 41:36 What Is Bitget, Explained Simply? 43:55 Gracy’s Breakdown Of Bitget’s Numbers 45:09 How Bitget Got Over 40% Of Its Management Team To Be Female 46:48 Where Women Excel In Terms Of Business Management Practice 50:45 Where Do Men Excel In Business Management Practice 52:20 Partnerships: Trezor Bitwise Sui 53:17 What Are Gracy’s KPIs As The CEO Of Bitget? 56:45 How Bitget Is Growing The Pie To Continue Expanding 59:31 What Can An Exchange Do To Capitalize On Stablecoin Growth 1:01:32 How Much Of Gracy’s Life Is Run On Crypto 1:03:40 What’s The Endgame For Bitget? 1:05:49 Why Can Anyone Become A CEO? 1:09:28 Running A 2,200 Employee Company While Taking Care Of A Kid 1:11:22 How Has Bitget Evolved Since Gracy Joined In 2022 1:13:25 Something Gracy’s Holding Onto That She Knows She Should Let Go Of 1:15:01 What She’ll Teach Her Son To End The “Failed Marriage Curse” 1:17:12 What The Voice In Gracy’s Head Tells Her In The Morning & Evening 1:20:05 One Thing Gracy’s Learned That You Can Take With You 1:24:31 Closing Thoughts

MR SHIFT 🦁

104,369 просмотров • 3 месяцев назад

77 Reasons Why I’ve Invested Over $8,000,000+ in MultiversX (EGLD) and Why EGLD Will Crush It in 2025 (My Investment Thesis). I publicly shared my portfolio on X. EGLD is A) Better than BTC B) Everything that ETH wants to be C) The GameStop of Crypto 1. EGLD is verifiably the most scalable (theoretically unlimited) L1 chain in the world, theoretically capable of over 10 million TPS (thanks to adaptive state sharding). 2. e-Gold is digital gold. It has the best tokenomics among all L1s, similarly scarce to BTC, with a maximum supply of 31.4 million coins. Currently, 27.68 million coins are in circulation. 3. EGLD will be the most decentralized cryptocurrency in the world thanks to sharding and minimal hardware requirements for running nodes. It’s already second only to Ethereum with 3,618 validator nodes. 4. EGLD has extremely low fees, around ~$0.002 per transaction. 5. EGLD is extremely secure. No wallet drains like on ETH/SOL; assets are owned natively (not via a smart contract). There is no MEV risk (front-running bots). 6. EGLD is the only chain in the world with an on-chain Guardian (two-phase verification), making it impossible for a hacker to steal your funds—even if they have your private keys (seed phrase). 7. EGLD is carbon-neutral and eco-friendly, not wasting energy like BTC and other PoW chains. It’s exceptionally efficient, scalable, global, and sustainable. 8. EGLD has the best UX in crypto. Download the xPortal wallet—it’s like discovering Apple in Web3. The interface is simple, flawless, and you barely realize you’re using crypto. Instead of addresses, you use HeroTags. The app features all dApps, everything runs smoothly, and the visuals are beautifully designed. The explorer, web wallet, etc. follow the same high-quality user experience. 9. EGLD supports native assets, unlike Ethereum, for example. 10. EGLD is the first chain to fully implement horizontal (theoretically unlimited) sharding without compromising on decentralization—unlike Solana and others that attempt vertical scaling, leading to multiple network downtimes (11+ times) and huge hardware demands for validators, ultimately harming decentralization. 11. EGLD makes setting up a validator agency extremely easy. Even complete IT beginners can do it. The UX and documentation are superb. I personally set up the “EGLDSqueeze” agency in about 30 minutes. Managing it is straightforward via the web wallet, which feels like managing a Facebook page. This simplifies decentralization enormously. 12. EGLD allows literally anyone (even your grandma) to participate in decentralization, since nodes can run on a Raspberry Pi or a relatively affordable phone. Imagine millions of people worldwide securing the network, validating transactions without even knowing it. This can’t be done with BTC, where setting up profitable mining operations is prohibitively expensive. 13. WASM-Based Virtual Machine: You can write smart contracts in your favorite language, compile them, and run them via the fastest VM in the world. 14. EGLD has been tested at an incredible 263,000 TPS using its sharding mechanism and low hardware requirements. Allegedly, by mid-next year (April), they’ll demonstrate 1,000,000 TPS. (For context: Mastercard handles around 5,000 TPS; BTC handles 5–7 TPS.) 15. EGLD is currently the most advanced L1 in terms of scalability, security, decentralization, UX, eco-friendliness, and tokenomics. It’s the only chain that has genuinely solved the Blockchain Trilemma and is ready to onboard 1 billion people into crypto—users who won’t even realize they’re interacting with crypto. 16. EGLD is perfectly positioned for AI projects—AI agents, AI tools, or a so-called “Truth Machine” that monitors other AIs on-chain, documenting what’s true and comparing different AI outputs (some of which may be censored or biased), ensuring people don’t get confused or scammed in an AI-driven world. 17. The EGLD team is the hardest-working team I’ve ever encountered. I had the honor of meeting many of them personally, and can attest that their pace—even during a bear market—is extraordinary. 18. EGLD’s development team is exceptionally active on GitHub, continually improving their network and actively committing code. 19. EGLD plans to introduce an update reducing block time to 600ms (down from ~6 seconds), which would make the chain essentially unrivaled. 20. EGLD is effectively the only usable L1 in Europe, and the team has direct connections within the EU government—extremely bullish for the project. 21. EGLD provides top-tier on-chain governance not only for the MultiversX (EGLD) protocol but also for DeFi projects (e.g., xExchange, MEX). 22. EGLD plans to expand to the US, likely opening offices in Austin, Texas. This could put them in direct contact with Elon Musk (if it hasn’t happened already), as he’s involved with If he’s done his research, he’d discover there’s simply no better L1 worldwide. 23. EGLD solved fully implemented sharding, perfect tokenomics, and top-tier architecture with just $5M, whereas other chains failed to do so even with $100M+. The second-best sharding network, NEAR, needed $100M, has worse tokenomics, and its sharding isn’t fully implemented yet. Its UX also doesn’t compare. Owning NEAR was like comparing a VW Golf R to a Porsche GT3—EGLD is the Porsche GT3. 24. According to Similarweb, EGLD has significantly high traffic relative to other chains with market caps 100x larger. The market cap vs. web traffic discrepancy is huge, which is a strong indicator of EGLD’s potential. 25. EGLD has the most active and dedicated community relative to its user base, with users who believe in the technology, have full faith in the team, and remain loyal despite price volatility—because they use the chain and know there’s nothing better. 26. Check other chains’ active user counts on X (Twitter) and compare it with the followers of EGLD’s founders and main network accounts, versus those with 30x, 50x, or 100x larger market caps. 27. Visit the MultiversX website to observe the futuristic design and presentation, then compare it to other chains that appear nearly a decade behind in design and branding. 28. EGLD hosts the xDay Global event, showcasing updates, new builders, projects in the ecosystem, and major announcements—similar to Apple’s Keynotes—delivered in a highly professional, goosebump-inducing atmosphere. The next event is in Korea, the second-biggest crypto market after the US. Check out their previous xDay after-movie to see why this is extremely bullish. 29. EGLD is moving forward with plans for the first regulated, audited EU stablecoin under MiCa regulation, made possible by acquiring xMoney, which I view as a “Stripe” for crypto/fiat, offering everything from user solutions to merchant services—potentially the future of payments. 30. Greg Siourouni recently joined EGLD, having been an executive director at SUI Foundation. He’s now co-founder of xMoney Global. xMoney (formerly UTrust, with token UTK) is owned and founded by the MultiversX Labs team. A stablecoin might be introduced soon, which would be massively bullish given xMoney’s roadmap. They recently announced integrations with Binance Pay—both ways. 31. EGLD prioritizes user safety, believing it’s the only feasible approach once the network scales to serve a billion people—many of whom are retail users with little to no security awareness. 32. EGLD offers “Sovereign Chains,” letting you effectively clone their chain without heavy development, set up your own validators, and leverage their unlimited scalability. Any blockchain (ETH, BTC, SOL) struggling with scalability, decentralization, or security could run an ultra-fast, scalable, and secure L2 on EGLD’s Sovereign Chain, meeting top enterprise requirements. No one else has really done this. The Sovereign Chain demo achieved astonishing TPS and has an SDK. 33. No downtime since inception. 34. No shard takeover attacks have occurred. 35. Extremely fast—soon 600ms block time will be in place. 36. ESDTs – The best token standard available: fungible, non-fungible, semi-fungible, DeFi assets—everything is native and highly customizable. 37. Top-tier composability of assets and smart contracts. 38. Integrated DNS at protocol level with HeroTags (nicknames) instead of long addresses. 39. Asynchronous calls are supported. 40. Cross-shard transfers, execution, reverts, and calls are seamlessly integrated. 41. The best staking system in the space. Secure Proof of Stake (SPoS) is far more efficient than Proof of Work (PoW). 42. Built-in Delegation and Staking Provider system, with over 125K delegators. 43. Complete support for liquid staked assets, fostering decentralization rather than centralization. 44. TransferRoles for ESDT and other advanced operations. 45. Composable tasks on-chain for more sophisticated DeFi workflows. 46. MultiTransfer and asset execution within one transaction. 47. Re-entrancy protection is built-in by design. 48. Storage for ESDT assets goes beyond a linear approach, optimizing performance. 49. No integer overflows thanks to integrated safeMath operations. 50. Integrated crypto opcodes in the VM, enhancing security and performance. 51. Support for BigFloats, BigInts, and BigDecimals, enabling advanced financial calculations on-chain. 52. No sandwich attacks, plus front-running and MEV protection. 53. Relayed Transactions, simplifying user interactions and fees. 54. Smart Accounts featuring data tries and multiple built-in functions. 55. Generalized Paymaster solutions, enabling flexible fee models. 56. Subscriptions for recurring or automated on-chain payments. 57. Web2-like usability with Web3 functionality, bridging mainstream adoption. 58. StakingV4 for improved decentralization. 59. Enhanced MEV protection rolling out to safeguard users. 60. Parallel execution is coming soon, boosting throughput. 61. 1 million TPS is on the roadmap, targeted for demonstration. 62. 600ms block time is also coming soon. 63. Reduced cross-shard processing is planned to improve efficiency. 64. ZK everywhere (PI²): “prove everything” approach is coming. 65. AsyncV3 is in development for more complex cross-contract interactions. 66. Scalability enhancements for Merkle Tries or a new data model are being explored. 67. Linear storage on the VM is forthcoming. 68. A dynamic language interpreter at the VM is also planned. 69. Rumors suggest that MultiversX (EGLD) is building a “Truth Machine” on their L1—an essential, game-changing tool for AI verification and societal impact. 70. The entire team features individuals with PhDs in mathematics and physics, and many are former engineers at Google, IBM, and similar companies. 71. Over 56% of the network’s supply is staked, showcasing strong community involvement. 72. More than 6,772,347 accounts have been created on the network. 73. A total of 476,627,710 transactions have been processed on-chain without any outages or hacks. 74. EGLD has built a massive ecosystem over time. While not as numerous in project count as Solana, its market cap is ~100x smaller, yet it has far superior tokenomics and technology. The projects that do exist, like Hatom Protocol, are top-tier in UX, security, and advanced features. Hatom will soon introduce USH, a truly high-quality, decentralized stablecoin. 75. On competing chains, automated transactions aren’t easily or cheaply executed, whereas on MultiversX, tools like let you do this for free (with near-zero fees). 76. No other chain combines such a strong team and long-term vision where every product meets extreme security and UX standards like MultiversX does. This is why I see it as the “next Apple” in Web3. 77. MultiversX has a new CMO – Adam Bates, a former CMO at the Cardano Foundation. He was behind the success of Cardano’s huge marketing campaign and has a very good relationship with Charles Hoskinson. Thanks to him, Beniamin Mincu (the founder of MultiversX) was likely introduced, and now they will probably discuss how both blockchains can help each other, as well as any other potential collaborations we don’t yet know about. This is also extremely bullish. #EGLD is undeniably the most Scalable, Advanced, Secure, and User-friendly L1 supercomputer ever created. It’s built to SHAPE THE FUTURE. 1) 2) 3) 4) 5) 27/6/2024 - EGLDSqueeze - SUMMARY: HERE IS NO 2ND BEST. EGLD IS ONLY ONE BLOCKCHAIN THAT CAN RULE THEM ALL. ✅ UNLIMITED SCALING ✅ SCARCE AS BTC ✅ PROGRAMMABLE AS ETH ✅ NO DOWNTIME AS SOL ✅ UI/UX OF Apple ✅ SHARDING DONE BEFORE NEAR & TON ✅ BEST WALLET xPortal WITH GUARDIAN Price prediction (NFA|DYOR): My reasoning is that the real market cap as of December 23, 2024...if we take into account the value of other cryptocurrencies such as BTC, SOL, ETH, AVAX, NEAR, TON, Cardano, BNB, XRP, and so forth, plus the existence of meme coins with valuations above 20 billion USD, or even games nobody plays anymore that still have valuations above 800 million shows that EGLD’s current market cap of approximately 942 million USD is incredibly low. From a technological standpoint, user experience, and other relevant aspects, compared to SOL, NEAR, TON, AVAX, and other L1 protocols, EGLD’s market cap should realistically be around 100 billion USD. Therefore, my prediction and investment thesis is a minimum of a 100x increase from its current price (+-SOL marketcap). MultiversX is ready to onboard 1 billion people to the blockchain. From a long-term perspective, it could even reach a market cap of 1 trillion USD, which is roughly half of where BTC is right now. That would be approximately a 1060x gain from the current market cap. 1 EGLD (MultiversX) is for $34 (only 31.4M max supply) think about this. Not financial advice. Again. There is no 2nd best L1. Position yourself where the puck is going, then wait at the goal until the goal gets there Apes together, strong. Ape alone, weak. We Don't Worry. We Just Win. Shape The Future

Daniel Veroc

50,029 просмотров • 1 год назад

The July 4th weekend All-In The All-In Podcast turned into a long argument about who owns the intelligence layer. The besties think enterprises just woke up to a trap they had been walking into, here's how the conversation went (save this): ◽️ The Palantir-Nvidia deal is a bet against the model-layer duopoly. Palantir will use Nvidia's Nemotron open models to build a custom frontier-quality model for US government agencies, and the agencies own the hardware, the data, and the weights. Sacks framed it as structural: an application company and a chip company both want a competitive model layer, so they are natural partners against a two-provider middle. ◽️ Alex Karp's CNBC "crashout" was actually the thesis. Karp argued enterprises have lost trust in the frontier labs and want to own their compute, models, data, and alpha. Sacks translated it as a new definition of enterprise AI safety: safety means the model provider cannot hoover up your proprietary knowledge and turn it into its next product. ◽️ Figma is the cautionary tale that made it real. Anthropic launched Claude Design into Figma's category, its chief product officer sat on Figma's board and resigned only 3 days before launch, and Figma's stock is down about 50% this year while Anthropic's valuation surged. Sacks listed Claude Science, Security, Legal, Financial, and Code as the same move: dominate the model layer, then take the lucrative verticals. ◽️ The playbook has a name, and it is Microsoft and Google. Sacks argued Anthropic is running the operating-system strategy: own the layer everyone builds on, then walk up the stack. His Google receipt is that fewer than half of searches now send you off-site, versus an early Google that prided itself on how fast it kicked you away. ◽️ The BCG number is what raises the stakes. Chamath cited a BCG return-on-capital-employed study: the cost of capital is back to its long-run 8 to 11%, and half of large US companies cannot earn returns above it. If you are already teetering on your cost of capital, handing your alpha to a provider that may compete with you is not a luxury risk, it is fatal. ◽️ The 16.4x number is the whole argument in one data point. Chamath ran a code-migration task through 8090's harness. Wrapping Claude was 1.4x cheaper and 1.5x faster than Claude Opus alone. Wrapping the best open-source model was 16.4x cheaper, at about 3x slower. For a background task, three extra hours to cut cost by 16x is not a close call. ◽️ Even at 100x cheaper, enterprises were saying no for the wrong reason. Chamath relayed an ex-Meta PM's point that companies reject open models over China and safety fears, when they could host those same open weights on their own GPUs in US data centers with nothing flowing back. The safety objection, she argued, is backwards: the leak is the data you hand the frontier labs. ◽️ Friedberg says the frontier labs are trying to commoditize their own customers. Anthropic has been signing up life-sciences companies to feed a new life-focused model in exchange for early access, and nearly everyone he has talked to now refuses, recognizing that data they spent billions generating becomes worthless once it is pooled with everyone else's. ◽️ The deployment topology is shifting from big hubs to distributed spokes. Friedberg's map: the old assumption was a few capital-advantaged mega-clusters plus inference clouds. The new one is large hubs, medium hubs (enterprise training clusters), and distributed spokes, including on-prem inference in your own building. Owning your weights is the point. ◽️ Chamath's endgame is running GLM himself. An industry contact told him that with harness post-training and telemetry, an open Chinese model like GLM could get as good as Anthropic's Mythos. His conclusion: take GLM, control it soup-to-nuts on US hardware with only US citizens touching it, and pay a fraction. ◽️ The Apple analogy sharpens why renting intelligence is different from renting distribution. Chamath argued Apple is the only platform that respected developers, deliberately keeping its stock apps basic to protect the ecosystem and collect its 30% tax. There is no 30% tax on open models, and worse, you cannot rent intelligence from the same place that rents it to your competitor without ending up identical to them. ◽️ Nvidia's open model is now good enough to matter. Calacanis claimed you cannot tell Jensen Huang's Nemotron from Claude on 95% of searches, and that Nvidia downplayed the model until now to avoid alarming its top customers. The gloves came off once OpenAI, Anthropic, and Elon all signaled their own silicon ambitions. ◽️ Sacks sized the duopoly: roughly $60B and $40B in ARR. Anthropic is around ~$60 billion of ARR, OpenAI at ~$40 billion, and no one else generates meaningful model-layer revenue. Sacks's policy line: the US does not ban monopolies, only anti-competitive tactics, but the government should do nothing to make the duopoly more likely. ◽️ The token deflation call: 90% a year for three years. Calacanis predicted token costs fall 90% annually for three years, putting the price of intelligence near free and making it rational to waste tokens on hardware you already own. Friedberg's version is a 70/20/10 split between big cloud, local, and other clouds. ◽️ A wave of platform lock-in spending is already landing. Calacanis flagged Microsoft standing up a roughly $2.5 billion forward-deployed-engineer effort and Amazon spending about $1 billion on the same, plus OpenAI's version. His read: enterprises will slam the door, because letting a provider's engineers study your business is how it ends up in their model. ◽️ The server-per-employee prediction. Calacanis expects every employee to get $10,000 to $20,000 of local compute, a Mac Studio or a high-RAM Dell, running a personal local model that syncs to a thin laptop. A server per person, so nothing leaks. ◽️ On jobs, the data does not show present-tense loss. Sacks cited a RAMP and Revelio Labs study of over 21,000 US firms: the heaviest AI spenders grew headcount about 10% over two years, and entry-level headcount grew even faster at 12%. Friedberg's harder claim: there is no AI job loss yet, only clunky, gradual value creation, and the media will not reverse its narrative because that destroys its credibility. ◽️ The displacement case is real but forward-dated. The counterpoint on the show was that customer support, entry-level data entry and BPO, and driving are the near-term displacements, with Waymo cited as present-tense evidence: in markets where it hits critical mass, Uber and Lyft stop recruiting drivers. Sacks noted most US entry-level support was already offshored, so the acute risk sits in those countries first. ◽️ The human-premium counternarrative. Friedberg argued that as automation spreads, human interaction gets a premium: the skilled bartender, the real driver, the human-in-the-loop tier. He cited the company (referenced as Klarna) that hyped replacing its whole support team with AI, then reversed a year later on brand grounds. ◽️ The export-control episode needed three conditions, and Sacks says do not over-read it. Commerce lifted controls on Anthropic's Fable 5 after two weeks, with Mythos 5 restored to US customers around June 26 once co-founder Tom Brown replaced Dario as lead negotiator. Sacks's three conditions: Dario boasting for months about a cyber weapon, Amazon reporting failed guardrails in testing, and Dario refusing to roll Fable back. His message to allies: this was a particular set of circumstances rather than the debut of a standing lever. ◽️ The import question nobody answered cleanly. Calacanis pressed on why the US blocks Chinese cars and drones but not Chinese open models like DeepSeek and Kimi. Sacks's answer: a forked open model run on US hardware stops being Chinese, and banning open source would isolate the US and impose a token tax on American enterprises, so let the market decide if American open models win. ◽️ The California fiscal story is a business-climate story. Friedberg walked through the numbers behind Newsom's "balanced" $351B budget: expenses exceed revenue and $20-40B is borrowed to close the gap, the budget grew 65% in six years ($215B to $355B), personal income tax is $142B of ~$211B revenue with the top 1% (150,000 people) paying $70B of it, and the corporate rate of 8.9% sits far above Texas at zero. ◽️ The tax base is leaving, and the state is now taxing everyone else. Friedberg cited 1 to 1.5% of adjusted gross income leaving each year (about 15% over a decade), at least 15 Fortune 500 HQs and ~2,100 firms gone since 2019, and a new 8% software sales tax hitting Word, Gmail, and ChatGPT subscriptions plus a health-insurance tax, on top of a now-permanent 14.4% top bracket. The liabilities behind it run $1.4T in debt, up to $1.5T in unfunded pensions senior to state bonds, and ~$40B/year in out-year deficits. Lastly, the line that framed the whole show: "You can't rent intelligence from the same place that rents it to your competitor." That is the sovereignty thesis in one sentence, and every number in this episode is an argument for it. ____ Follow Fireside Alpha for more summaries on key business and technology conversations.

Fireside Alpha

55,334 просмотров • 24 дней назад