Loading video...

Video Failed to Load

Go Home

Huawei Cloud aC8, the latest general computing-plus Elastic Cloud Server (ECS), redefines performance for compute-intensive workloads such as film rendering, gaming, and AI inference. With an impressive capacity of up to 384 vCPUs, aC8 maintains stable performance even under extreme workloads. Built with unmatched scalability and flexibility, aC8 supports...

21,239 views • 11 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Today we announced our new Fairwater datacenter in Atlanta, connected with our first Fairwater site in Wisconsin and our broader Azure footprint to create the world’s first AI superfactory. Fairwater exemplifies our vision for a fungible fleet: infra that can serve any workload, anywhere, on fit-for-purpose accelerators and network paths, with maximum performance and efficiency. AI workloads have evolved beyond large-scale pre-training. Today, they encompass fine-tuning, reinforcement learning (RL), synthetic data generation, evaluation pipelines, and more. Fairwater is built to support this full lifecycle: Max density: Fairwater’s two-story design and liquid cooling system lets us place racks in three dimensions and pack them with GPUs as densely as possible, minimizing cable runs and improving latency and effective bandwidth. Fleet: Each Fairwater DC can integrate hundreds of thousands of the latest NVIDIA GPUs into a single coherent cluster. This provides flexible infra that can support the full spectrum of workloads, and ensure no GPU is left unnecessarily idle. And that’s on top of the more than 100,000 GB300s coming online this quarter alone for inference across the rest of our fleet. For us, it’s all about turning every gigawatt into the maximum number of useful tokens. Not every GW is created equal! Planet-scale: Every Fairwater DC will connect through our continent-spanning AI WAN to prior generations of AI supercomputers, forming a truly fungible pool of compute. This enables developers to scale beyond the capacity of a single site and dynamically land workloads on the right infra for their needs. Together, these innovations let us bring together different generations of silicon and AI systems across DCs and geos into a single elastic system that scales seamlessly across training and inference workloads And this elastic AI capacity is all available alongside all the other cloud services (compute, storage, databases, app services) that AI agents and workloads need. This is what we mean when we talk about building a fungible fleet – a single, unified platform that pushes the limits of performance per watt and per dollar. Read more:

Satya Nadella

907,624 views • 8 months ago

$AMD $AMZN partnership will 🚀 in 2026 🔥 Amazon/AMD partnership is hidden among hot headlines from OpenAI $NVDA $ORCL... TLDR: Amazon refused to bid up the overpriced $NVDA chips among other hyperscalers, and decided to work closely with $AMD. Amazon is expected to spend up to $10-$20B a year on 2026 EPYC breakthrough Gen and Future Gen. Dr. Su confirmed "we have plenty for other large customers". For its 2026 EPYC "Venice" processors, AMD is using a multi-node manufacturing strategy: the CPU core complex dies (CCDs) are built on TSMC's 2 nm-class node (N2), while the I/O die (IOD) uses the N3P (3 nm) process. Context: Andy Jassy Amazon Web Services has been working with AMD on EPYC processors since November 2018. With this "secret weapon" breakthrough(patented), this long time partnership has expanded to New breakthrough 2026 EPYC Gen. AMD's 6th Gen EPYC "Venice" processors, slated for 2026, introduce New Chiplet design breakthrough. a revolutionary chiplet interconnect fabric that redefines server scalability for AI. This isn't just faster silicon; it's a paradigm shift for AWS, enabling hyper-efficient, rack-scale AI inference that slashes costs and latency while boosting throughput. AMD to benefit AWS's $100B+ AI opportunity along with $ORCL $MSFT $GOOGL $META Saudi, UAE ,38+ countries and startups. In early October, Amazon/AWS announced the new EC2 M8a instances as their latest-generation, general-purpose compute instances now powered by AMD EPYC 9005 "Turin" processors. Amazon announced the M8a as having up to 30% higher performance and up to 19% better price performance over M7a. With my testing of both at 32 vCPUs, the new AMD EPYC Turin instance provided 1.59x the performance over the prior-generation EPYC Genoa instance! How will this impact AWS AI Inference? ~Cost Efficiency: Inference is 80%+ of AI workloads and latency-sensitive (e.g., chatbots need <1s responses). "Secret weapon" enables 35x better inference perf (per AMD's CDNA roadmap tie-in), cutting AWS's energy use by 50%+ in clusters. With $118B 2025 capex, this could save $20–$30B annually in OPEX, boosting margins to 35%-40%. ~Scalability for Agentic AI: Supports "Helios" rack-scale platforms (up to 128 GPUs + EPYC hosts), delivering 3.58x FP6 perf for distributed inference. AWS can run 700K+ more tokens/sec in 1,000-node clusters (via EPYC 9575F boosts), enabling real-time apps like personalized search or fraud detection at enterprise scale. ~Adoption Catalysts: Early partners like Oracle signal broad uptake; AWS's existing AMD instances G4ad with Radeon GPUs) pave the way. By 2026, EPYC could power 40%+ of AWS AI infra, outpacing Nvidia's GPU lock-in via open standards (ROCm 8 software). Lastly, Amazon’s trajectory toward a $320 stock price is not a speculative leap but a grounded projection rooted in its unmatched fundamentals and strategic AI leadership. With Amazon Web Services poised to surpass $100 billion in annual revenue by 2026, driven by explosive AI inference demand, Amazon is redefining cloud computing’s future. The adoption of AMD’s 2026 EPYC processors with "Secret" architecture is a game-changer, slashing costs by up to 50% and boosting inference throughput 3x, enabling AWS to dominate enterprise AI workloads with unmatched efficiency. This technological edge, combined with Amazon’s e-commerce dominance and high-margin advertising growth, supports a valuation rerating to 22x EV/EBITDA, and it is still a discount to historical highs. Trading at $222, $AMZN is undervalued for its 15–20% revenue CAGR and 25%+ EPS growth through 2030.

Mike

511,082 views • 9 months ago

Interview with Nebius Co-Founder Roman Chernin Please like & share this video so that all $NBIS investors on X will see it! :) If you prefer watching on YouTube: Timestamps: 00:00 - Why AI Infrastructure Is So Hard to Understand 00:24 - Market Fragmentation and What Actually Differentiates Providers 01:30 - Consolidation, Segmentation, and the Future AI Cloud Landscape 02:56 - What Analysts and VCs Still Get Wrong About AI Infrastructure 05:34 - Nebius Cloud: Product Readiness and Customer Proof Points 07:42 - Why Inference Workloads Are Exploding 09:11 - Training vs. Inference: How AI Models Actually Reach Production 10:10 - Why Inference Market Share May Concentrate Around a Few Winners 12:36 - Customer Use Cases: Coding, Enterprise AI, and Real-World Adoption 14:01 - Why Integrated Training and Inference Matter Strategically 16:01 - Building Scalable AI Infrastructure With High Utilization 18:24 - Token Factory: Inference as a Managed Service 20:24 - Revolut Case Study: AI-Driven Product Enhancements 22:56 - Token Factory Performance Optimization and Competitive Advantage 25:07 - Scale, Capacity, and Efficiency as Growth Drivers 28:36 - Why Inference Capacity Could Become the Next Major Bottleneck 30:10 - How Nebius Benchmarks Performance Across Providers 33:14 - The Future Size and Shape of the Inference Market 36:38 - Value-Based Pricing: Moving Beyond Cost per GPU Hour 40:55 - How Nebius Wins Deals: Quality, Performance, and Customer Experience 44:53 - Autonomous AI Platforms and the Rise of Agent-Based Models 47:28 - Tavily, Agentic Applications, and the Next Layer of the AI Stack 50:45 - Strategic Trade-Offs: Scaling, Product Roadmap, and Customer Relevance 55:40 - Final Thoughts: Adapting to the Next Shift in AI Workloads Nebius Roman Chernin

Daniel Koss

203,706 views • 2 months ago

$AMD $MSFT Partnership is MASSIVE in 2026 🚀 If you were excited about my thread on $AMD $AMZN AWS long time partnership, you will be even more excited about what Microsoft gonna do with 2026 AMD EPYC "Venice". Historical Context: The relationship between AMD and Microsoft began in the early 2000s, with Microsoft initially focusing on Intel's x86 architecture for its Windows operating system and server products. However, AMD's entry into the server market with its Opteron processors in 2003 marked the beginning of a competitive dynamic that eventually led to collaboration. The partnership intensified with the launch of 3rd Generation EPYC "Milan" in 2021, powering Azure's N2D and C2D VM families. By 2025, Microsoft had integrated 5th Generation EPYC "Turin" into new compute-optimized instances, reflecting a strategic shift towards AMD for cost and performance benefits. This "Secret Weapon" breakthrough will mark another inflection point for AMD Microsoft Azure relationship, will probably be more aggressive than EPYC "Milan" moment in 2021. We can call it EPYC "Venice" moment 2026" 1. Technical performance of AMD EPYC "Venice" (2026) AMD's 6th Gen EPYC "Venice" processors, slated for 2026, introduce New Chiplet design breakthrough. a revolutionary chiplet interconnect fabric that redefines server scalability for AI. This isn't just faster silicon; it's a paradigm shift for Microsoft Azure , enabling hyper-efficient, rack-scale AI inference that slashes costs and latency while boosting throughput. ~Up to 256 Zen 6 cores, a 70% performance increase over "Turin," optimized for AI and HPC. ~Memory and Bandwidth: 1.6 TB/s per socket, doubling "Turin's" capability, with support for MR-DIMM/MCR-DIMM. ~Efficiency: 1,500-1,700W power draw, a 50% reduction, aligning with Microsoft's sustainability initiatives. ~Interconnect: PCIe 6.0 and a new chiplet fabric for rack-scale AI, reducing latency and enhancing scalability. 2. Why $MSFT will adopt $AMD YPYC Share to 50%+ in 2026. AMD EPYC Share: ~30-35% of Azure's x86 CPU-based business while Intel Xeon share is 65% Microsoft's Azure has been progressively integrating AMD EPYC, with "Venice" expected to expand this footprint: A. Dominance of AI Inference Workloads ~AI inference constitutes 80% of AI workloads in cloud environments, with latency-sensitive applications like chatbots, recommendation engines, and fraud detection requiring sub-second response times. ~"Venice's" 35x inference performance uplift directly addresses these requirements, outperforming Intel's offerings and custom Arm solutions in multi-threaded scenarios. B. Cost Efficiency and Operational Savings ~Azure's 2025 capex of $118B is under pressure to deliver returns. "Venice" can reduce operational expenses by $20-30B annually due to its power efficiency and performance gains, improving Azure's margins to 35-40%. ~The cost per inference operation is significantly lower with "Venice," estimated at 24-31% less than Intel-based alternatives, enhancing Azure's competitiveness against AWS and GCP. C. Scalability for Enterprise AI: ~"Venice" supports rack-scale AI deployments, enabling Azure to scale AI services for enterprise customers. For example, a 1,000-node cluster can process 700,000+ tokens per second, crucial for large-scale AI applications like personalized marketing and predictive analytics. ~This scalability is particularly important as Azure aims to capture the $100B+ AI opportunity by 2026, as stated by Microsoft CEO Satya Nadella. D. Reduction of Nvidia Dependency ~While Nvidia ( $NVDA) dominates AI accelerators, AMD's integrated EPYC-GPU solutions (MI450 with "Venice") offer a balanced approach, reducing Azure's reliance on Nvidia's high-cost GPUs. ~"Venice" enables hybrid inference models, where CPU-based inference handles 80% of workloads, and GPU acceleration is reserved for training and complex tasks, optimizing resource allocation. 3. Financial Implication: ~Revenue from Azure could reach $15-18B annually by 2026, part of a total revenue projection of $70-100B ~Profit margins could improve to 55-60%, boosting net income to $20-25B, supported by scale economies and reduced production costs. Intel could respond by giving more aggressive discounts, but this breakthrough has been a decade long of $AMD R&D, or rethinking chiplet design, a complete new approach. "Venice's" lead in AI inference and efficiency is challenging to match. Broader Industry: Other hyperscalers ( Amazon Web Services , GCP) and enterprises will follow Azure's lead, standardizing EPYC technology and pressuring Intel further. This could lead to a broader industry shift towards AMD, enhancing its ecosystem and bargaining power. Conclusion: The strategic adoption of AMD's 6th Generation EPYC "Venice" processors by Microsoft Azure in 2026 marks a pivotal moment in the evolution of cloud computing, particularly for AI inference capabilities. "Venice's" groundbreaking chiplet design, offering a 35x performance uplift for AI inference tasks, a 50% reduction in power consumption, and unparalleled scalability, positions Azure to leapfrog its competitors in the race for AI dominance. This technical superiority, combined with significant cost savings potentially $20-30B annually in operational expenses; aligns perfectly with Microsoft's ambitions to capture the $100B+ Revenue AI opportunity by 2026. The shift to 50% x86 market share for AMD within Azure is not merely a technical transition but a strategic realignment that redefines the competitive landscape. Historically, Microsoft's partnership with AMD has evolved from niche deployments to a core component of Azure's infrastructure, and "Venice" accelerates this trend. The 30-35% AMD EPYC share in 2025 is expected to double, driven by new VM families like C4D and H4D, which will dominate AI-intensive and HPC workloads. This migration is incentivized by "Venice's" efficiency gains, reducing dependency on Intel and Nvidia, and enhancing Azure's sustainability profile. Not Financial Advice!

Mike

141,018 views • 9 months ago

We are back again :) After three weeks of quiet building. Introducing Genesis World 1.0, our latest simulation platform, the second release in our full-stack suite. Open-sourced. Robotics is still bottlenecked by the 1× speed of the physical world. Every model, checkpoint, and data recipe eventually needs to be tested on physical hardware, slowly, expensively, and with limited coverage. One hour in reality can become 100 days in simulation. That is how robotics model iteration moves from a wall-clock bottleneck to a compute problem. To make this work, simulation has to be both fast and trustworthy. Over the past year, we rebuilt the entire stack: a GPU-accelerated cross-platform compiler, penetration-free multi-physics contact solvers, unified rigid and deformable physics, and a photo-realistic renderer purpose-built for physical AI applications. We built Nyx, a high-performance path-traced rendering engine for robotics application. Genesis World 1.0 achieves near realtime performance with our latest development for penetration-free IPC solver, supporting various types of deformables beyond rigid bodies. It supports contact-rich, dexterous manipulation simulation across different embodiments: unitree, sharpa, wuji, genesis hand and various types of grippers. Under the hood is Quadrants, our effort in pushing forward cross-platform GPU-accelerated computation. Quadrants started as a fork of Taichi, and we rebuilt most of the critical parts for optimizing simulation workloads, giving 10x faster launch time and up to 4.6x runtime performance compared to the initial Genesis release. Together, they bring us to an unprecedentedly low sim-to-real gap, enabling zero-shot real-to-sim model evaluation and much faster iteration of GENE. All available today. Genesis World 1.0: Quadrants: Nyx:

Genesis AI

308,392 views • 1 month ago

If intelligence is the log of compute… it starts with a lot of compute! And that’s why we’re scaling our GPU fleet faster than anyone else. Just last year, we added over 2 gigawatts of new capacity – roughly the output of 2 nuclear power plants. And today we’re going further, announcing the world's most powerful AI datacenter, located in southeastern Wisconsin. Fairwater is a seamless cluster of hundreds of thousands of NVIDIA GB200s, connected by enough fiber to circle the Earth 4.5 times. It will deliver 10x the performance of the world’s fastest supercomputer today, enabling AI training and inference workloads at a level never before seen. For AI training workloads, you need compute at exponential scale. That’s why we designed the datacenter, GPU fleet, and network together as one integrated system. This ensures a single job can run from day 1 at exponential scale across thousands of GPUs. Fairwater uses a liquid-cooled closed-loop system for cooling GPUs that requires zero water for operations after construction. And we’re matching all of the energy that is consumed with renewable sources. And of course, it is just one of several similar sites we’re lighting up across our 70+ regions. We have multiple identical Fairwater datacenters under construction in other locations across the US, in addition to our AI infrastructure already deployed in over 100 datacenters around the world, powering model training, test-time compute, RL tuning, and real-time inference at global scale. Too often during times like this, people go with the current and only later wonder, how did we get here? With Fairwater, we're charting a new path: doing the hard engineering work, bringing compute, network, and storage into one highly scaled cluster, and designing closed-loop energy systems to meet real-world computing needs. And partnering with local communities to ensure it's thoughtfully done in a way that is sustainable, creates new jobs, and expands opportunity. We are thrilled to see this take hold in Wisconsin, and we are just getting started.

Satya Nadella

2,022,537 views • 10 months ago

Hey everyone, today I want to introduce a project that’s aiming to redefine how we access compute for AI — it’s called GPUAI. 🔶 GPUAI: Unlocking Global GPU Power for the AI Era GPUAI isn’t just another GPU marketplace or leasing service. It’s a fully decentralized protocol that connects idle GPU resources around the world — from gaming PCs to data center clusters — and transforms them into a high-performance compute network for AI workloads. 🧠 Why does it matter? Right now, the biggest bottleneck in AI isn’t algorithms — it’s access to compute. Training and running models requires massive GPU power, but it’s locked up in centralized cloud platforms, expensive and hard to access for smaller teams. With GPUAI, anyone can tap into a global GPU pool that’s: ✅ Fully decentralized ✅ Reputation-based and smart contract coordinated ✅ Encrypted and secure ✅ Token-incentivized — meaning contributors get rewarded in $GPUAI 📈 For developers, it’s a flexible way to access GPU compute for training, inference, and more — without cloud lock-in. 💰 For GPU owners, it’s a chance to monetize idle hardware that would otherwise go unused. The protocol is live, the apps are active, and the ecosystem is growing fast. 🌐 Try it yourself at 📖 Learn more on 🎮 Play our community games at This is real infrastructure for the future of AI, not hype. Follow them and explore their mission of decentralized computing at Tell me what you think - if you have a GPU, you can start profiting now. #GPUAI #Web3Infrastructure #AIComputing #DePIN #Decentralization

The Crypto GEMs

69,984 views • 1 year ago

No single vendor will win the AI race, but open ecosystems might. Real velocity in AI comes from interoperability, not lock-in. And AMD just made all of its software open source. At last week’s Advancing AI 2025, we sat down with AMD’s VP of AI Software Anush Elangovan and Sharon Zhou VP of AI at AMD, to discuss their case for why an open, multi-partner ecosystem will accelerate AI innovation faster than any proprietary alternative. AMD’s announcements last week double down on this OSS focus and their commitment to AI infrastructure, including: ✅ Open Source Ecosystem: ROCm 7, AMD’s latest open-source AI software stack, introduces kernel-level improvements for GEMM operations, optimized attention mechanisms, and expanded support for distributed inference. The update brings substantial speedups for inference workloads, with average performance increases of 3.2x to 3.8x ✅ Hardware: New MI355X GPU delivers up to 40% more tokens per dollar vs competition & the MI350 Series has seen a 35x generational leap in AI inference performance ✅ Infrastructure Investments: Oracle just committed to zettascale (‼️) clusters with up to 131,072 MI355X GPUs and AMD showcased their new $10 billion partnership with Saudi Arabian AI firm HUMAIN to build AI infrastructure, including data centers, powered by AMD chips. ✅ Partnership Momentum: 7 out of 10 top AI companies now run production workloads on AMD Instinct accelerators (including Meta, OpenAI, Microsoft & xAI) By inviting interoperability and contribution at every layer, AMD is enabling developers to build faster, optimize deeper, and deploy with flexibility. Listen to Anush and Sharon’s Chain of Thought Podcast episode with host Conor Bronsdon in the next tweet to get all the details and a deep dive into AMD’s strategy 👇

Galileo

78,922 views • 1 year ago

How have the fundamentals of building large, distributed software systems changed the last decade? A conversation with Martin Kleppmann (author of Designing Data-Intensive Applications) - given that the second, updated edition of the book was just released. Timestamps: 00:00 Early career 05:46 Building Rapportive 10:47 Working at LinkedIn 14:09 Writing Designing Data-Intensive Applications 23:00 Reliability, scalability, and repeatability 26:24 DDIA: the second edition 30:50 Tradeoffs of using cloud services 39:02 How the cloud changed scaling 42:53 The trouble with distributed systems 49:02 Ethics for software engineers 52:45 Formal verification 1:00:12 Academia vs. industry 1:03:50 Local-first software 1:09:50 Computer science education 1:18:32 Martin’s current research and advice Brought to you by: • Statsig – ⁠ The unified platform for flags, analytics, experiments, and more. • Sonar – The makers of SonarQube, the industry standard for code verification and automated code review. Check out Sonar's new architecture management capabilities that ensure both humans and AI agents respect your system’s blueprint. • WorkOS – Ship enterprise features – SSO, directory sync, RBAC, audit logs – in days, not months. Three things worth considering, as discussed with Martin, in this episode: 1. Multi-region and multi-cloud are risk/cost trade-offs, not best practices. Martin does not believe that there is a “best practice” in deciding whether to go multi-region or multi-cloud. This decision is a tradeoff between risk and costs. It’s a business decision to be made. Designing Data-Intensive Applications gives engineers the vocabulary to articulate the tradeoffs, not to dictate answers. 2. Replication for fault tolerance is more relevant for most engineers these days than sharding. Though the book has a full chapter on sharding, Martin said that the cloud has reduced the need for manual sharding for the majority of teams. This is also because machines are increasingly bigger, and more workloads fit on a single machine. Sharding across machines is increasingly a specialist concern; replication for fault tolerance, however, is still relevant at every scale. 3. Knowing system internals as a superpower for application developers. Martin maintains that Designing Data-Intensive Applications is not a book for people who build databases or even infrastructure, but it’s helpful for application developers to develop an intuition for making good design decisions and debugging performance issues we will eventually encounter.

Gergely Orosz

79,406 views • 3 months ago

New Model 3 Performance launching today 🏎️ → 0-60 mph in 2.9 510 hp / 741 Nm 163 mph top speed — Performance-tuned chassis Same quiet & comfortable cabin plus bespoke chassis hardware for improved stiffness and higher performance baseline. More power, lower energy consumption New Performance 4th gen drive unit can deliver: +22% continuous power +32% peak power +16% peak torque compared to previous Model 3 Performance. All with lower total energy consumption! Forged & staggered 20" wheels + Pirelli P Zero 4 tires Better traction out of corners while limiting traction control interventions. Also, better comfort, lower rolling resistance & increased range. Better Track Mode Track Mode V3 now integrates motor controls, suspension controls, powertrain cooling, & our Vehicle Dynamics Controller (VDC) under a single, unified system. This gives you a more predictable, stable & consistent experience in various track environments. New Adaptive Damping system Adjusts to driver & road inputs in real time to optimize ride & handling, while also improving ride comfort. Controlled via in-house software, which means it keeps improving via future over-the-air software updates. More aerodynamic exterior design 5% reduced drag, 36% lift reduction & 55% improvement in front-to-rear lift balance compared to previous Model 3 Performance. New Sports Seats Same functionality & comfort as before, but with much better lateral support for cornering & dynamic driving

Tesla

34,247,009 views • 2 years ago

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 views • 10 months ago

Dear ICP community, the Internet Computer has now been running strong for 5 years 👏👏👏 Here is a celebratory preview of ICP "cloud engines," the sovereign frontier cloud technology the network shall soon provide from Main points: — Cloud engines enable anyone to spin up their own sovereign frontier cloud. The technology involves an extraordinary inventive step, in which cloud is created from a mathematically secure network of nodes. The nodes run as part of the Internet Computer network ( but are selected and configured by the cloud engine's owner. — The frontier cloud provided by engines is strongly focused on enabling AI agents to build and update online applications and services for us. The world is changing fast, and nearly all new online apps and services are already being built with the help of AI, and thus cloud engines target the future of cloud. — Software hosted on cloud engines is tamperproof, which means that it is immune to infrastructure hacks, because it runs inside a mathematically secure network protocol, rather than on computers directly. This means that AI agents, and those building with them, don't need to have a security team in the loop, or to trust someone else's security team. This is crucial, because in the future, non technical people will demand the freedom to build with full automation — where they just need to issue instructions to AI about what to build, and don't need to worry about anything or anyone else. Of course, apps and services running on engines are also vastly safer from the new breed of hacker being enabled by frontier AI. (The cloud engines themselves are also "tamperproof." Even if a hacker gains physical access to some portion of a cloud engine's nodes, and can make arbitrary changes, the computations and data of the hosted apps and services cannot be corrupted or interrupted so long as the network's fault bounds aren't exceeded. The recent hack of Vercel, a major cloud platform, which gave hackers access to the apps it hosted, provides additional perspective on the importance of this advantage.) — Software hosted on cloud engines is guaranteed to run, so long as a sufficient number of the engine's nodes are running. This means that AI can build applications and services without the need to have a human systems admin team constantly tinkering with the underlying platform to keep it running, which is again crucial, because in the future, non technical people will expect the freedom to use AI to build without the support of others. — New frontier programming language technology, in the form of the Motoko language developed by Caffeine Labs, leverages seminal "orthogonal persistence" technology that unifies program logic and data to deliver further unlocks for AI (Motoko is the first computer language being developed that targets agents that are writing software rather than humans engineers per se). Nowadays, AI can build and update production apps at a prodigious rate, even at the speed of conversation. But it can also make mistakes, and there's a risk that an update it creates might be "lossy" in the sense it causes some transformed data to be lost. Again, in this new world, it's both undesirable and impractical for everyone to have to have a systems admin team on-hand to detect lossy updates and roll them back, but Motoko provides a solution: it can detect new software updates are lossy before they are applied, reducing potentially catastrophic errors by AI to harmless coding retries. — Software hosted on cloud engines is "serverless" but unlike traditional serverless software, directly it directly incorporates data through "orthogonal persistence." Another key purpose is simplify backend software logic and fuel the modeling power of AI by increasing abstraction (sorry for the technical language!!!). Put simply, this enables AI to produce more sophisticated backends, faster, and at dramatically lower costs, as measured by the number AI API tokens consumed during coding. (Tip for the technical: orthogonal persistence is a new paradigm where "the program is the database," and data lives inside program variables, which is possible because it's as if hosted software runs forever in persistent memory). — An expanding database of skills at shall make it possible to develop and directly deploy apps and services to your cloud engines directly from Claude Code, Perplexity, Codex and other AI platforms. Further, your account on can be connected, so that new apps and updates created through conversation automatically appear hosted from your cloud engine. In the future, R&D is going to be very seamless. You converse with AI, and your secure and unstoppable apps or services are created or updated. Cloud engines are designed to directly support this "self-writing cloud" future where we can work hands-free. — Tech sovereignty is becoming a huge issue worldwide, with governments and corporations seeking to create sovereign tech stacks owing to geopolitical tensions. Increasingly, people are realizing that tech provided by foreign nations can come with hidden backdoors and kills switches, from the base platform, right up through hosted apps and services. ICP technology is open source, and those building on ICP using AI own their own source code. When you have the source code, you can verify that there are no backdoors, and when you own the source code thanks to AI, you can update it at will, freeing you from vendor lock-in. But cloud engines take sovereignty much further... — You create a cloud engine by selecting the nodes that will be combined. You can choose the class of nodes used, and their number, but more importantly, you can choose who operates the nodes, and where they are located. Almost any configuration is possible, because the Internet Computer scales the security privileges afforded to hosted software within the network according to configuration (software hosted on cloud engines can directly interoperate with software on other engines and traditional subnets, but base restrictions are applied according to security rules). A cloud engine can be created within a region such as Europe, to comply with regs such as GDPR, or completely within a sovereign state like Switzerland or Pakistan. But cloud engines go further still... — Sovereignty is also about freedom from vendor lock-in. Cloud engines are essentially ICP (Internet Computer Protocol) network configurations, and this means the underlying compute nodes they combine can be swapped out without interrupting their hosted apps and services. This is a big deal. In addition, cloud engines now support nodes that are instances running on Big Tech's clouds, in addition to nodes that are dedicated specialized hardware, as per the Gen I and Gen II nodes that dominate the Internet Computer today. For example, it is possible to have an engine running across different AWS data centers, say, and then reconfigure the engine to run across a mixture of AWS, Google, Azure and Hetzner for even more resilience, without the users of hosted apps and services noticing a thing. That's true freedom. — Sovereign AI is becoming increasingly important too, and cloud engines allow special "AI nodes" to be added to them, so that hosted software can perform inference on hardware provisioned by the owner from a location the owner has selected. Even though the AI nodes are only accessible within the cloud engine, they can still benefit from the forthcoming Internet Intelligence Gateway (IG), which will make it possible to validate inference performed on key frontier open weights LLMs, even when the inference is performed on completely independent AI clouds. When the results of inference are received, this technology can verify that neither the prompt+context (input) nor the inference result (output) have been modified, and that the results were produced by the precise LLM expected. This ensures that AI clouds don't cheat by running inference on cheaper models than are being paid for, and bad actors aren't modifying the inputs or outputs to surreptitiously insert advertising into results, say, or change facts, or insert malware when code is being generated. What's super cool about this technology is the cost of the verification is scalable. A very valuable additional security can be achieved with only 1-2% of extra cost. — Scaling apps and services when they hit capacity limits is another thorny problem that cloud engines help the world address. Engines make scaling possible without rewriting or reconfiguring software. The query workload capacity of hosted software can be horizontally scaled simply by adding new nodes to an engine, and nodes can also be added in geographical proximity to demand. Meanwhile, update workload capacity can first be scaled-up by swapping an engine's nodes out for the next class up, and then when no larger class of node is available, horizontally scaled-out by "splitting" the engine into two, which doubles available capacity. (Technical tip: horizontally scaling update capacity by splitting engines requires multi-canister architectures). — For those who have been following how Caffeine builds apps that can efficiently store large numbers of files, I should mention that apps built on cloud engines will also support the new ICP Blob Storage cloud network (since cloud engines currently have up to about 3 TB of memory, which apps storing large amounts of files can easily exceed). We are also working on allowing blob storage nodes to be added to cloud engines, to enable sovereign mass blob storage within an engine, similarly to how AI nodes can be added currently. — Lastly, but certainly not least, I should mention that cloud engines are multi-blockchain capable, and ready for digital assets, thanks to the clever math at their core. For example, an e-commerce service built on a cloud engine can securely accept and custody stablecoin payments, or a multi-chain DEX could be hosted. Further, engines can support software autonomy (software orchestrated and controlled by other autonomous software, in a decentralized way) and can themselves be orchestrated by SNS technology, and thus run autonomously too. Today, though, the focus is on *mainstream* cloud. This year, the cloud industry will generate approximately one trillion dollars in revenue. That number is already huge, but is expected to grow to two trillion dollars by 2030. After years of continuous development, which have seen more than $500m spent on R&D, the Internet Computer network is now tacking directly toward this mainstream cloud market with cloud engine technology. In their first version, cloud engines are not meant to be a cloud panacea. For example, currently they are not ideal for working with big data. You should use something like DataBricks for that. Cloud engines are carefully targeted at enabling AI to produce traditional online applications and services, including SaaS, in a safer and more productive way, which represents a new market segment with tremendous potential. Of course, DFINITY will continue to work relentlessly to push forward ICP's capabilities, so expect further developments. It's worth mentioning that this cloud segment isn't just about creating new apps and services using AI, it's also about replacing legacy systems and apps built on super expensive SaaS services. Caffeine Labs is working to produce technology (Caffeine Snorkel) that can study an enterprise's legacy systems and app built on SaaS, create replacement systems and apps, and migrate the data, while supporting key stakeholders through the process over email and chat, with full automation. Thus the legacy systems and SaaS markets shall also be addressed by cloud engines. Zooming out, and reasoning in a more metaphysical way, we believe, as we always have, that there is room for a new kind of cloud created by mathematical networks, that provides seminal advances in the fields of security and resilience, as well as true sovereignty and freedom from lock-in. That this same technology, with the help of additional technologies like orthogonal persistence and Motoko, enables AI to build for us without the need for so much oversight, and to create more backend sophistication while consuming fewer AI API tokens, enables ICP to bring game-changing advances to the world. Cloud engines will work synergistically with the Intelligence Gateway, which will enable apps and services running on engines to seamlessly leverage AI, wherever that AI is running, while providing verifiability at extremely low cost for open weights frontier models. We believe that cloud engines represent an inflection point in the storied history of the Internet Computer project, and I'm very proud to be sharing the details with you on the network's fifth birthday 💪 I'll be back with more news soon!!

dom | icp

266,920 views • 2 months ago

🇦🇲 Armenia launches one of the world’s most advanced AI data centers Armenia has officially joined the ranks of countries with their own artificial intelligence infrastructure. #EleveightAI has announced the launch of the #SouthCaucasus’ first AI Factory in the town of Gagarin, powered by #NVIDIA Blackwell B300 chips — one of the world’s most advanced architectures for generative AI. The project is now valued at up to $120 million for the first phase, significantly above the previously reported $70 million. The facility is designed to scale up to 35 MW of computing capacity and is considered supercomputing-class infrastructure. According to the company, Armenia has the potential to become a new global hub for #AI computing — an alternative location to the United States and Western #Europe. Among the advantages cited are lower energy costs, the availability of technical talent, and the region’s growing international connectivity. Eleveight AI CEO Arman Aleksanyan stated that the goal of the project is to transform Armenia from a consumer of AI into a country where AI is developed, trained, and deployed. The company is also considering future expansion into #CentralAsia'n and European markets. Twenty percent of the center’s computing capacity will be provided free of charge to Armenian universities, research institutions, and non-profit organizations. The project is already being described as part of a broader strategy to turn #Armenia into a regional technology and AI hub. As part of that strategy, another AI data center — developed by FireBird — is expected to launch in #Hrazdan, with planned investments of $500 million in the first phase and $4 billion in the second. NVIDIA Blackwell B300 chips are considered among the most advanced AI processors in the world and are designed for generative AI, supercomputing, and high-performance computing workloads. Access to such technology is subject to #US export-control regulations and is granted under strict licensing and compliance requirements. #IT #technology #MiddleEast #EU

Arthur Maghakian

63,776 views • 1 month ago

$AMD is ready to break $1 Trillion MC| $TSM 2nm🧵 TLDR FY 2026(Excluding China AI Revenue) AI GPUs: $35-$50B EPYC Data Center: $15B-$17B Client Segment: $12-$13B Gaming: $6B Embedded: $4B-$5B Total Revenue $70-$100B Non-GAAP net income $18B-$25B Non-GAAP EPS $10.97-$15.40 Foward P/E 55x-70x= $603-$1,078 The semiconductor industry is at a pivotal juncture, with advanced process nodes like TSMC's 2nm technology becoming the battleground for leadership in artificial intelligence and high-performance computing (HPC). Amid this landscape, AMD stands poised to secure early production and higher allocation of its Venice (EPYC ) and MI450 (Instinct GPUs) on TSMC's 2nm process. This strategic advantage is not merely a product of timing but a culmination of a robust partnership, market demand, technical superiority, and geopolitical dynamics. The AI and HPC markets are experiencing unprecedented growth, with inference workloads projected to constitute 80-90% of AI compute by 2030. AMD's EPYC processors and Instinct GPUs are uniquely positioned to capitalize on this trend, particularly given the demand from hyperscalers such as OpenAI , $META , $MSFT, $AMZN, and $ORCL. With $TSM starting 2nm Mass Production in Taiwan is ensuring AMD to meet FY2026 $70B to $100B revenue, driven by non-GAAP net income of $18B to $25B highlights the scale of this opportunity, starkly contrasting with analyst revenue consensus of $39-$45B. This discrepancy arises from analysts' failure to account for major orders, notably from OpenAI(Today SoftBank secured OpenAI a massive cash balance of $55-$62B).OpenAI is raising $100B, so this left $77B from UAE, Saudi, $MSFT, and others. $AMD is on track to receive higher allocation of EPYC Venice and Mi450 in 2026. AMD's acquisition of Xilinx has significantly strengthened its position in AI inference, particularly through adaptive computing technologies like FPGA-based AI Engines. The upcoming Zen 6 "Venice" generation (on TSMC 2nm, launching with MI450 in 2026) promises ~1.7× performance uplift, enhanced vector/AI capabilities, greater thread density, and open firmware innovations positioning EPYC to maintain its inference leadership while powering massive hybrid AI superclusters. TSMC's Fab 22 in Kaohsiung, Taiwan, is now the epicenter of 2nm mass production, a earlier strategic move to meet soaring demand from $AMD and $AAPL. Early production slots are typically reserved for customers with the highest revenue potential and strategic importance. AMD's early tape-out of Venice and the MI450's role as the first AMD GPU on 2nm place it at the forefront of this allocation. The 2nm process offers 10-15% higher performance or 25-30% lower power use compared to 3nm, a critical advantage for AI and HPC applications(TSMC claimed) Moreover, TSMC's recent 20% yield improvement in Versal production, as mentioned in related discussions, indicates efficient scaling. Higher yields translate to more chips produced per wafer, reducing costs and increasing allocation for key customers like AMD. This efficiency is particularly important given the aggressive timelines of customers like OpenAI, who require rapid scaling to meet their computational needs. The reopening of the China market adds another layer of demand pressure. Vendors and hyperscalers are begging for allocation of AMD's MI308X, MI300X, and MI355X, and the 2nm capacity will be critical to meet this need. TSMC's early production of 2nm ensures AMD can capitalize on this opportunity, securing higher allocation to fulfill these orders. Dr. Lisa Su's emphasis on disciplined supply chain planning for multiple gigawatt-scale customers, such as OpenAI, demonstrates AMD's readiness to scale. TSMC's confidence in AMD's ability to absorb this capacity is evident in the early 2nm production allocation. This discipline is particularly important in a market where demand outstrips supply by 10-12x. TSMC's competitors, such as Samsung and Intel, are still in the early stages of their 2nm and equivalent processes. Samsung's 2nm GAA transistors and Intel's 18A process are not yet in mass production, giving TSMC and AMD a first-mover advantage. Nvidia's acquisition of Groq Inc. is a defensive move to diversify into inference, but it does not immediately address the 2nm gap. AMD EPYC Venice and future Gen are already ahead of lowest cost for Inference along with MI450 has TCO of $0.65 to $1.00 per million inference tokens, significantly lower than Nvidia's Rubik (H2 2026) at $0.70 to $1.20 and Broadcom's XPU (2027-2029) at $0.70 to $1.30. Additionally, the MI450's TDP is estimated at 1000-1800W, compared to Nvidia's 2300-3600W (Ultra), reducing operational costs and energy consumption(TSMC 2nm vs TSMC 3nm). The MI450 features 432GB of HBM4 memory and 19.6 TB/s bandwidth, surpassing Nvidia's Rubik (288GB HBM4, 16 TB/s) and Broadcom's XPU (192/256GB HBM4, 7 TB/s est). This enhanced memory and bandwidth capacity is essential for handling the complex, data-intensive workloads of large language models and other AI applications. AMD's full-stack vision, combining EPYC hosts with Instinct accelerators, offers the lowest total cost of ownership (TCO) and thermal design power (TDP). This synergy is unbeatable for both training and inference, further justifying TSMC's prioritization. The 2nm process amplifies these advantages, ensuring AMD can maintain its competitive edge over rivals like Nvidia, whose Rubin GPUs are still on N3P (a 3nm derivative). Today, TSMC just secured $AMD to join the top 10 largest companies in the world as it begins 2nm mass production in Taiwan. AMD and Apple are to receive highest allocation. The long-standing partnership with TSMC, massive demand from hyperscalers, technical advantages of 2nm, and disciplined supply chain planning all point to AMD's favored position. The 2nm process's early mass production at Fab 22, combined with AMD's revenue potential and competitive edge, justifies TSMC's prioritization. This allocation is critical for AMD to meet aggressive demand, capture market share, and solidify its position as a leader in AI and HPC, especially in the inference-dominated future. Dr. Lisa Su "We will multiple customers/hyperscalers at GW scale" Not Financial Advice!

Mike

43,219 views • 6 months ago

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,046 views • 7 months ago