Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

#M5StackNew 🎉LLM-8850 Kit (4G) Released LLM-8850 Kit is a high-performance AI accelerator kit designed for edge AI and embedded computing scenarios. It consists of the LLM-8850 Card AI accelerator card, which is based on #Axera AX8850 SoC, and the LLM-8850 PiHat adapter board, enabling the Raspberry Pi platform to...

11,024 Aufrufe • vor 3 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

After 8+ years on the Tesla Autopilot team and 3 years at Intel, I started Apex Compute to design a new architecture for efficient AI inference. For the past 9 months, we’ve been building our custom inference accelerator. Today we’re releasing Unified Engine v1. Last June we raised our seed round with Maxitech , DeepFin Research, Soma Capital and an incredible group of angel investors. In less than 9 months, we completed our RTL architecture and brought our first pre-silicon prototype to life on FPGA. Our architecture combines systolic array and vector processing in a single compute engine with multiple architectural optimizations, achieving very high FLOPs utilization. A single engine is super lean and it uses less than 90K LUTs and 1 MB Block RAM. It may also be one of the smallest logic-footprint compute engines developed so far. Our Unified Engine v1 supports: -matrix-matrix multiplication (~95% FLOPs utilization) -softmax (~90% FLOPs utilization) -broadcast and element-wise operations -RMSNorm / LayerNorm -block quantization/dequantization (fp4, int4) -multi-engine synchronization and many other operations. We even implemented memory-efficient attention similar to FlashAttention, reaching ~90% FLOP utilization. Full benchmarks and the software stack are available on our GitHub: We have basic compiler written in Python and it supports PyTorch tensors directly to easily test and transfer tensors between the accelerator and host using bf16, fp4 and int4 formats. Our FPGA prototype can already run LLM inference and outperform NVIDIA Jetson Orin Nano, even on a mid-tier FPGA setup (6.4x lower memory bandwidth, 18% slower clock speed at 4.5 Watts). Check the side-by-side comparison video below. Our GitHub includes low-level operator implementations, examples for tiled matrix multiplication, operation chaining, tensor parallelism, attention kernel and a full Gemma 3 1B model implementation. Many more models(Vision Transformers and VLA) are coming soon. Our accelerator IP is AXI-ready for deployment on any AMD(Xilinx) FPGA platform today. Even better, our two-engine prototype runs on an entry-level AMD(Xilinx) FPGA as a PCIe accelerator card. You can purchase it here for $50 to experiment our pre-silicon prototype on your desktop PC or Raspberry Pi 5. We will be releasing hardware bitstream updates as the architecture gets new features. More to come soon! We are expanding our team and looking for compiler engineers and floating-point hardware design engineers. If you're interested, please send me a DM.

Hasan

37,748 Aufrufe • vor 6 Monaten

Dear ICP community, the Internet Computer has now been running strong for 5 years 👏👏👏 Here is a celebratory preview of ICP "cloud engines," the sovereign frontier cloud technology the network shall soon provide from Main points: — Cloud engines enable anyone to spin up their own sovereign frontier cloud. The technology involves an extraordinary inventive step, in which cloud is created from a mathematically secure network of nodes. The nodes run as part of the Internet Computer network ( but are selected and configured by the cloud engine's owner. — The frontier cloud provided by engines is strongly focused on enabling AI agents to build and update online applications and services for us. The world is changing fast, and nearly all new online apps and services are already being built with the help of AI, and thus cloud engines target the future of cloud. — Software hosted on cloud engines is tamperproof, which means that it is immune to infrastructure hacks, because it runs inside a mathematically secure network protocol, rather than on computers directly. This means that AI agents, and those building with them, don't need to have a security team in the loop, or to trust someone else's security team. This is crucial, because in the future, non technical people will demand the freedom to build with full automation — where they just need to issue instructions to AI about what to build, and don't need to worry about anything or anyone else. Of course, apps and services running on engines are also vastly safer from the new breed of hacker being enabled by frontier AI. (The cloud engines themselves are also "tamperproof." Even if a hacker gains physical access to some portion of a cloud engine's nodes, and can make arbitrary changes, the computations and data of the hosted apps and services cannot be corrupted or interrupted so long as the network's fault bounds aren't exceeded. The recent hack of Vercel, a major cloud platform, which gave hackers access to the apps it hosted, provides additional perspective on the importance of this advantage.) — Software hosted on cloud engines is guaranteed to run, so long as a sufficient number of the engine's nodes are running. This means that AI can build applications and services without the need to have a human systems admin team constantly tinkering with the underlying platform to keep it running, which is again crucial, because in the future, non technical people will expect the freedom to use AI to build without the support of others. — New frontier programming language technology, in the form of the Motoko language developed by Caffeine Labs, leverages seminal "orthogonal persistence" technology that unifies program logic and data to deliver further unlocks for AI (Motoko is the first computer language being developed that targets agents that are writing software rather than humans engineers per se). Nowadays, AI can build and update production apps at a prodigious rate, even at the speed of conversation. But it can also make mistakes, and there's a risk that an update it creates might be "lossy" in the sense it causes some transformed data to be lost. Again, in this new world, it's both undesirable and impractical for everyone to have to have a systems admin team on-hand to detect lossy updates and roll them back, but Motoko provides a solution: it can detect new software updates are lossy before they are applied, reducing potentially catastrophic errors by AI to harmless coding retries. — Software hosted on cloud engines is "serverless" but unlike traditional serverless software, directly it directly incorporates data through "orthogonal persistence." Another key purpose is simplify backend software logic and fuel the modeling power of AI by increasing abstraction (sorry for the technical language!!!). Put simply, this enables AI to produce more sophisticated backends, faster, and at dramatically lower costs, as measured by the number AI API tokens consumed during coding. (Tip for the technical: orthogonal persistence is a new paradigm where "the program is the database," and data lives inside program variables, which is possible because it's as if hosted software runs forever in persistent memory). — An expanding database of skills at shall make it possible to develop and directly deploy apps and services to your cloud engines directly from Claude Code, Perplexity, Codex and other AI platforms. Further, your account on can be connected, so that new apps and updates created through conversation automatically appear hosted from your cloud engine. In the future, R&D is going to be very seamless. You converse with AI, and your secure and unstoppable apps or services are created or updated. Cloud engines are designed to directly support this "self-writing cloud" future where we can work hands-free. — Tech sovereignty is becoming a huge issue worldwide, with governments and corporations seeking to create sovereign tech stacks owing to geopolitical tensions. Increasingly, people are realizing that tech provided by foreign nations can come with hidden backdoors and kills switches, from the base platform, right up through hosted apps and services. ICP technology is open source, and those building on ICP using AI own their own source code. When you have the source code, you can verify that there are no backdoors, and when you own the source code thanks to AI, you can update it at will, freeing you from vendor lock-in. But cloud engines take sovereignty much further... — You create a cloud engine by selecting the nodes that will be combined. You can choose the class of nodes used, and their number, but more importantly, you can choose who operates the nodes, and where they are located. Almost any configuration is possible, because the Internet Computer scales the security privileges afforded to hosted software within the network according to configuration (software hosted on cloud engines can directly interoperate with software on other engines and traditional subnets, but base restrictions are applied according to security rules). A cloud engine can be created within a region such as Europe, to comply with regs such as GDPR, or completely within a sovereign state like Switzerland or Pakistan. But cloud engines go further still... — Sovereignty is also about freedom from vendor lock-in. Cloud engines are essentially ICP (Internet Computer Protocol) network configurations, and this means the underlying compute nodes they combine can be swapped out without interrupting their hosted apps and services. This is a big deal. In addition, cloud engines now support nodes that are instances running on Big Tech's clouds, in addition to nodes that are dedicated specialized hardware, as per the Gen I and Gen II nodes that dominate the Internet Computer today. For example, it is possible to have an engine running across different AWS data centers, say, and then reconfigure the engine to run across a mixture of AWS, Google, Azure and Hetzner for even more resilience, without the users of hosted apps and services noticing a thing. That's true freedom. — Sovereign AI is becoming increasingly important too, and cloud engines allow special "AI nodes" to be added to them, so that hosted software can perform inference on hardware provisioned by the owner from a location the owner has selected. Even though the AI nodes are only accessible within the cloud engine, they can still benefit from the forthcoming Internet Intelligence Gateway (IG), which will make it possible to validate inference performed on key frontier open weights LLMs, even when the inference is performed on completely independent AI clouds. When the results of inference are received, this technology can verify that neither the prompt+context (input) nor the inference result (output) have been modified, and that the results were produced by the precise LLM expected. This ensures that AI clouds don't cheat by running inference on cheaper models than are being paid for, and bad actors aren't modifying the inputs or outputs to surreptitiously insert advertising into results, say, or change facts, or insert malware when code is being generated. What's super cool about this technology is the cost of the verification is scalable. A very valuable additional security can be achieved with only 1-2% of extra cost. — Scaling apps and services when they hit capacity limits is another thorny problem that cloud engines help the world address. Engines make scaling possible without rewriting or reconfiguring software. The query workload capacity of hosted software can be horizontally scaled simply by adding new nodes to an engine, and nodes can also be added in geographical proximity to demand. Meanwhile, update workload capacity can first be scaled-up by swapping an engine's nodes out for the next class up, and then when no larger class of node is available, horizontally scaled-out by "splitting" the engine into two, which doubles available capacity. (Technical tip: horizontally scaling update capacity by splitting engines requires multi-canister architectures). — For those who have been following how Caffeine builds apps that can efficiently store large numbers of files, I should mention that apps built on cloud engines will also support the new ICP Blob Storage cloud network (since cloud engines currently have up to about 3 TB of memory, which apps storing large amounts of files can easily exceed). We are also working on allowing blob storage nodes to be added to cloud engines, to enable sovereign mass blob storage within an engine, similarly to how AI nodes can be added currently. — Lastly, but certainly not least, I should mention that cloud engines are multi-blockchain capable, and ready for digital assets, thanks to the clever math at their core. For example, an e-commerce service built on a cloud engine can securely accept and custody stablecoin payments, or a multi-chain DEX could be hosted. Further, engines can support software autonomy (software orchestrated and controlled by other autonomous software, in a decentralized way) and can themselves be orchestrated by SNS technology, and thus run autonomously too. Today, though, the focus is on *mainstream* cloud. This year, the cloud industry will generate approximately one trillion dollars in revenue. That number is already huge, but is expected to grow to two trillion dollars by 2030. After years of continuous development, which have seen more than $500m spent on R&D, the Internet Computer network is now tacking directly toward this mainstream cloud market with cloud engine technology. In their first version, cloud engines are not meant to be a cloud panacea. For example, currently they are not ideal for working with big data. You should use something like DataBricks for that. Cloud engines are carefully targeted at enabling AI to produce traditional online applications and services, including SaaS, in a safer and more productive way, which represents a new market segment with tremendous potential. Of course, DFINITY will continue to work relentlessly to push forward ICP's capabilities, so expect further developments. It's worth mentioning that this cloud segment isn't just about creating new apps and services using AI, it's also about replacing legacy systems and apps built on super expensive SaaS services. Caffeine Labs is working to produce technology (Caffeine Snorkel) that can study an enterprise's legacy systems and app built on SaaS, create replacement systems and apps, and migrate the data, while supporting key stakeholders through the process over email and chat, with full automation. Thus the legacy systems and SaaS markets shall also be addressed by cloud engines. Zooming out, and reasoning in a more metaphysical way, we believe, as we always have, that there is room for a new kind of cloud created by mathematical networks, that provides seminal advances in the fields of security and resilience, as well as true sovereignty and freedom from lock-in. That this same technology, with the help of additional technologies like orthogonal persistence and Motoko, enables AI to build for us without the need for so much oversight, and to create more backend sophistication while consuming fewer AI API tokens, enables ICP to bring game-changing advances to the world. Cloud engines will work synergistically with the Intelligence Gateway, which will enable apps and services running on engines to seamlessly leverage AI, wherever that AI is running, while providing verifiability at extremely low cost for open weights frontier models. We believe that cloud engines represent an inflection point in the storied history of the Internet Computer project, and I'm very proud to be sharing the details with you on the network's fifth birthday 💪 I'll be back with more news soon!!

dom | icp

328,655 Aufrufe • vor 4 Monaten

In a newly released technical update, SpaceX's leadership team, which includes communications manager Dan Huot, Director of Satellite Engineering Ian Dahl, and CEO Elon Musk, detailed a highly ambitious infrastructure roadmap to design, manufacture, and operate specialized artificial intelligence computing satellites at scale. Positioned as a major strategic pillar to dramatically elevate civilizational energy and processing capacity on the Kardashev scale, this strategy moves past traditional communications architectures into massive orbital server arrays. Here is the complete breakdown of the core technologies and timelines driving this space-based intelligence revolution: 🛰️ AI1 satellite power and compute capacity Ian Dahl and Elon Musk introduced the baseline performance targets for the first-generation AI1 satellite, explaining how its custom hardware is engineered to operate like an orbital data center server rack. Ian Dahl noted that their direct operational experience with xAI guided them to target a 150-kilowatt peak power capacity. To manage active machine learning workloads continuously, Elon Musk explained that the satellite is optimized to maintain a sustained average compute power envelope of 120 kilowatts, which directly mirrors the real-world performance of a terrestrial NVIDIA server rack. The official presentation slides outline several key operational metrics for this payload configuration: ⚡ The custom architecture delivers a 150 kW peak compute payload. 🔋 The system maintains a 120 kW sustained average compute payload under active workloads. ⚖️ The hardware achieves a highly optimized power-to-weight density of 70 kW per ton. 🔄 The layout features a completely interchangeable compute provider design. "We thought that the right place to start is around the 150 kilowatt peak power level. But as we look at the workloads with our experience with xAI, we see that we can support about 120 kilowatts of average compute. The 150 kilowatt peak power level roughly matches what, say, an NVIDIA GV300 rack would do. A more reasonable operating envelope would be around 120 kilowatts average power, but it can peak up to 150. So it is basically thinking about it as a rack of compute in space." --- 📐 AI1 satellite dimensions and thermal efficiency specs Elon Musk detailed the physical layout of the AI1 satellite, highlighting the massive dimensions required to accommodate its immense power and cooling hardware. He shared specific design criteria, explaining that the engineering relies on a custom 150 kW solar array paired with a high-capacity deployable liquid radiator thermal management system. The technical specifications of this vehicle layout include: 📏 The structural frame features a massive 70-meter wingspan. ↕️ The vehicle spans a total deployed height of 20 meters. ☀️ The onboard solar array delivers an efficiency of 250 W/m² using technology manufactured in Bastrop, Texas. 🌡️ The thermal system utilizes a 110 m² deployable liquid radiator to cleanly dump waste heat. 🔄 The cooling architecture incorporates redundant pumping loops for mission safety. 🛡️ The exterior contains integrated micrometeoroid shielding to protect the fluid lines. 🧭 The double-sided radiators achieve a dissipation rate of 1400 watts per square meter while remaining oriented knife-edge to the sun. "The assumptions here are 250 watts per square meter for the solar array and about 1400 watts per square meter for the radiators. The radiators are double-sided, radiating on both sides, and they're oriented knife-edge to the sun. They have about a 70-meter wingspan, so these are fairly large." --- 🧩 Simplified design architecture built on Starlink V3 tech Elon Musk explained that despite the satellite's imposing size, its internal architecture is fundamentally much simpler than a standard Starlink satellite. Because it lacks heavy phased array and parabolic communications antennas, the entire vehicle layout is completely streamlined around a few essential structural modules: 🎛️ The hardware framework is arranged around a centralized compute module. ☀️ Large deployable solar arrays extend outward to capture orbital energy. 🌡️ A deployable liquid-radiator thermal management system controls active operational temperatures. 🔄 The engineering team heavily leverages the component evolution and manufacturing experience gained from developing the Starlink V3 vehicle platform. "The AI satellite is actually much simpler than a Starlink satellite. A Starlink satellite has gigantic phased array antennas, parabolic antennas, and a lot of laser links, making it much more complicated. An AI satellite is essentially a lot of solar cells, a radiator, and you still need some laser links, but you don't have all of the super complex antennas that you have on a Starlink satellite. A lot of this is technology we've already made for the Starlink V3 satellites." --- 🔌 Interchangeable compute reference designs and high connectivity Elon Musk outlined a modular hardware approach for the satellite's payload, allowing it to house a variety of industry-standard processing units depending on client requirements. This interchangeable compute rack is supported by a high-bandwidth connectivity loop that links separate orbital units together or transmits data directly back to Earth. The core network parameters include: 🧠 Reference designs are fully established to seamlessly accommodate NVIDIA Reuben chips. 💾 The system architecture is built to support alternative setups using NVIDIA GB300 chips. 💻 Custom hardware layouts are explicitly designed to integrate Google TPUs. 🌐 The onboard communications setup delivers roughly 1 terabit of laser link connectivity. ⏱️ The network closes the communication loop directly with the main Starlink constellation at an ultra-low latency of only 3 milliseconds. "Our current reference design is for NVIDIA Reuben chips, or it could be either GB300 or Reuben chips. We'll also have a reference design for TPUs. Essentially, you can put up any existing chips into orbit. There would also be probably something on the order of a terabit of laser link connectivity from the satellite. Then you can connect these racks of compute to each other by the laser links or directly to the Starlink constellations. Light travels 300 kilometers per millisecond, so that's about three milliseconds away." --- 🏭 The "gigasat" AI satellite and solar production hub in Bastrop, Texas Dan Huot highlighted that the primary production hub for this entire hardware ecosystem is anchored at their sprawling complex in Bastrop, Texas, officially designated as the Gigasat factory. Elon Musk verified that construction is already actively underway on the solar manufacturing facility to feed the project's supply line, with plans moving forward to construct the adjacent AI satellite assembly lines. The physical footprint and timeline of this manufacturing hub are defined by the following benchmarks: 🗺️ The company has over 1,000 acres of land currently owned or under contract for the site. 🏢 The manufacturing complex boasts a massive structural building potential exceeding 11 million square feet. ⚙️ The facility will vertically integrate production to manufacture solar ingots, wafers, solar cells, and completed AI satellites. 📅 Both the solar and AI satellite production lines are targeted to be operational at a viable volume by the end of next year. "We're going to be building a lot of satellites and we're going to be building them here in Bastrop. We already have the solar manufacturing facility under construction, and then we will be building out the AI sat production building soon. We expect to have the AI sat production, the solar production, and all of that operating at some reasonable volume by the end of next year." --- 🏢 The 100-million-square-foot "terafab" chip factory Elon Musk revealed a massive, long-term scaling strategy to build an immense chip manufacturing facility dubbed the "terafab" to completely bypass global semiconductor volume constraints. This manufacturing infrastructure is designed to transition the company into next-generation industrial scaling by producing highly specialized computing components at an unprecedented volume. The scale of this infrastructure project is defined by several extraordinary engineering and production benchmarks: 🏭 The colossal factory is projected to span approximately 100 million square feet, making it ten times larger than the current Tesla Gigafactory Texas. ⚡ The facility is structurally engineered to achieve a massive manufacturing output of 1 terawatt per year once fully operational. 📦 This unprecedented physical footprint provides the capacity required to manufacture 1 billion full-reticle equivalent chips annually. 🔌 Each individual chip manufactured by the facility is designed to run at a power capacity of 1 kilowatt. 🇺🇸 The total scaled output of the facility represents an energy footprint that is exactly double the current annual electricity consumption of the entire United States. "In order to get to the next order of magnitude, you need a gigantic chip factory. To give you a sense of scale here, we expect that the terafab is going to be around 100 million square feet, which is 10 times the size of the Tesla Gigafactory Texas. From a logic die standpoint, that's like having a billion chips per year with a kilowatt per reticle, scaling to a terawatt per year. That is twice the current electricity consumption of the United States." --- 📶 Next-generation high-volume Starlink terminals Dan Huot and Elon Musk introduced their next-generation Starlink user terminals, which have been redesigned specifically to achieve massive manufacturing throughput. Elon Musk pointed out that these newer models will be produced in vastly higher volumes than current hardware designs to fulfill their long-term global deployment targets: 📈 The upgraded user hardware is manufactured at a much higher volume capacity than existing units. 🌍 The company's ultimate target is to successfully deploy a few hundred million of these next-generation terminals worldwide. "In fact, these are the new Starlink terminals, which we made in much higher volume than the current terminals. Ultimately, we think there's probably going to be a few hundred million Starlink terminals out there." --- 📈 Aspirational timeline for orbital AI compute scaling Elon Musk laid out an ambitious, multi-year execution timeline detailing how the company plans to progressively scale space-based processing power. The roadmap targets an initial run-rate by the end of next year and sets an aggressive pace to increase total operational capacity sequentially through a structured, multi-phase timeline: 1️⃣ The initial target aims to hit an annualized run-rate of 1 gigawatt of space AI compute by the end of next year. 2️⃣ The capacity scales to an annualized rate of 10 gigawatts within the next two and a half years. 3️⃣ The operational envelope expands to reach 100 gigawatts in three and a half years. 4️⃣ The long-term deployment plan scales directly to a full terawatt capacity per year using the output of the terafab. "The goal is to get to roughly an annualized rate of a gigawatt per year by the end of next year in terms of space AI compute. Then aspirationally, we want to scale that by an order of magnitude per year. In two and a half years, hitting an annualized rate of 10 gigawatts a year in space, and in three and a half years, maybe a hundred gigawatts, going beyond that with the terafab to scale to a terawatt per year." --- 🌕 Ultimate scaling via lunar production and mass drivers Elon Musk explained that scaling three orders of magnitude past a single terawatt forces a transition completely off-planet to avoid the logistical penalty of Earth's deep gravity well. The vision relies on establishing manufacturing infrastructure directly on the moon to leverage localized resource loops and zero-atmosphere physics: 🌙 The company plans to establish localized raw production lines on the moon to fabricate solar panels, photovoltaics, and radiators from lunar materials. ⚡ Manufacturing components locally avoids the massive fuel and mass penalties of transporting heavy structural materials from Earth. 🧲 Because the moon has no atmosphere and only one-sixth of Earth's gravity, the facility will utilize an electromagnetic mass driver to launch completed satellites. 🚀 Operating essentially as a linear electric motor rail gun, this mechanism will shoot fully assembled AI satellites straight into deep space without relying on chemical rockets. "The only way that we can really see that you can achieve that is on the moon with a mass driver, essentially where you do local production of photovoltaics, solar panels, and radiators on the moon. Because the moon has no atmosphere and only one-sixth Earth's gravity, you can accelerate the AI satellites into deep space without a rocket. You can basically shoot them into space using an electromagnetic gun, like a rail gun type—it's basically a linear electric motor."

Ming

22,203 Aufrufe • vor 3 Monaten

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,267 Aufrufe • vor 9 Monaten