Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

in the lectures below, i hold your hand through low-level LLM systems engineering. it includes everything up to TODAY! 1) pytorch tensors 2) large matmul on cpu vs gpu 3) JAX (and why xAI uses it instead of pytorch) 4) raw cuda kernels and global threading indexing 5) triton...

57,855 Aufrufe • vor 9 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

$AMD $5 Trillion MC Is Inevitable Long Term👑 This thread will focus more on Inference! 2026 EPYC "Venice" $TSM 2nm to save Large GW Scale Inference by 40% more than Prior Turin gen. Context: EPYC Turin achieves ~$0.001 per million tokens for batch inference vs $0.02-$0.12/ million tokens as I wrote the thread below. Venice is going to lower cost down to $0.0005-$0.0006/Million Tokens. OpenAI spent roughly $20B on Inference and Training, where 80-90% of that was for Inference per Analysts. AKA Renting Compute is Expensive AF! In this thread, I want to focus on why most analysts and investors are underestimating the role EPYC "Venice" and future Gen on overall Data center revenue. And $TSM ramping up 2nm supply early is a confirmation that AMD will be a major buyer long term. I will also link the thread the Gap between AMD Analysts & Reality and 2nm Ramp Thread so you have more comprehensive view of what I'm writing here. Before I go into detail this is my 2026 Projection: AI GPUs: $35-$50B EPYC Data Center: $15B-$17B Client Segment: $12-$13B Gaming: $6B Embedded: $4B-$5B Total Revenue $70-$100B Non-GAAP net income $18B-$25B Non-GAAP EPS $10.97-$15.40 Foward P/E 55x-70x= $603-$1,078 AMD's Analysts are projecting $0 Revenue for MI450 and sluggish EPYC Growth. Meaning, all analysts are either full of 💩 or Sexist, you decide! Analysts are also projecting 0% growth on AMD "Secret Weapon" Chip as $MSFT said we are at significant Windows refresh and upgrade cycle. Do you think TSMC would allocate more 2nm supply to $AMD at $0 MI450 revenue and sluggish EPYC? 1. EPYC is going to be the leader in lowest Inference! Current Turin cost saving is 95% vs $NVDA or 98-99% on Inference cost when you factor in renting Inference compute from Amazon Web Services, Microsoft Azure, or $NVDA Neocloud pets. TSMC claimed: 10-15% higher performance at iso-power, 25-30% lower power at iso-speed, and ~15% higher transistor density compared to 3nm. This reduces operational expenses (energy, cooling) while increasing throughput per chip. EPYC Turin achieves ~$0.001 per million tokens for batch inference (via vLLM on models like Llama 3 70B), driven by high core counts and low hardware costs. EPYC Venice offers ~1.7x overall performance and up to 70% more compute capability per core, with up to 256 cores (512 threads). Enhanced vector/AI instructions and open-source firmware (openSIL) optimize for inference workloads. AMD Incorporates AI Engines (now part of AMD's XDNA) for on-chip acceleration, improving efficiency for low-latency and edge inference. This reduces reliance on discrete GPUs, lowering system complexity and TCO. Venice SKUs are projected at $3,000-$15,000 ($5,000 for 256-core flagship), far below NVIDIA Rubin ($50,000-$90,000) or AMD's own MI450 GPUs ($40,000-$50,000). High memory bandwidth (up to 1.6 TB/s) supports efficient batch inference. Venice is designed exactly for Large customers that want to lower Inference Cost and MI450 Helios is for Customers that want Training at lowest TCO, TDP as well as lower Upfront 1GW scale(Full build $35-$40B vs $NVDA $55B-$80B). 2. Real World Example: OpenAI's 2025 inference spend reached ~$20B, escalating to even higher total compute rental (mostly inference) amid token volume growth(from video generating). By 2026, with usage doubling (consistent with industry trends: token demand grows 2-5x YoY), assume OpenAI processes ~1,800 billion million-tokens annually $NVDA Blackwell at $0.02-$0.12 is $36B(most optimized) Rubin is projected to be at $0.01/million tokens or $18B annual Inference Cost vs $AMD Venice $0.0005/million tokens or $0.9B annual Inference Cost => Massive saving for OpenAI or anyone that are paying 80-90% Annual Bill for Inference compute. In short, it is unsustainable to pay this much rent vs owning for all current AI players for the medium to long term. Rubin excels in low-latency decode (if Groq integration from $20B deal in 2027-2028), but Venice dominates batch (80% of inference by 2030). Actual savings depend on deployment scale (OpenAI's 6GW AMD plans), electricity rates, and software maturity. If Rubin only hits $0.03, savings swell to $53.1B vs. $17.1B. 3. Will running Inference on Venice and future Gen slow down response generation in 2026 and beyond? Human perception of "fast enough" for chat, agents, search augmentation, summarization, coding assistance is roughly Meaning, EPYC may generate $100B a year on data center revenue, Hence $MSFT $AMZN $META $GOOGL OpenAI xAI and 42+ Countries are leaning AMD for Inference, because the cost saving is MASSIVE! 4. Regular users (you, me, people using ChatGPT, Claude, Gemini, Grok, Perplexity...) are extremely unlikely to notice any slowdown and in many cases might even experience slightly faster or more consistent response times if the industry heavily shifts toward AMD EPYC for inference. What actually happens when companies save massively on inference? When OpenAI , Anthropic , Gemini , Grok Meta .... save billions on the batch/enterprise/RAG layer using EPYC Venice, they typically do one or more of these things with the savings, none of which make your chat slower but enhancing their bottom line(Profit) ~Keep prices the same → make more profit ~Lower subscription prices / increase free tier limits ~Train bigger & better models more frequently ~Offer longer context windows ~Add more reasoning steps / tool calls / agents per query ~Improve multimodal capabilities ~Build more data centers / reduce throttling during peaks In practice the consumer experience usually gets better, not worse, when inference becomes dramatically cheaper. Prime example is $META leaning AMD heavily or currently AMD largest customer. or Grok 2 to Grok 3 heavily used AMD for Inference saving. And most Grok Users reported Groke responses snappier, not slower. 5. What does this mean for potential Revenue? Noted that TSMC is massively ramping 2nm supply for $AMD both MI450 and EPYC. EPYC Conservative projection: FY2025: $10.5B(best Est) FY2026: $16B FY2027: $29B FY2028: $49B FY2029: $75B FY2030: $100B Large customers: $META OpenAI $MSFT $AMZN $GOOGL xAI (Apple?) Smaller customer: $DELL $HPE $SMCI and 42+ other countries. The roadmap to $5 Trillion is very much inevitable as Inference Cost from Renting or owning $NVDA are too high, but $NVDA will still dominate Training market share, where MI families are likely to take 15-20% market share, but the TAM is also expanding Rapidly. Most Institutions are projecting $2-$3Trillion TAM by 2030. $NVDA said $4 Trillion. Dr. Lisa Su said $1 Trillion+ by 2030. So you decide on how much TAM. If you enjoy this kind of analysis, Slap the Like/Repost and Bookmark to please the X Algo as it is Free.99! If you want to support my work further, consider subscribe to see more in-depth analysis! Alright, that is it. Not Financial Advice!

Mike

102,223 Aufrufe • vor 7 Monaten

After 8+ years on the Tesla Autopilot team and 3 years at Intel, I started Apex Compute to design a new architecture for efficient AI inference. For the past 9 months, we’ve been building our custom inference accelerator. Today we’re releasing Unified Engine v1. Last June we raised our seed round with Maxitech , DeepFin Research, Soma Capital and an incredible group of angel investors. In less than 9 months, we completed our RTL architecture and brought our first pre-silicon prototype to life on FPGA. Our architecture combines systolic array and vector processing in a single compute engine with multiple architectural optimizations, achieving very high FLOPs utilization. A single engine is super lean and it uses less than 90K LUTs and 1 MB Block RAM. It may also be one of the smallest logic-footprint compute engines developed so far. Our Unified Engine v1 supports: -matrix-matrix multiplication (~95% FLOPs utilization) -softmax (~90% FLOPs utilization) -broadcast and element-wise operations -RMSNorm / LayerNorm -block quantization/dequantization (fp4, int4) -multi-engine synchronization and many other operations. We even implemented memory-efficient attention similar to FlashAttention, reaching ~90% FLOP utilization. Full benchmarks and the software stack are available on our GitHub: We have basic compiler written in Python and it supports PyTorch tensors directly to easily test and transfer tensors between the accelerator and host using bf16, fp4 and int4 formats. Our FPGA prototype can already run LLM inference and outperform NVIDIA Jetson Orin Nano, even on a mid-tier FPGA setup (6.4x lower memory bandwidth, 18% slower clock speed at 4.5 Watts). Check the side-by-side comparison video below. Our GitHub includes low-level operator implementations, examples for tiled matrix multiplication, operation chaining, tensor parallelism, attention kernel and a full Gemma 3 1B model implementation. Many more models(Vision Transformers and VLA) are coming soon. Our accelerator IP is AXI-ready for deployment on any AMD(Xilinx) FPGA platform today. Even better, our two-engine prototype runs on an entry-level AMD(Xilinx) FPGA as a PCIe accelerator card. You can purchase it here for $50 to experiment our pre-silicon prototype on your desktop PC or Raspberry Pi 5. We will be releasing hardware bitstream updates as the architecture gets new features. More to come soon! We are expanding our team and looking for compiler engineers and floating-point hardware design engineers. If you're interested, please send me a DM.

Hasan

37,648 Aufrufe • vor 4 Monaten

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,179 Aufrufe • vor 7 Monaten

$NVDA $GFS NVIDIA’s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIA’s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groq’s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groq’s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groq’s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The company’s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groq’s solution is to move “control” into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groq’s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groq’s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groq’s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that “no useful models” fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groq’s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groq’s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth “upwards of 80 TB/s” and contrasts that with off-chip HBM bandwidth “about 8 TB/s,” asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes “TruePoint numerics,” including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the company’s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in “available speed” for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groq’s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as “tokens-as-a-service,” aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groq’s documentation states its API is designed to be “mostly compatible” with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groq’s go-to-market is not simply “sell chips,” but “sell predictable unit economics per token,” with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native “edge within the cloud” strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groq’s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groq’s marketing materials and partnerships provide signals about demand vectors. The company’s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groq’s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, which—if realized in deployments—can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groq’s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groq’s core proposition—deterministic, compiler-scheduled inference with predictable latency—aligns directly with the segment where GPU generality is least valued and where “good enough” programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIA’s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groq’s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The “technology acquisition” rationale is unusually strong in this specific case because Groq’s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIA’s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groq’s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groq’s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIA’s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chips—while not costless—could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship “inference capacity” in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIA’s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groq’s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groq’s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groq’s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a “preemptive acquisition” angle. The reporting identifies recent investors in Groq’s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groq’s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groq’s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groq’s compiler-led deterministic model is philosophically and practically different from CUDA’s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIA’s own rapidly improving inference performance complicates the “need” for Groq but does not eliminate the “option value.” NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIA’s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is “special” about Groq’s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows “hundreds” (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio manager’s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about “NVIDIA needs more inference speed” and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 Aufrufe • vor 7 Monaten

$AMD $620/share is too conservative for 2026 🧵 Some quick facts before I dive into this super long thread: $META allocated 42% GPUs to $AMD and 58% to $NVDA OpenAI allocated 6GW(38%) to $AMD and 10GW to $NVDA My $620 PT below by end of 2026 was only for 10-15% market share. I believe $AMD is going to have much much higher market share than I projected. The AI accelerator market is exploding, projected to reach $500 billion by 2028(is now heading $1Tril), driven by insatiable demand for training and inference compute in large language models (LLMs), recommendation systems, and autonomous systems. Nvidia ($NVDA) has long held a stranglehold, commanding over 90% market share through its CUDA ecosystem and superior rack-scale solutions. However, AMD is mounting a formidable challenge, leveraging cost advantages, open-source software momentum, and hyperscaler partnerships to erode Nvidia's moat. Recent deals—such as Meta's ($META) allocation of 42% of its GPU capacity to AMD and OpenAI's commitment to 6GW of AMD compute (versus 10GW for Nvidia)—signal a tipping point. At the forefront is AMD's Instinct MI450 series, a next-generation AI GPU slated for H2 2026 launch, which promises "no-excuses" leadership in training, inference, and distributed workloads. This analysis dissects how AMD will capture more market share and why hyperscalers like $Meta , xAI , Oracle , and others are poised to become voracious buyers of the MI450. AMD's AI GPU revenue has surged from negligible levels in 2022 to an estimated $4-5 billion in 2025, capturing ~6% of the data center GPU market. This growth stems from the Instinct MI300X, which offers 141GB of HBM3 memory and competitive FP8/FP16 performance at 20-30% lower cost than Nvidia's H100. Hyperscalers, facing NVIDIA 's overcharging, have turned to AMD for diversification. Meta, for instance, plans 600,000 H100-equivalent GPUs by end-2024, with ~42% (or 250,000+ units) sourced from AMD's MI300 series for inference tasks like image editing and AI assistants. Similarly, OpenAI's recent multi-year deal commits to 6GW of AMD compute—equivalent to ~300,000-400,000 MI450 GPUs—starting with 1GW in 2026, explicitly to counterbalance its 10GW Nvidia allocation. These aren't one-offs. Microsoft Azure, Amazon AWS, and Oracle Cloud Infrastructure (OCI) have integrated MI300X for AI workloads, with Oracle deploying 30,000 MI355X units in zettascale clusters. xAI, Elon Musk Musk's AI venture, ran 30% of Grok-1's production traffic on MI300X GPUs and has confirmed ongoing purchases. Collectively, these partners represent over $400 billion in projected AI infrastructure spend through 2028, with AMD targeting up to 40% market share. For those that subscribed, I wrote a specific thread on how AMD "secret weapon" is going to change the game in 2026 with an improved designs on all its products, yes AMD has patent on it. Software is the linchpin. AMD's ROCm platform, once derided as "half-baked," now supports day-zero integration for Llama-4, DeepSeek V3, and GPT-OSS models—closing the CUDA gap. Benchmarks show MI355X (MI450 precursor) outperforming Nvidia's B200 in inference by 1.5-2x on memory-bound tasks, at 25-35% lower TCO. For training, MI450's rack-scale IF128 configuration (128 GPUs, 1.4 PB/s intra-rack bandwidth) rivals Nvidia's VR200 NVL144, enabling clusters like xAI's Colossus (scaling to 1M GPUs). My below thread projected Etimated conservative FY 25 revenue: $34-$36B Estimated conservative FY 26 revenue: $55B-$62B Below is why $AMD is revenue is going to be much higher after OpenAI deal. 1. OpenAI 1GW in 2026. With high demand for MI355X at $30,000k+ per unit, with MI450 is likely to be sold in the $45k-$55k. We can safely calcuate 1GW would require roughly 400,000 MI450 GPUs. or Roughly ~$20B revenue in 2026 alone from OpenAI. That would mean $AMD would hit $56B just from one partnership(OpenAI) in 2026 2. $META, the biggest spender on AI Infrastructure right now, Daddy Zuckerberg bought 250,000+ MI300, and is buying MI355X for recommendation engines and Llama training. It is very unlikely for Daddy Zuck to slow down AMD Chips, due to its Inference superiority to NVDA Chips. Most likely we will see at least 300,000-400,000 MI355X ordered from now toward end of H1 2025. And another 300,000-500,000 MI450 by H2 2025. Or ~$20B from just Meta in H2 alone, excluded H1. 3. xAI : Musk confirmed "AMD GPUs work very well" for Grok's small/medium models, with 30% of Grok-1 on MI300X. xAI's Colossus (200K+ GPUs, targeting 1M) and Oracle partnership (via OCI's MI355X cluster) position it for MI450 trials in H1 2026. With $6B funding and Grok integration into Oracle services, xAI could allocate 10-20% ($10B-$15B) to MI450 for distributed inference. We haven't heard the detail from Daddy Elon Musk yet, but most likely not going to be spending less than OpenAI or Sam Altman 4. Oracle ($ORCL): A multi-billion-dollar MI355X deal powers OCI's AI superclusters, with $500B+ remaining performance obligations. Larry Ellison's zettascale ambitions and xAI/OpenAI integrations make Oracle a MI450 anchor tenant—projected 50-100k units ($15B+ spend) for enterprise AI platforms. $ORCL is likely to spend more on the new "secret weapon" due to its capability in AI inference and cost advantage for $500B backlog. 5. Others ( Microsoft , Amazon , Saudi+other countries): Microsoft (Azure MI300X for training) and Amazon ($148B 15-year spend) test MI450 via Stargate ($500B with Oracle/SoftBank). Emerging buyers like G42 (5GW UAE campus), Crusoe, and Hot Aisle add 5-10GW demand. These potentially would add $15B-$30B in 2026 alone. We also need to factor in $TSM supply constraint( $NVDA is TSMC favorite), so $AMD market cap/growth is being tamed by TSMC. So what are you saying Mike, well $AMD 2026 revenue could hit $90-$100B by end of 2026 or nearly 185% growth YoYo. So what does that mean for valuation? I have no idea how Mr. Market gonna value AMD in 2026 with 3 digits growth. My Conservative $620 was my best projection until today with OpenAI partnership. I'm telling you as one of the biggest AMD bull, that I will leave it to "smart money" and other investors to do the price discovery while I'm chilling and writing DDs daily. Lastly, AMD's MI450 isn't hype—it's a calibrated strike at Nvidia's vulnerabilities, amplified by hyperscaler bets like Meta's 42% allocation and OpenAI's 6GW lifeline. By prioritizing inference efficiency, rack-scale innovation, and open ecosystems, AMD will siphon 10-15% share in 2026, scaling to 20%+ as TCO trumps CUDA loyalty. Meta, xAI, Oracle et al. aren't passive; they're active co-designers, betting billions on MI450 to fuel AGI pursuits without Nvidia's premium. For investors, this is AMD's inflection Per Dr. Lisa Su Not Financial Advice!

Mike

711,006 Aufrufe • vor 10 Monaten

$AMD $NVDA & the AMD Bear SemiAnalysis 🧵 Here are some facts: $META allocated 42% AI GPUs to $AMD OpenAI allocated 6GW(38%) to $AMD 1. Model-Specific Bias: Llama 3.3 70B graph favored NVIDIA due to TRT-LLM optimizations, highlighting throughput and latency where Blackwell excels. In contrast, the GPT-OSS 120B chart shifts focus to cost and interactivity, where MI355X shines. This selective model choice clearly suggests SemiAnalysis tailors benchmarks to reinforce narratives—NVIDIA’s dominance in speed (Llama 3.3) and AMD’s niche in cost (GPT-OSS). GPT-OSS 120B, with its sparse attention mechanisms (similar to DeepSeek-V3.2-Exp), shows AMD’s CDNA 4 architecture, while Llama 3.3’s dense attention favors NVIDIA’s Tensor Cores. SemiAnalysis’ decision to emphasize Llama 3.3 initially could reflect its AMD bear stance. 2. The way Data is presented The Llama 3.3 graph focused on raw performance metrics (throughput vs. latency), downplaying cost, where AMD holds an edge. This new chart, buried in follow-up posts, reveals AMD’s strength but receives less prominence, suggesting a curated narrative. Labeling variability (e.g., B200 with/without TRT) and the lack of uniform scaling across graphs indicate potential cherry-picking of configurations to favor NVIDIA’s optimized setups. 3. Historical Context: SemiAnalysis’ past critiques of AMD’s R&D and ROCm (web results from May 2025) align with a bearish outlook. Their own hype/brand around NVIDIA’s 15x ROI contrasts with muted coverage of AMD’s cost advantages, reinforcing bias. Despite AMD’s participation in InferenceMAX, the benchmark’s framing (e.g., prioritizing Blackwell’s ROI) reflect SemiAnalysis’ market predictions rather than balanced analysis. Lastly, AMD’s Instinct MI355X proves superior in inference and cost per million tokens for the GPT-OSS 120B model, offering a 25% cost advantage over NVIDIA’s H200 at moderate-to-high interactivity levels. This efficiency, driven by AMD’s memory bandwidth and FP4 support, makes it a better choice for cost-sensitive, multi-user deployments over a three-year horizon. However, SemiAnalysis’ sole focus(presentation graph) on Llama 3.3—where NVIDIA excels demonstrates a pattern of cherry-picking models and data to favor NVIDIA , consistent with its historical AMD bearish stance. This selective presentation risks misleading stakeholders by overshadowing AMD economic strengths. My personal take: I would trust Dr. Lisa Su, and Greg Brockman Sam Altman take on AMD and how they viewed and allocated 6GW for AMD over SemiAnalysis . At the end of the day, Large customers pay when it works. $Meta allocated 42% AI GPUs to $AMD for a reason. And the "secret weapon" will improve energy consumption by 20-50%, meaning at 6GW, OpenAI would be able to deploy 25-50% more MI450 at a much better cost advantage, higher memory bandwidth, and the queen of Inference! Oh and ROCm 8 is expected to be on par with CUDA in 2026.

Mike

104,194 Aufrufe • vor 9 Monaten

[CLIP] by Hand ✍️ The CLIP (Contrastive Language–Image Pre-training) model, a groundbreaking work by OpenAI, redefines the intersection of computer vision and natural language processing. It is the basis of all the multi-modal foundation models we see today. How does CLIP work? Goal: 🟨 Learn a shared embedding space for text and image [1] Given ↳ A mini batch of 3 text-image pairs ↳ OpenAI used 400 million text-image pairs to train its original CLIP model. Process 1st pair: "big table" [2] 🟪 Text → 2 Vectors (3D) ↳ Look up word embedding vectors using word2vec. [3] 🟩 Image → 2 Vectors (4D) ↳ Divide the image into two patches. ↳ Flatten each patch [4] Process other pairs ↳ Repeat [2]-[3] [5] 🟪 Text Encoder & 🟩 Image Encoder ↳ Encode input vectors into feature vectors ↳ Here, both encoders are simple one layer perceptron (linear + ReLU) ↳ In practice, the encoders are usually transformer models. [6] 🟪 🟩 Mean Pooling: 2 → 1 vector ↳ Average 2 feature vectors into a single vector by averaging across the columns ↳ The goal is to have one vector to represent each image or text [7] 🟪 🟩 -> 🟨 Projection ↳ Note that the text and image feature vectors from the encoders have different dimensions (3D vs. 4D). ↳ Use a linear layer to project image and text vectors to a 2D shared embedding space. 🏋️ Contrastive Pre-training 🏋️ [8] Prepare for MatMul ↳ Copy text vectors (T1,T2,T3) ↳ Copy the transpose of image vectors (I1,I2,I3) ↳ They are all in the 2D shared embedding space. [9] 🟦 MatMul ↳ Multiply T and I matrices. ↳ This is equivalent to taking dot product between every pair of image and text vectors. ↳ The purpose is to use dot product to estimate the similarity between a pair of image-text. [10] 🟦 Softmax: e^x ↳ Raise e to the power of the number in each cell ↳ To simplify hand calculation, we approximate e^□ with 3^□. [11] 🟦 Softmax: ∑ ↳ Sum each row for 🟩 image→🟪 text ↳ Sum each column for 🟪 text→ 🟩 image [12] 🟦 Softmax: 1 / sum ↳ Divide each element by the column sum to obtain a similarity matrix for 🟪 text→🟩 image ↳ Divide each element by the row sum to obtain a similarity matrix for 🟩 image→🟪 text [13] 🟥 Loss Gradients ↳ The "Targets" for the similarity matrices are Identity Matrices. ↳ Why? If I and T come from the same pair (i=j), we want the highest value, which is 1, and 0 otherwise. ↳ Apply the simple equation of [Similarity - Target] to compute gradients of for both directions. ↳ Why so simple? Because when Softmax and Cross-Entropy Loss are used together, the math magically works out that way. ↳ These gradients kick off the backpropagation process to update weights and biases of the encoders and projection layers (red borders).

Tom Yeh

67,858 Aufrufe • vor 2 Jahren

Neuroscientist Dr. Jeff Beck from Noumenal Labs discusses the fundamental nature of representation, understanding, and modelling, comparing biological intelligence with current artificial intelligence. Jeff argues that *how* information is represented dictates predictive ability and that LLMs, while impressive at symbol manipulation and pattern matching (like next-word prediction), lack the *grounded*, causal understanding of the world inherent in biological systems. Timestamps: 00:00 - Cat visual cortex experiments & discovering orientation sensitivity (slide projector analogy) 01:49 - Representation choice and neural coding (orientation vs. feature intensity) 02:30 - Choice of representation impacts predictions; generative models 03:15 - Importance of choosing the right generative model for predictions 03:35 - The problem: We don't know the brain's true generative model 03:55 - Theory of Mind (ToM) in LLMs 04:05 - Jeff Beck's ToM tests on early ChatGPT (stapler example) 05:40 - ChatGPT recognizing the ToM test vs. passing it 06:32 - Analogy: LLMs recognizing known problems vs. generalizing (sum/product riddle) 07:25 - Do LLMs implicitly build world models? Vicarious experience analogy 07:59 - The difference: Grounding symbols in reality outside language 08:35 - AI Alignment: Difficulty in capturing human reward functions & belief formation 09:21 - Nightmare scenario: Humans as "complacent value function selectors" 09:44 - Hope: AI enhancing human understanding, not replacing thought 10:08 - Philosophy of science: Science realism vs. modeling pockets of regularity 10:39 - Noise in models as ignorance or deliberate exclusion (design choice) 11:00 - Design choices in science, controlled experiments, and induced bias 11:29 - Are there true, discoverable mathematical laws of the universe? 11:41 - Is there a "true" ground truth distribution (P)? Beck's answer: No (with nuance) 12:55 - Ontological vs. Epistemological divide: Perfect models vs. models of regularities 13:21 - Are scientific models "false by definition"? The Bayesian perspective 14:07 - "All knowledge is conditional"; Are foundational theories (e.g., FEP) true or just perspectives? 14:51 - FEP as a mathematical framework, not a theory; models are just models 15:54 - Legibility vs. Utility: Useful but illegible AI models 16:01 - Prediction vs. Explanation: Trusting black boxes can be unsatisfying 16:30 - Why understanding AI matters: Ensuring alignment with human decisions/values 17:08 - Line-of-sight legibility as an alignment approach 17:14 - Benefits of explainable AI: Human understanding and value alignment verification 18:21 - RL components: Prediction engine, reward function, policy; the alignment challenge 19:25 - Trusting AI = Trusting its policy aligns with our reward + its superior beliefs 19:57 - Language: Intrinsic representation vs. pointers between shared minds 20:21 - Why language works: Shared internal models and common grounding 21:00 - Basis of shared understanding: Not linguistic, but shared experience/intuitive physics 22:44 - Consciousness and language as lossy, simplified summaries of complex brain processes 23:22 - Evidence for simplification: Brain regions, perception vs. representation; limits of language models 24:32 - Counterpoint: Language captures complex/ambiguous human concepts 24:54 - Language as massive compression: The information bottleneck (Meister's paper) 26:22 - Implication: Language/actions are poor representations of internal understanding 27:03 - Can language models understand? The mimicry argument (Piantadosi) 27:33 - Beck's skepticism: LLMs excel at prediction/mimicry, not true understanding 28:09 - LLM explanations replicate structure but lack grounding 28:54 - Beck's test for LLM understanding: Genuine novelty beyond training data 29:19 - Summary: Symbol manipulation is not understanding; grounding is key 30:06 - Abstraction and Idealization in scientific modeling ("The Brain Abstracted") 30:45 - Revisiting Newton: Intuitive physics is correct for our world; idealizations are simplifications 32:01 - Sophistication & boundaries: Nested systems vs. one complex system? 32:32 - The boundary problem in FEP/Markov Blankets: Where to partition? 33:41 - Beck's research: Finding principled partitions based on interaction dynamics 35:44 - Beyond direct experience: Imagination, language, and learning 36:16 - Human creativity: Creating new *things* by combining modeled objects (Systems Engineering) 37:43 - Goal for AI: Automating systems engineering for creative combination 38:17 - Sutton's "Reward is Enough" paper 38:25 - The challenge of "Reward is Enough": Defining and obtaining the *right* reward function 39:02 - Difficulty of eliciting individual reward functions 39:52 - The core alignment problem: Accessing and representing individual reward functions 40:13 - Impossibility: Disentangling beliefs and rewards from observed actions 41:51 - Argument analogy: Disagreements stem from different beliefs or values 43:00 - Prerequisite for value inference: Understanding belief formation 43:13 - Building aligned systems: Sparsity of data, meta-models vs. base system modification 43:46 - Proposed solution: AI layer that models the human's belief formation system 44:40 - Alignment process: Align beliefs first, then address value differences 45:00 - Conclusion CC Maxwell Ramstead

Machine Learning Street Talk

17,828 Aufrufe • vor 1 Jahr

how to set up hermes agent step by step. built-in memory, 40+ tools, works on your phone, and what to think of hermes vs openclaw: 1. hermes is a personal AI agent that runs in your terminal. think of it like open claw but with built-in memory, 40+ tools out of the box, and 90% cheaper token costs. you install it with one command. 2. the 3 problems with open claw that hermes solves: no memory (you keep repeating yourself), constant gateway restarts, and zero visibility into what you're spending on tokens. 3. hermes remembers everything. every completed task gets saved to memory. it searches through past logs to find solutions. over time it literally gets smarter at your specific workflows. 4. connect it to open router. you see exact costs per model per task. free models rotate weekly. one founder went from $130 every five days on open claw to $10 on hermes. same output. 5. it comes preloaded with skills. apple notes, imessage, find my, browser, web search, image generation, cron jobs. no hunting for plugins. 6. connect it to obsidian so it reads your entire vault. connect it to gstack for your dev environment. create custom skills for your specific workflows. 7. the biggest money saver: have it write code once for recurring tasks. then it runs without burning tokens every time. stop paying an LLM to do the same scrape or report daily. 8. run it on android via telegram. name your agents. talk to them like coworkers. in this episode imran shows you how to set this up. 9. you can run it bare metal, in docker, or serverless on modal. pick your risk level. i begged imran to come on The Startup Ideas Podcast (SIP) 🧃 and walk through the full installation live. he made it impossibly clear. if you've heard of Hermes Agent and want the clearest explanation of how to get set up like a pro let me know what you want me to cover on the next ep this is the best personal agent setup video on the internet right now. watch

GREG ISENBERG

618,145 Aufrufe • vor 3 Monaten

Davide Asnaghi (Davide Asnaghi) is the co-founder and CEO of Diode Computers, Inc., a Brooklyn-based startup using AI to design and manufacture circuit boards in the United States. Before Diode, Davide worked on Apple’s Special Projects Group and spent time in Hong Kong and Shenzhen studying Asia’s electronics manufacturing ecosystem. That experience convinced him that PCB design, despite powering everything from smartphones and satellites to medical devices and autonomous systems, remained one of the most overlooked layers of the tech stack. Since its founding just two years ago, Diode has landed Physical Intelligence and Saronic as customers and partnered with Anthropic to help Claude become a better electrical engineer. The company’s ultimate ambition: to make hardware as nimble as software. In our conversation, we explore: 1. Why the West outsourced PCB manufacturing to Asia in the 2000s and why bringing it back matters for American competitiveness 2. What Shenzhen’s manufacturing culture does better than Silicon Valley (and what the U.S. can learn from it) 3. How Diode’s models can one-shot much of schematic design and compress hardware timelines from months to weeks 4. The three-week YC pivot that transformed Diode from a design validation tool into a full-stack manufacturer 5. Why circuit boards are the “forgotten middle child” between silicon and software 6. How Diode partners with Anthropic to make LLMs better electrical engineers 7. What it takes to build a hardware company in 2025—and why the talent bar must stay incredibly high 8. How Italian, American, and Chinese cultures shaped Davide’s approach to entrepreneurship and manufacturing Thank you to the partners who make this possible .TECH domains: An identity for builders at their core: Guru_HQ: The AI source of truth for work: Brex: The intelligent finance platform: (0:00) Intro (4:15) Why Davide calls himself a copper merchant (5:53) Diode’s mission to rebuild PCB manufacturing in the U.S. (7:58) What success looks like (9:00) Growing up in northern Italy and spending a year in Minnesota (13:14) Why Italy produces fewer venture-backed founders (15:30) Why Hong Kong accelerated Davide’s learning (19:09) Silicon Valley vs. Shenzhen (22:05) What Davide learned in Apple’s Special Projects Team (24:11) Why Davide left Apple after two years (26:54) Meeting his co-founder, Lenny (29:32) How Davide uncovered the need for better PCB design and manufacturing (33:23) PCB manufacturing in Asia, and Diode’s approach (41:29) The YC pivot that changed Diode’s business (44:39) Inside Diode’s customer journey (48:10) Where the value is in electronics manufacturing, and Davide’s AGI thesis (51:30) What separates a working board from a great one (55:32) Where Diode fits in the electronics stack (59:55) Diode’s early near-death moment and long-term vision (1:02:30) Diode’s exceptionally high bar for hiring (1:04:48) Where Davide gets his best ideas (1:07:00) Final meditations The Generalist

Mario Gabriele

37,137 Aufrufe • vor 2 Monaten

$AMD Massive Rotation from $NVDA $INTC🧵 Not Financial Advice! DYOR! 5-10 minutes before the bell today, last trading day of May 2026, massive rotation out of $INTC and $NVDA into $AMD. I wrote this thread this morning on what $TSM said on Energy Efficiency is now TOP Priotity and why AMD is the biggest winner. Of course I did not have influence on this rebalancing, I was just pointing out why Dr. Su saw this coming years ago. (Check the picture to understand more). I been talking about Agentic AI for like 3-4 years now. OpenClaw broke the CPU:GPU Ratio 1:4 narrative to 1:1 to 5:1 in late Jan and Feb 2026. I will link various threads where you can understand the full picture from supply chain, to TSMC expansion, and different Wafer Ratio for EPYC Venice and MI455X. Energy efficiency is a structural, long-term driver behind institutional rotation from $NVDA and $INTC into $AMD (with spillover strength in $AVGO for complementary networking/custom silicon). This isn't just short-term rebalancing, it's a massive bet on the shift from AI training (performance-at-any-cost) to inference, deployment, and embodied/agentic systems (where total cost of ownership, power draw, and scalability dominate). Precisely What I been writing about $AMD for years now, probably at least more than 5,000 threads.This is the FOMO from Institutions to own $AMD. Do know that AMD is the least owned Semi Stock among vs Peers. AI infrastructure is moving beyond massive training clusters to widespread inference for Agentic AI (running models 24/7) and embodied AI (robots, autonomous agents, edge devices). These workloads prioritize: ~Tokens-per-watt and performance-per-watt ~Lower total power consumption for data centers facing grid constraints ~Better economics at scale (cost-per-token, TCO) ~Thermal and power efficiency for on-device/robotics use Hyperscalers are now thinking more about Margin, Profitability, and $/M Tokens At $516/share. AMD Fwd PEG Ratio is still 35/100+= 0.35 AKA very cheap IMO for the growth and potential. A. Why institutions rotated out of $NVDA? Because Agentic AI is going to dominated by CPUs for years to come, moving violently to 5-10-20:1 CPU:GPU Ratio as enterprises are demanding more than 10-20 agents to run tasks. Now, that does not mean training is going away, Inference is just going to grow much faster. B. Why instiutitons rotated out of $INTC? Because AMD x86 unit share is only at 30-31% but Revenue share is already at 46.2% according to Mercury Research. And Dr. Su wants 50-60% market share, and that would mean 60-70%+ Revenue share where the CPUs TAM Is now already at $200B in 2026 and projected to be $500B by 2030. C. Why $AMD? Because AMD secured meaningful 2nm Capacity, Advanced Packaging and Memory through 2027-2028. And TSMC is expanding 2 primary 2nm Fabs toward 60-65k WPM each, and speeding up 5 2nm Fabs in Taiwan. With total up to 12 2nm Fabs through 2027/2028. 2nm Capacity is expected to be 140k+ WPM toward end of 2026, and 220-240k WPM by end of 2027. Apple has secured 35-45k WPM. And AMD does not have to worry about allocation competition until late 2027 from $AVGO for $META and $GOOGL(This may change) D. Agentic AI will evolve to 24/7 Autonomous Agent, and that will become the foundational layer for Robotic or Physical AI. Agentic AI (autonomous systems that plan, reason, use tools, self-correct, pursue long-horizon goals, and adapt) provides the high-level cognitive architecture. It turns raw perception and low-level control into useful, general-purpose behavior in the physical world. Physical AI (or Embodied AI) refers to AI that senses, understands, and acts directly in the real world through robots, actuators, and sensors. Agentic capabilities are what make this scalable and useful beyond narrow, scripted tasks. Reactive/programmed machines → To proactive, goal-oriented autonomous agents. How does this work? Autonomous Agent layer is the brain ~Vision-Language-Action models or robotics foundation models. ~Agentic loops: Planning, chain-of-thought reasoning, reflection, tool use (simulators, APIs), multi-step task decomposition. ~Persistent 24/7 operation with Memory, world modeling, continuous learning. Institutions may not like $AMD from 2022-2025, but they cannot stop this evolution and it is inevitable. Part of my main thesis for AMD to get to $5 Trillion Market Cap Long Term. Conclusion: Institutions are rotating capital toward AMD not merely for tactical rebalancing, but because Dr. Lisa Su and her team anticipated this exact inflection years in advance and have been methodically engineering AMD’s platform to dominate it. Dr. Su has long championed the convergence of Agentic AI as the high-level cognitive foundation for Physical AI and robotics. As far back as her 2023/2024 CES keynote and earlier strategic commentary, she described Physical AI (including humanoid robotics and edge autonomy) as “the next big thing”; a natural extension of agentic workflows moving from digital reasoning to real-world action. She emphasized that enabling persistent, 24/7 autonomous agents requires a full-stack approach: high-performance CPUs for orchestration and motion control, dedicated accelerators for real-time vision and multimodal inference, and open software ecosystems for rapid development. This vision aligns precisely with the structural drivers we’ve discussed. As AI shifts from training to massive-scale inference and embodiment, energy efficiency, total cost of ownership, and heterogeneous compute become first-order advantages. AMD’s Instinct MI350/MI355 series, Ryzen AI Embedded processors, and EPYC platforms deliver superior performance-per-watt and balanced CPU + GPU + NPU integration ideal for power-constrained robots that must run sophisticated agentic reasoning loops without excessive thermal or battery drain. Dr. Su has repeatedly highlighted the rising importance of CPUs in agentic systems (moving toward 1:1 or even CPU-heavy ratios with GPUs), positioning AMD’s strengths in orchestration, memory handling, and efficiency as critical for the next phase of growth. AMD is engineered for the deployment realities of embodied agents: scalable, efficient, and deployable at the edge and in physical systems. The institutional flows out of NVDA and INTC into AMD reflect recognition of this prepared leadership. Dr. Su didn’t just see the future of Agentic AI powering robotics, she has spent years building the silicon, software, and partnerships to make it practical and economically viable. This rotation signals confidence that the companies best positioned for the physical, always-on intelligence layer will capture the highest-volume opportunities in the coming decade. Not Financial Advice! DYOR!

Mike

104,109 Aufrufe • vor 2 Monaten

$AMD $5 Trillion is Inevitable LT| Agentic AI🧵 Agentic AI is the new $5 Trillion TAM 🚨🚨🚨 This thead will do Comp with $INTC and how to quantify this massive Agentic AI demand spike, and forcing Jensen to rush a CPU design. Global Agentic AI Market size is estimated to be $3-$5Trillion TAM by 2030(McKinsey) Quantifying the demand from agentic AI for AMD involves assessing the broader market growth for agentic systems, their unique computational requirements (particularly for CPUs in orchestration and reasoning tasks), and AMD's positioning very well through products like EPYC processors and partnerships. AMD EPYC Venice is the most superior choice in 2026-2027 for most Agentic AI workloads Agentic AI refers to autonomous AI agents that perform multi-step tasks, involving sequential logic, tool integration, and decision-making workloads that heavily rely on CPUs for handling orchestration, memory management, and context switching, rather than just GPU-parallelized training or batch inference. Agentic AI is often cited as 40-100x more "hungry" than traditional AI due to its continuous, 24/7 operation and complex workflows. This stems from factors like chain-of-thought reasoning (multiple LLM calls per query), API/tool interactions, memory management, and orchestration loops, which can generate 10-100x more tokens and require real-time responsiveness. For example, a single agentic query might trigger 5-20 model inferences, making it 10-20x more compute-intensive than simple chatbots, and the always-on nature compounds this to 40-100x overall. Nvidia's CEO has highlighted this as driving "easily 100x more computation" for inference in agentic/reasoning setups. AMD's EPYC Venice (6th Gen EPYC, codenamed "Venice") and Intel's Xeon 7 Diamond Rapids represent the pinnacle of server CPU technology in 2026, both targeting high-performance data center workloads like AI inference, agentic AI orchestration, cloud computing, and HPC. Venice builds on AMD's Zen 6 architecture, emphasizing core density and efficiency, while Diamond Rapids leverages Intel's Panther Cove P-cores for balanced performance. Both chips adopt similar advancements like 16-channel DDR5 memory and PCIe Gen 6, but differ in core counts, process nodes, and overall design philosophy. Intel has faced acute supply constraints across its Xeon lineup, including legacy nodes (Intel 7/3) and the ramping 18A process for next-gen parts. Intel shortage is expected with lead times up to 6 months or longer. 1. AMD EPYC Venice vs Intel Xeon 7 Diamond Rapids Architecture AMD: Zen 6 chiplet design with 8 CCDs and dual IODs Intel: Panther Cove P-cores; multi-die architecture with 4 compute tiles Core/Thread Count AMD: Up to 256 cores / 512 threads (Zen 6c variant) Intel: Up to 192 cores / 192 threads Process Node AMD: TSMC N2 (2nm) Intel: Intel 18A (1.8nm-class); in-house fab Memory Support AMD: 16-channel DDR5; up to 1.6 TB/s bandwidth. Intel: 16-channel DDR5 ; up to 1.6 TB/s bandwidth I/O and Connectivity AMD: PCIe Gen 6 (up to 128 lanes); twice the CPU-to-GPU bandwidth Intel: PCIe Gen 6 (up to 128 lanes); LGA 9324 socket Power (TDP) AMD: Starting 400-500W, potentially lower due to efficiency gains from TSMC 2nm Intel: Starting 400-500W, as it targets competitive efficiency Performance Projections AMD: Up to 70% uplift vs. 5th Gen Turin (1.7x in multi-threaded/AI tasks) Intel: ~40% faster than Granite Rapids (Xeon 6, 128-core). Lags AMD in per-core perf and 40-50% behind Venice core-for-core comp Target Workloads AMD: AI inference/orchestration, HPC, cloud virtualization. Partnerships Intel: Hyperscale AI, general enterprise. Custom silicon Pricing: AMD: estimated $10k-$20k for top SKUs Intel: estimated $8-$18k Availability: AMD: Significant Ramp H2 2026 due to higher allocation from TSMC Intel: H1-H2 2026 delayed, but trying to catch up Overall: ~Venice's 256 cores provide a 33% edge over Diamond Rapids' 192, making it superior for massively parallel tasks like AI training/inference or virtualization ~TSMC's N2 vs. Intel 18A debates rage on which is "better," but AMD's mature chiplet approach yields better density ( 32 cores/CCD vs. Intel's 48/tile). Venice's redesign reduces latency, aiding agentic AI where CPUs handle orchestration ~ Early projections show Venice widening AMD's lead matching or exceeding Diamond Rapids' perf with fewer watts in multi-threaded benchmarks. Intel's no-SMT design (to prioritize AI) handicaps it vs. AMD's 512 threads, though Clearwater Forest (E-core) could compete in density-focused niches. ~Power & Cooling: Both push above 400-500W, demanding liquid cooling. ~AMD been taking market share now above 40%. AMD EPYC Venice emerges as the superior choice in 2026 for most server workloads. Its higher core/thread count (256/512 vs. 192/192), stronger per-core performance, and architecture optimized for AI-driven tasks (agentic orchestration with GPU integration) provide decisive advantages in throughput, scalability, and efficiency. Projections indicate Venice delivering 1.7x the performance of prior gens while widening the gap over Intel ( 40-70% leads in multi-threaded benchmarks). AMD's fabless model with TSMC ensures reliable scaling, and its ecosystem ( open ROCm) appeals to AI adopters. Intel's Diamond Rapids is competitive in single-threaded enterprise apps and custom hyperscale ( NVLink), with potential fab advantages for supply/security. However, without SMT and lower density, it falls short in core-for-core battles—exposing Intel to another generation of AMD dominance unless 18A yields surprise efficiency gains. For data centers prioritizing raw compute ( AI, HPC), Venice wins; for Intel-centric ecosystems or specialized I/O, Diamond Rapids holds ground. Real benchmarks post-launch will confirm, but logic points to AMD pulling ahead. 2. Market size , Potential Revenue and Supply Global Agentic AI market size is projected to be $3-$5 Trillion by 2030 according to McKinsey, where consensus points to 40-50% CAGR driven by small to large enterprise demand. I also wrote a full thread on how and why Agentic AI is so explosive that AMD will blow all anlaysts estimate for subscribers. Link below if you are interested. AMD's data center segment hit a record $5.4B in Q4 2025 (up 39% YoY), with EPYC shipments ramping due to agentic demand. With 2GW of deployment in H2 2026, AMD AI data center revenue has $40-$50B+ at the lowest or most conservative projection; or Total Revenue in the $77-$94B For FY2026. However, Agentic AI massive demand spike could send EPYC revenue 3x to 4x in the next few years, potentially surpassing MI series GPU demand as enterprises prioritize CPU-dense Rack setups. This is pushing $NVDA Jensen to rush a CPU design and acquired Groq, a new CPU player due to this massive TAM. Noted that this is just popping just in weeks, highlighting we are just so early in this AI Supercycle and the pace of adoption is insane, and clearly productivity will skyrocket. Why? Because Agentic AI is 24/7 Smart AI agent working for you or your businesses is a mad compelling, and it is estimated to be 40-100x more Inference Hugnry! Many experts already said it is impossible to project this kind of Inference Demand. AI CapEx is expected to ramp up even more in 2027-2028-2029 and 2030 as Global Agentic AI is going to scale to $3-$5 Trillion TAM by 2030. The nature of Agentic is driving higher CPU/GPU ratio, with CPUs handling 50-90% of Agentic workflows. For example, The current Helios Rack: 18 compute trays per rack with 72 GPUs + 18 CPUs. The beauty of this $META and $AMD long term partnership is, that it is absolutely flexible to adjust racks to higher CPU rato or equal to service different needs. Helios rack can be easily swap to 2 GPUs 2CPUs or even CPUs only trays for dedicated orchestration/head nodes. You see, the beauty of this open rack-scale is flexibility and evolvability. If Agentic AI demand pushes much higher, AMD should be able to adjust variant trays without abandoning Heilos Rack. We can't talk just about massive Agentic AI demand without talking about the Supply side or TSMC. TSMC, AMD's primary foundry for advanced nodes ( Zen 6/Venice on N2/2nm), is addressing AI-driven shortages through massive expansions. TSMC accelerates fab construction with up to 10 facilities targeted for 2026. TSMC is accelerating its domestic manufacturing expansion, with industry sources indicating that as many as ten fabs could be under construction or preparing to begin operations across Taiwan’s major science parks. TSMC Capex: $52-56B in 2026 (up 37% YoY), with $45B already approved for new/upgraded capacities. 70-80% for advanced processes (2nm/A16), 10-20% for packaging (CoWoS quadrupling to 120-140K wafers/month by late 2026). In addition, Taiwanese companies (led by TSMC) commit to at least $250B in direct investments in US-based advanced semiconductor, AI, and energy production/innovation capacity.Taiwan provides $250B in government credit guarantees to facilitate additional investments and build a full US semiconductor ecosystem (including industrial parks). TSMC completed a second land purchase in Arizona (January 2026) for gigafab scaling, with an additional $100B+ (potentially four more modules) to further expand and qualify for tariff exemptions. AMD with secured 12GW from OpenAI and $META and massive Agentic AI will mean higher priority acess to 20-30% more wafers on TSMC advanced nodes, as TSMC has multi-year agreements with AMD for AI chips. Dr. C. C. Wei, CEO of TSMC quote: "I spend a lot of time in the last three or four months talking to my customer and then customers. Customer. I want to make sure that my customers demand are real. I talk to those cloud service providers, all of them. Their answer is. I'm quite satisfied with their answer. Actually they show me the evidence that the AI really help their business. So they grow their business successfully and he or she in their financial return. So I also double check their financial status. They are very rich." Amid shortages, the US buildout ensures AMD can ramp production of Instinct GPUs and EPYC CPUs without the constraints hitting competitors like Intel. By diversifying away from Taiwan (85% of advanced nodes today), the agreement mitigates supply disruptions, ensuring stable flows for AMD's chips. Scaling production and securing supply will matter for AMD the most in the next 5-10 years growth. The growth could be 80-100% YoY or higher; or it could be in the 60%. The aggressive TSMC supply ramp is reassuring the higher growth point. Conclusion: AMD stands at a pivotal inflection point in 2026, where the explosive rise of agentic AI demanding 40-100x more inference compute through its 24/7, multi-step orchestration positions the company to potentially triple its EPYC CPU revenue to $45-60B+ by 2028 while scaling Instinct GPUs to tens of billions annually by 2027. Agentic AI demand could push AI CapEx closer to $1 Trillion in 2027, far higher than most estimates. Dr. Lisa Su, AMD's visionary CEO, is masterfully securing supply to harness this massive demand by prioritizing operational execution and deep TSMC collaboration, ensuring readiness for the second-half 2026 AI ramp. Dr. Su has explicitly called out surging EPYC demand for agentic tasks where CPUs power head nodes and traditional workloads alongside GPUs while guiding for data center dominance through proactive capacity planning and partnerships like Nutanix ($150M investment for open agentic platforms) or providing tens of millions CPUs for OpenAI, $META, $ORCL, $AMZN, $MSFT, $GOOGL and others. Her strategy includes multi-year TSMC agreements for advanced nodes (N2 for Venice CPUs and future Instincts), diversifying beyond Taiwan to mitigate risks, and unveiling innovations like the MI455X GPU at CES 2026, which she touted as enabling "the next trillion-dollar market opportunity" in physical AI. Dr. Su's forward-looking vision predicting AI reaching 5 billion users emphasizes "AI everywhere," backed by hardware like Ryzen AI chips, all while declaring demand "going through the roof" and committing to scale without bottlenecks. TSMC's aggressive ramp-up, fueled by $52-56B in 2026 capex (up 37% YoY) and 10+ new fabs across Taiwan, the US (Arizona cluster expanding to 6+ modules with $165B+ investment), Japan, and Europe, provides profound reassurance for AMD's supply stability. The January 2026 US-Taiwan agreement committing $250B in investments and credit guarantees for US reshoring accelerates this, granting tariff relief (15% rates with 1.5-2.5x exemptions) tied to capacity buildouts, enabling TSMC to potentially double output over the decade to meet AI wafer hunger. This translates to 20-30% higher wafer allocations on key nodes, sidestepping Intel-like shortages and empowering Dr. Su's team to deliver on hyperscaler demands without disruption. Ultimately, this synergy cements AMD's leadership in the agentic era, promising sustained growth, $5T+ valuations at scale, and a resilient path forward as AI reshapes the world. This is NOT Financial Advice! Video source: AMD CES 2026

Mike

44,460 Aufrufe • vor 5 Monaten

$AMD's heading to $5T MC LT| Lowest $/M tokens 🧵 The real reason why Institutions are FOMOing into AMD while other Semi stocks are underperforming ($NVDA $AVGO) Not Financial Advice! DYOR! Under Dr. Lisa Su’s leadership, AMD has transformed from a distant challenger into a formidable force in AI infrastructure, delivering the industry’s most compelling TCO story for high-volume inference. Her clear vision open ecosystems, aggressive annual roadmaps, rack-scale innovation, and relentless focus on tokens-per-dollar has positioned AMD’s Helios racks as the go-to solution for hyperscalers and AI natives struggling with exploding token costs, collapsing the cost down to $0.0003-$0.0005/M tokens. I will link various threads on this analysis to supply chain and wafer ratio if you are interested in understanding the full picture. In the last 3-4 months, explosive Agentic AI demand significantly increased Inference demand for Agentic AI models with 5-10 agents. If you are a listener of CNBC or Bloomberg, u should know enterprises and companies are complaining abt cost of token, and how it starts to spike up way too much to make sense. The fact that most data center today are run by $NVDA Chips, where the cost is way too high for Training or Inference. 1. Token cost Here are some quick comp, so u understand why $META OpenAI Anthropic $MSFT $AMZN Softbank $GOOGL and many more small to medium AI Natives are buying AMD CPUs and GPUs as much as they want, or pretty much AMD chips are sold out for the next 3-5 years. Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens 2. Why Hyperscalers and AI Natives Are Choosing AMD Token consumption (especially Agentic) is outpacing even NVIDIA’s efficiency gains, making diversification mandatory for economic viability. Massive deals reflect this reality like $META, OpenAI, $MSFT, Softbank, $AMZN, Oracle, LumaAI, G42... Dr. Lisa Su’s Vision in Action: Since taking the helm, Su has driven AMD’s turnaround with disciplined execution, annual GPU cadence (MI300 → MI350 → MI400), full-stack software (ROCm 7), open ecosystems (UALink, OCP designs), and customer-centric rack-scale solutions like Helios. Her emphasis on “tokens per dollar” and TCO has turned AMD into the pragmatic choice for sustainable AI scaling. Power/Energy Efficiency: ~Helios Rack-level is estimated at 120kW-140kW with 50% more HBM4 where Inference and Training cost matter ~Rubin Rack-Level is estimated at 160kW-230kw AMD Helios shines in owned TCO, memory density, and energy flexibility at hyperscale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B 3. Superior CPUs to pair with GPUs on massive scale 5-10-20GW Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. Conclusion: NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always-on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. Not Financial Advice! DYOR! Video source: Microsoft Build 2026

Mike

145,778 Aufrufe • vor 2 Monaten

$AMD| The FOMO to buy AMD Chips is NOW 🧵 Not Financial Advice! DYOR! Research Purpose Only! The Inference Queen is the biggest winner in Agentic AI where all other CPUs are struggling to compete with a 2yr old EPYC Turin and EPYC Venice is in mass production phase. AMD stresses deployability today on standard x86 platforms (no proprietary architectures required), full software compatibility, and open standards. This positions Venice + Helios as a practical, high-density alternative to competing solutions while underscoring that agentic AI shifts the balance toward CPU-rich racks alongside GPUs, and most importantly, lowering the cost of token to accelerate adoption and innovation. Context: The Wall Street Journal yesterday came out with an article that OpenAI is condiering drasstically lowering the token prices to win more customers from Anthropic. The narrative "they" are trying to exacerbate the current AI selloff won't last long. This is a fundamental misunderstanding of what is going on, or what I already discussed for months and years. Followers and Subscribers already knew this for years, that this day would come, where token cost will bcome the central discussion among enterprises as there is no such thing as unlimited budget or Tokenmaxxing when they use $NVDA chips or In-house Hyperscalers chips. I will link various threads if you are interested in understanding the full picture from supply chain to recent TSMC Rapid 2nm expansion up to 12 Fabs total by 2027/2028. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. The OpenAI-AMD 1GW Helios deployment (starting H2 2026) represents a pivotal vertical integration move that directly supercharges the inference economics. This isn't incremental; it's a structural shift toward ownership of massive, optimized rack-scale capacity, enabling the lowest token costs and triggering the enterprise adoption flywheel. We need to be honest, $AMD is the only company that made a big bet on Inference since the day Chatgpt became sensational where $NVDA and others were betting big on Training. At the end of the day, Token bill from Anthropic has to obey economics. Meaning the bills rise, companies have to get more out of it to justify the cost. It cannot be an unlimited inference budget, and it has to show up on efficiency, profitability and operating leverage. 1. Tokenomics After you understand this, you will understand why Citi cited Anthropic is likely to sign a deal with $AMD along with Hyperscalers, AI Labs, Sovereign AI like Softbank 5GW in France and many other countries. However, OpenAI and $META are now wanting faster deployment, and they are AMD shareholders now, they have prioritized allocation. Anthropic and Hyperscalers just cannot compete when Helios Rack lower token cost to$0.0003–$0.0005 per million tokens at GW scale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens Now, OpenAI, META and Hyperscalers can lower Inference cost even further with $AMD EPYC Venice "dense rack" or Agentic AI Rack. AMD published a detailed technical blog emphasizing that the future of agentic AI autonomous, multi-step AI systems requiring heavy orchestration, databases, caching, APIs, and control planes demands massive CPU-dense rack-scale infrastructure, not just GPUs. The catalyst prominently positions their upcoming 6th Gen EPYC "Venice" processors as the key enabler for next-generation dense racks, delivering leadership throughput under real-world power, cooling, and density constraints. ~EPYC Venice (Zen 6 architecture, up to 256 cores / 512 threads per socket) is projected to deliver exceptional rack-level performance. In AMD’s modeled 100 kW rack comparisons, Venice-powered systems are expected to achieve ~3.30x the throughput of NVIDIA’s Vera (88-core Olympus) baseline across a broad mix of agentic-supporting workloads. ~This builds on current-generation 5th Gen EPYC "Turin" (up to 192 cores), which already delivers ~2.37x rack throughput vs. Vera and ~1.6x vs. Intel’s Xeon 6980P (128 cores). ~ Liquid-cooled Turin deployments already support >27,000 CPU cores per rack today. Venice is architected to push this beyond 36,000 cores in the same rack class, dramatically increasing concurrent agent capacity and overall infrastructure efficiency. 2. Ownership vs renting compute from Hyperscalers matter to OpenAI and only owning $AMD chips can meaningfully lower token cost for enterprises. ~Eliminates cloud overhead: No provider margins, utilization buffers, or egress fees. Direct control over power contracts, cooling, scheduling, and orchestration at dedicated facilities. ~Helios optimizations at GW scale: Rack-level density (1.4+ exaFLOPS FP8 per rack), high HBM4 bandwidth, EPYC orchestration for agentic workloads, and superior TCO/TDP. AMD's long-standing focus on tokens per dollar/watt shines here 20-40%+ efficiency edges in inference-heavy scenarios. ~At 1GW+ optimized deployment, inference hits $0.0003–$0.0005 per million tokens (community/analyst models tied to Helios metrics). This is dramatically lower than typical rented/cloud equivalents, especially for high-volume output tokens in agentic flows. High token bills today, enterprises running heavy agentic/coding/analysis workloads can face $50-100M+/month at current API rates (flagship models $5-30+/M output, scaled to massive volumes). Post-Helios compression, same volume will drop to $10-15M/month (or better) via lower underlying costs passed through as pricing flexibility, volume tiers, caching, or batch discounts. ROI thresholds collapse. More companies greenlight pilots → production → massive scaling. Agentic AI (autonomous workflows) multiplies token demand exponentially, but affordability removes the friction. OpenAI gains flexibility, Unlike more cloud-dependent rivals (Anthropic), they can lower effective pricing, offer aggressive enterprise bundles, or absorb volume without margin destruction directly tackling "high token bill" complaints while maintaining profitability as usage explodes. 3. Agentic AI Models shifted CPU:GPU Ratio to 1:1 toward 3-5:1 with Explosively Token-Hungry Workloads Agentic AI (autonomous, multi-step agents with planning, tool use, iteration, and self-correction) is fundamentally more compute and token intensive than conversational or single-turn generative AI. Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. ~Agents often generate 10–100x+ more tokens per task due to iterative reasoning chains, multiple tool calls, verification loops, and long-context orchestration. ~Goldman Sachs forecasts token consumption multiplying 24x by 2030 (to 120 quadrillion tokens/month) largely driven by agentic adoption in consumer and enterprise. ~Enterprise data shows agent-pattern workloads growing at 680% annualized rates, projected to surpass conversational AI in token volume by Q3 2026. ~Daily enterprise agent token consumption is already in the billions, with complex workflows (coding, workflows, analysis) amplifying this dramatically. 4. Competitive Edge: Winning Customers from Anthropic Anthropic’s Claude models (especially Opus/Sonnet) excel in complex reasoning and agentic coding, commanding premium positioning. However, their higher underlying costs (heavier reliance on third-party cloud with margins) limit pricing flexibility compared to OpenAI’s owned Helios capacity. Anthropic is on track to generate $10.9 billion in Q2 revenue. The company expects to achieve its first-ever quarterly adjusted operating profit of $559 million. However, sustaining full-year profitability remains challenging due to immense computing and model training costs The truth is, Anthropic has no choice but to buy as much $AMD chips as possible if they want to compete with OpenAI or get investors attention. This 5% adjusted operating profit to revenue ratio is just pathetic. Current pricing dynamics (2026): OpenAI already undercuts on many tiers ( flagship output tokens significantly cheaper than equivalent Claude Opus). Nano/mini models offer 5–10x advantages for volume work. Anthropic holds edges in long-context flat pricing and certain reasoning quality. OpenAI after Helios Rack Ownership, At $0.0003–$0.0005/M effective costs, OpenAI gains massive headroom to: ~Aggressively discount high-volume agentic tiers or bundles. ~Offer “unlimited” enterprise plans or usage-based models that Anthropic struggles to match without margin erosion. ~Target cost-sensitive, high-throughput agent deployments (dev tools, automation platforms) where token bills explode. Enterprises facing $ millions in monthly agentic bills will migrate to the provider delivering better economics at scale. OpenAI’s combination of strong models (o-series reasoning) + lowest TCO positions it to erode Anthropic’s enterprise share, especially as agentic becomes the dominant token consumer. Cheaper tokens expand the total addressable market dramatically. This feeds the data/model improvement loop, justifying further capex. AMD benefits from proven scale pulling in more customers (Meta, Oracle, Microsfot, Amazon, Softbank, TensorWave, LumaAI ... already aligned on Helios). Conclusion: Dr. Lisa Su has been laser focused on inference economics since at least 2022–2023, repeatedly emphasizing that the real battleground for AI scalability would be TCO, power efficiency (TDP), and ultimately tokens per dollar and per watt not just raw training FLOPS. While many viewed inference as a secondary, commoditized workload, Dr. Su architected AMD’s roadmap around rack-scale systems optimized for high-volume, sustained inference that would dominate as models matured and usage exploded. Helios represents the culmination of that multi-year bet: a fully integrated, open platform designed precisely for the economics of massive token throughput. This deep, strategic partnership with OpenAI starting with the 1GW Helios deployment in H2 2026 and scaling to 6GW, is the embodiment of that shared vision. Both companies foresaw a future where agentic AI models evolve to become extraordinarily token-hungry: autonomous agents executing complex, iterative workflows with planning, tool use, verification loops, and long-context reasoning. These workloads can consume 100x+ more tokens per task than traditional chat or single-turn generation, driving exponential demand as capabilities improve and enterprises deploy them at scale. By owning and optimizing this massive Helios capacity at GW scale, OpenAI achieves inference costs as low as $0.0003–$0.0005 per million tokens. This structural cost advantage allows OpenAI to absorb the coming token explosion profitably, dramatically lower effective pricing for enterprises, and win high-volume agentic workloads from higher-cost competitors like Anthropic. What was once a prohibitive monthly token bill becomes an affordable accelerator for productivity and innovation. The OpenAI-AMD alliance validates Dr. Su’s prescient strategy and turns the Agentic flywheel into reality: Collapsing inference costs → explosive token consumption → richer data and better models → accelerate greater demand. This partnership doesn’t just address today’s economics, it positions both leaders at the center of the infrastructure buildout that will power AI’s next decade. By delivering the lowest inference economics at scale, OpenAI not only solves enterprise bill pain but gains a decisive weapon to win share from higher-cost rivals like Anthropic. And that is why OpenAI and $META will deploy EPYC Dense Rack Not Financial Advice! DYOR! Research Purpose Only!

Mike

84,951 Aufrufe • vor 1 Monat

Seoul came alive with shared excitement today, thanks to Shaw and SB. I believe we all felt the same vibrant energy in the room together. Here’s the full video of the meetup, hope you can feel it through the digital air, too 🙏: Part A. Fireside chat 1. Introduction (00:03): 🎙️ Shaw, Founder at Eliza Labs and Creator of elizaOS 2222222222222222222222222🎙️ SB An, Data and Tech Lead at Hashed 2. elizaOS Overview (00:48): - Shaw describes elizaOS as an open-source agent framework he's been developing for about a year, aiming for an open and collaborative project that shares all its developments. 3. ai16zDAO and the Autonomous Investor (01:08): - Shaw explains that ai16zDAO has been working on an autonomous investor or trader project, leveraging the community to gain alpha and generate profits. 4. Origins of ai16zDAO (02:25): HAW discusses how the idea for ai16zDAO emerged after his introduction to daos.fun and its founder, @ who suggested the name "ai16z." 5. Autonomous Trading vs. Investing (04:48): - Shaw clarifies the distinction between autonomous trading, involving buying and selling tokens, and investing, which focuses on pre-validating and funding projects. 6. Marketplace of Trust (06:17): - Shaw introduces the "marketplace of trust," where ai16zDAO aims to identify effective traders and utilize their signals to guide the autonomous trading system. 7. elizaOS Features and Differentiators (11:25): - Shaw highlights that elizaOS is built using web technologies, addresses the social loop, and supports multiple blockchain networks, setting it apart from other agent frameworks. 8. ai16z Token and Launchpad (13:22): - Shaw outlines plans for the ai16z token, including launching a platform that allows projects to deploy tokens paired with the ai16z token. 9. Team Growth and Expansion (15:17): - Shaw shares that the ai16zDAO team has rapidly grown to around 50 members, with further expansion plans. 10. Asia Tour and Opportunities (17:23): - Shaw discusses the enthusiasm from the Asian community, which constitutes a significant portion of the ai16zDAO team and contributors. 11. Contribution Opportunities (19:10): - Shaw outlines ways for developers and non-developers to contribute to the ai16zDAO ecosystem, including through a retroactive funding program and the AI agent dev school. Part B. Q&A Session Question #1 (20:52 - 25:07, Speaker: Steve Lee): - What differentiates your use of AI for investment decisions compared to quant funds using LLMs and ML? - Is ai16zDAO focused mainly on liquid token investments rather than traditional VC? How do you measure performance? Answer #1: - ai16zDAO uses a "marketplace of trust" model, leveraging collective intelligence rather than LLMs for direct investment decisions. - Investments are smaller (e.g., $50K) and aimed at media engagement while maintaining a focus on treasury management and ecosystem growth. - Question #2 (25:07 - 31:38, Speaker: johncho.& (k/acc)): - How does elizaOS compare to competing frameworks (@0xzerebro, AI Rig Complex, etc.)? - How do crypto-oriented frameworks compete with traditional AI frameworks like those from OpenAI or Anthropic? Answer #2: - elizaOS emphasizes developer accessibility using TypeScript and has a growing community, providing first-mover advantages. - Traditional frameworks focus on models, while crypto frameworks enable functionality like social media interactions and decentralized integrations. - Question #3 (31:38 - 33:20, Speaker: johncho.& (k/acc)): - Do you see OpenAI and Anthropic as competitors? Would you expand to their domain? Answer #3: - OpenAI and Anthropic excel in model training but avoid riskier, user-facing applications like social connectors. - elizaOS focuses on delivering real-world functionality quickly, addressing gaps left by traditional frameworks. - Question #4 (33:20 - 38:15, Speaker: Kevin, Hashed Open Research): - What human functions is elizaOS targeting in the near future? - What functions remain challenging but offer opportunities? Answer #4: - Key opportunities include gaming and DeFi, where agents can simplify complex interfaces and provide actionable insights. - Challenges include secure payments and complex multi-chain trades, which require robust solutions like TEE. - Question #5 (38:15 - 41:54, Speaker: YK): - What feature updates or integrations are you most excited about? Answer #5: - Focus on improving memory systems, connectors (e.g., call, text, Reddit), and refining v1 while building a streamlined v2 with better modularity and tools. - Question #6 (42:50 - 44:34, Speaker: Zo): - Can I call myself an ai16zDAO member as a token holder? How else can I participate? Answer #6: - Participation includes holding tokens, contributing to GitHub, joining workgroups, and integrating personal projects with elizaOS. - Question #7 (45:32 - 47:30, Developer at PUBG: BATTLEGROUNDS) - How can we verify that AI agent outputs are untampered and authentic? Answer #7 (With Wenfeng Wang/Phala Network) - TEE (Trusted Execution Environment) ensures secure and verifiable agent outputs. - Remote attestation provides cryptographic proof of AI-generated outputs. - TEE enables secure computation and proof verification, similar to ZK proofs, ensuring trust in agent processes. - Question #8 (51:58 - 54:24, Mike at Orca ☀️): - Should we improve elizaOS directly or create complementary frameworks? Answer #8: - Collaborate on DeFi agents by integrating bots with Eliza OS, lowering entry barriers, and unlocking income opportunities for users. - Question #9 (54:35 - 58:44) - Will elizaOS expand beyond Solana? - How is 찌 G 跻 じ MBA, CFA, FRM, CFP, NGMI, HFSP, HENTAI 🛡️ being worked on? Answer #9: - elizaOS is chain-agnostic with plugins for multiple ecosystems (e.g., EVM, Solana). Unified wallet abstraction will simplify multi-chain interactions. - Dedicated teams like 찌 G 跻 じ MBA, CFA, FRM, CFP, NGMI, HFSP, HENTAI 🛡️ and Placeholder operate independently while collaborating within the DAO ecosystem, focusing on trust mechanics and KOL signal tracking for autonomous trading. - Initial token buybacks (4%+ of supply) were manual, but automation is planned to enhance efficiency and scale the ecosystem. - Question #10 (59:05 - End, University Professor): - Would you collaborate with academia on next-gen AI agent research? Answer #10 - Open to collaborations, particularly on the "marketplace of trust" research and optimizing collective intelligence models for better investments. h/t co-hosts of the event Fragmetric and #Hashed 🧡

Jun Kim

87,540 Aufrufe • vor 1 Jahr

$AMD is easily a $1,200 stock IMO| CPUs TAM 🧵 Not Financial Advice! DYOR! In this thread, I want to discuss the actual TAM for CPUs data center for just 2026, where many are giving different ranges, where I don't agree with. I will explain in detail why I disagree with these research firms and financial analysts using Math. And this thread should not be treated as Financial Advice. I'm just explaining my research and thought process so we can have a discussion. In 2024/2025, I gave out $620 PT for FY2026 was too conservative for AMD potential. At the time, It was early and many were just laughing, that PT was unrealistic and the AI world is run on GPUs only. Today, most of these folks are laughing with me. That is ok, I dont offer financial advice, and I do not need everyone to agree with me. I respect other opinions. If you enjoy this kind of thread, slap the like/repost/bookmark. If you want to support my work further and gain more in-depth analysis, consider subscribe! In early 2026, hyperscalers, enterprises, and OEMs are scrambling as Intel and AMD server CPUs are largely sold out for the year, with prices jumping 10–20% and lead times stretching from weeks to months (or longer for certain SKUs). What was once a GPU dominated story has flipped: the shift to explosive Agentic AI with its multi-step reasoning loops, tool calling, multi-agent orchestration, real-time data movement, and reinforcement learning, is dramatically tightening CPU:GPU ratios from the old training-era 1:4–8 all the way to 1:1 to 5:1 or even CPU-heavy configurations. CEOs across NVIDIA, AMD, Intel, Google, Meta, Microsoft, and public companies have been sounding the alarm on CNBC, Bloomberg, and earnings calls. CPUs are “cool again,” and in many agentic deployments they are becoming the new bottleneck alongside (or even ahead of) GPUs and custom ASICs. In 2025, roughly 12-15m AI GPUs + AI ASICs GPUs shipped, and is expect to be 15-20m units by 2026, where it suggesting Training demand is not going away. The actual TAM is structural, multiplicative demand that has already forced AMD to double its long-term server CPU TAM forecast to >$120 billion by 2030 (>35% CAGR), with Dr. Lisa Su noting Q2 2026 server CPU sales expected to surge 70%+ year-over-year and demand “far exceeding expectations.” At the same time, AMD’s secured 30–40% share of TSMC’s initial 2nm capacity (behind only Apple’s >50%) positions it to ramp Zen 6-based EPYC Venice exactly when this agentic wave hits hardest but even that aggressive five-fab 2nm expansion (with plans scaling toward 11 total advanced facilities) cannot instantly close the gap in the near-term. Supply constraints on wafers, advanced packaging, and power are compounding the squeeze, just as hyperscalers forward-buy and lock in long-term deals. 1. The actual potential TAM Various sources and institutions are giving $50-$160-$200B CPUs TAM toward 2030, and i disagree, where supply is severely behind vs Demand by at least 2-3 years or even longer by some estimates. The actual TAM will probably be 15-20m for FY2026. The typical average selling price from low to high end is $5,000 to $15,000, but due to rising memory, and different inflationary pressures on Semi, it would be more logical to think between $7,000-17,000. A. CPU:GPU Ratio at 1:1 A basic calucation at mid range =12,000 x 15-20m CPUs= $180-$240B TAM B. CPU:GPU Ratio at 5:1 = $12,000 x 75m-100m CPUs= $900B-$1.2T TAM Of course TSMC cannot even supply 20% of this massive inflection TAM in 2026. But do we think of Demand for TAM or Supply for TAM? Hence we are seeing massive 2nm Ramp from TSMC for $AMD. IMO, conservatively, I would take down 15-20% on 1:1 or $135-$192B TAM for just 2026. Im not even talking about 2030. We are just months into this, it is impossible to estimate Cagr atm, but this is 1-5 agents running tasks, I wrote a thread on 24/7 autonomous agents thread, where companies could use 50-250 agents to run tasks for them 24/7. It would require a different structural CPU:GPU to bring down the cost of token as well as handling the Orchestration bottleneck. GPUs would be useless and sit idle waiting for CPU due to highly CPU-intensive nature. The cost per Million tokens must come down more rapidly for this 50-250 autonomous agents to work, otherwise the token cost would be too enormous. Helios Rack is estimated to bring inference cost down to $0.0003-$0.0005/M tokens with 18 EPYC Venices along with 72 MI455x and other chips+ Components. A heavier or CPUs dense rack would bring down inference cost further. EPYC Verano(2027 gen 7 AI-optimized) is expected to drive inference costs meaningfully lower than the Venice baseline likely to the $0.00002–$0.00025 per million tokens range (or even sub-$0.00015 in highly optimized agentic/batch workloads). Verano have higher core counts than Venice, LPDDR5X SOCAMM2 memory support, more AI optimized and Next-Gen rack density & efficiency. 2. $AMD secured at least 30-40% of TSMC 2nm capacity and Memory from Samsung through 2028-2030. 2 2nm fabs are entering ramping phase toward 60-65k wafers per months and 5 dedicated 2nm fabs entering mass production/ramp in 2026. Will link sub threads below if you are interest for full detail. Apple is reported to secure 50%+ 2nm capacity for Iphone 18 and Mac chips and AMD secured at least 30-40% capacity while $NVDA $AVGO $ARM $AMZN $GOOGL and others are on 3nm. This broader aggressive ramp from TSMC to target up to 11 fabs is to address $AMD massive growth ahead. Where $ARM is facing massive CPUs supply constraints as they have to compete with other Mega Cap players on 3nm allocation. And $INTC is also facing supply constraints for data center CPUs and PC per management with lead times extrended to longer than 12 weeks. Dr. Su is aiming for higher than 50%+ Market share, and I believe it is achievable in 2026 or 2027 as AMD has the strongest CPUs offerings. Dr. Su did not want to take advantage of the shortage and she said during the Q1 earning call, AMD is prioritizing Units shipped while guiding margin to be inching 60%. If Jensen were in charge, I'm sure margin would be 70-75% in this kind of severe CPUs shortage condition. But that is not how Dr. Su operates for more than a decade. She wants most market share. So we will see it in revenue growth, but as TSMC ramps faster and faster, AMD Operating and FCF margin will massively improve vs prior decade. A significantly higher margin profile than before. 3. How I came up with $1,200 withint 12-18 months? At $1,200/ share, that would be around $2 Trillion MC. I expect FY2027 revenue to be $124-$144B where data center revenue dominates overall revenue. AI GPUs: I will stick to the lowest end so show u that I'm conservative at $18B for each GW vs $NVDA Rubin is $30B+ (most likely Helios Rack in the $20B+ due to memory price rising). We know deals with OpenAI and Meta are around 12GW and additional multi-customers at multi-GW scale were hinted and will be revealed as we get to July 22-23 2026 Advancing AI event. For now I will conservatively add a bit more to this model. (3-6GW Helios Rack Range) EPYC Venice is reported to be in $15,000-$20,000. However large customers will likely to enjoy $10-$12k discount. I expect AMD to be able to ramp 7m EPYC Venice for entire 2026 and 3-4m of EPYC Verano(higher price than Venice). If we take an average selling price of $10,000 to be on the conservative side. Take down another 30% to be even more conservative on projection. I like to be conservative. That would be ~ 7m EPYC CPUs(Venice + Verano) for FY2027 or 583,000 units per month or 15,000 additional 2nm wafers per month which is completely reasonable for current TSMC Ramp, and I may be too conservative here. EPYC Verano and MI500 series will also be on 2nm. AI GPUs: 3GW x $18B= $54B EPYC CPUs: $10k x 7m CPUs= $70B = Data center revenue alone is $124B Other segments= probably in the $20-$25B FY 2027. FY2027 revenue = $124-$149B At 7m EPYC CPUs for entire 2027, that would be more than 50% market share when we comp it to availability from supply side, not from total Demand. It is possible that TSMC could significantly ramp even more capacity in 2027, so we will see. Metric Q1 2026 FY2027 Gross Margin 55-56% 60-62% Operating Margin 25-26% 32-35% Net Income Margin ~22% 26-30% FCF Margin 25% 28-30% At $124-$149B Revenue FY 2027 Net Income would be $32-$44B EPS would be $20-$27 (GAAP) Non-GAAP would be $25-$31 At $1,200 a share or $2T valuation that would be: 13.4-16x Price to Sales (P/S) 38-48 P/E At this kind of growth of AI SuperCycle, I think it is very reasonable valuation. If we use today at $406/share or $661B MC: 2027 P/S = 4.4x-5.3x 2027 P/E = 13x-16x Is AMD today expensive or cheap to you? Above is already a very conservative where I trimmed 20-30% of doable units. Meaning, there could be upside if TSMC is able to ramp meaningfully like they are planning. Conclusion: A $1,200 per share valuation IMO for AMD in FY2027 is not expensive at all; it is, in fact, conservative when viewed against the structural explosion in agentic AI demand we have mapped out. With server CPU TAM potentially scaling into the $100–$200B+ range in just CPU:GPU 1:1 Ratio for just 2026. AMD positioned to capture 50%+ share thanks to its 2nm TSMC allocation advantage and full-stack leadership, the company could realistically deliver $124–149B in total revenue and $25–$31+ non-GAAP EPS. At those levels, $1,200 implies a 2027 P/E = 13x-16x. Entirely reasonable for a company that will have become the clear Inference Queen (and in many workloads the preferred) AI infrastructure provider, with operating margins expanding above 30% and tens of billions in high-margin rack-scale AI revenue. Dr. Lisa Su was right presciently so about the Agentic AI inflection all the way back to her early 2022–2023 commentary on the coming shift from pure training to inference and orchestration-heavy workloads. While the broader market only fully woke up to this in 2026 when she doubled AMD’s long-term server CPU TAM forecast to >$120B by 2030 (with >35% CAGR), Dr. Su and her team have consistently positioned the company at the center of the CPU renaissance. The explosive demand we are seeing today, sold-out lines, rising ASPs, and hyperscalers forward-buying entire gigawatts of Helios-class systems is exactly the outcome she forecasted years ago. Not Financial Advice! DYOR!

Mike

301,322 Aufrufe • vor 2 Monaten

AGI? One day, but not yet. The only AI that works well right now is the one behind the screen [12-17]. But passing the Turing Test [9] behind a screen is easy compared to Real AI for real robots in the real world. No current AI-driven robot could be certified as a plumber [13-17]. Hence, the Turing Test isn't a good measure of intelligence (and neither is IQ). And AGI without mastery of the physical world is no AGI. That’s why I created the TUM CogBotLab for learning robots in 2004 [5], co-founded a company for AI in the physical world in 2014 [6], and had teams at TUM, IDSIA, and now KAUST work towards baby robots [4,10-11,18]. Such soft robots don't just slavishly imitate humans and they don't work by just downloading the web like LLMs/VLMs. No. Instead, they exploit the principles of Artificial Curiosity to improve their neural World Models (two terms I used back in 1990 [1-4]). These robots work with lots of sensors, but only weak actuators, such that they cannot easily harm themselves [18] when they collect useful data by devising and running their own self-invented experiments. Remarkably, since the 1970s, many have made fun of my old goal to build a self-improving AGI smarter than myself and then retire. Recently, however, many have finally started to take this seriously, and now some of them are suddenly TOO optimistic. These people are often blissfully unaware of the remaining challenges we have to solve to achieve Real AI. My 2024 TED talk [15] summarises some of that. REFERENCES (easy to find on the web): [1] J. Schmidhuber. Making the world differentiable: On using fully recurrent self-supervised neural networks (NNs) for dynamic reinforcement learning and planning in non-stationary environments. TR FKI-126-90, TUM, Feb 1990, revised Nov 1990. This paper also introduced artificial curiosity and intrinsic motivation through generative adversarial networks where a generator NN is fighting a predictor NN in a minimax game. [2] J. S. A possibility for implementing curiosity and boredom in model-building neural controllers. In J. A. Meyer and S. W. Wilson, editors, Proc. of the International Conference on Simulation of Adaptive Behavior: From Animals to Animats, pages 222-227. MIT Press/Bradford Books, 1991. Based on [1]. [3] J.S. AI Blog (2020). 1990: Planning & Reinforcement Learning with Recurrent World Models and Artificial Curiosity. Summarising aspects of [1][2] and lots of later papers including [7][8]. [4] J.S. AI Blog (2021): Artificial Curiosity & Creativity Since 1990. Summarising aspects of [1][2] and lots of later papers including [7][8]. [5] J.S. TU Munich CogBotLab for learning robots (2004-2009) [6] NNAISENSE, founded in 2014, for AI in the physical world [7] J.S. (2015). On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning (RL) Controllers and Recurrent Neural World Models. arXiv 1210.0118. Sec. 5.3 describes an RL prompt engineer which learns to query its model for abstract reasoning and planning and decision making. Today this is called "chain of thought." [8] J.S. (2018). One Big Net For Everything. arXiv 1802.08864. See also patent US11853886B2 and my DeepSeek tweet: DeepSeek uses elements of the 2015 reinforcement learning prompt engineer [7] and its 2018 refinement [8] which collapses the RL machine and world model of [7] into a single net. This uses my neural net distillation procedure of 1991: a distilled chain of thought system. [9] J.S. Turing Oversold. It's not Turing's fault, though. AI Blog (2021, was #1 on Hacker News) [10] J.S. Intelligente Roboter werden vom Leben fasziniert sein. (Intelligent robots will be fascinated by life.) F.A.Z., 2015 [11] J.S. at Falling Walls: The Past, Present and Future of Artificial Intelligence. Scientific American, Observations, 2017. [12] J.S. KI ist eine Riesenchance für Deutschland. (AI is a huge chance for Germany.) F.A.Z., 2018 [13] H. Jones. J.S. Says His Life's Work Won't Lead To Dystopia. Forbes Magazine, 2023. [14] Interview with J.S. Jazzyear, Shanghai, 2024. [15] J.S. TED talk at TED AI Vienna (2024): Why 2042 will be a big year for AI. See the attached video clip. [16] J.S. Baut den KI-gesteuerten Allzweckroboter! (Build the AI-controlled all-purpose robot!) F.A.Z., 2024 [17] J.S. 1995-2025: The Decline of Germany & Japan vs US & China. Can All-Purpose Robots Fuel a Comeback? AI Blog, Jan 2025, based on [16]. [18] M. Alhakami, D. R. Ashley, J. Dunham, Y. Dai, F. Faccio, E. Feron, J. Schmidhuber. Towards an Extremely Robust Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms. Preprint arxiv 2404.08093, 2024.

Jürgen Schmidhuber

72,331 Aufrufe • vor 1 Jahr

Dean Koontz has published more than 140 novels, 74 works of short fiction, and sold more than 500 million books. Simply put, he’s one of the most prolific writers alive today. Some highlights from our chat: 1. Dare to love the English language. 2. Characters come alive when they're given free will. Instead of constraining them in an outline, let them go where they want. You know they’re alive once they start surprising you. He says: “I give the characters free will like God gave it to us.” 3. Everything a writer believes about life and death, culture and society, relationships and the self, God and nature will wind up in their books. A writer’s body of work, therefore, reveals the intellectual and emotional progress of its creator, and over time, becomes a map of their soul. 4. To think you understand the world is to be foolish in the extreme. The world is too complex for us to understand it. To see reality clearly is to be utterly enchanted by its staggering complexity. 5. Where should you look? Well, the supernatural enters the world in mundane ways, and rarely the great and glorious flashes of drama. 6. Dean writes his novels page-by-page, and doesn’t move onto the next page until he nails the existing one. There’s no messy first draft. Because of that, he’s basically done with his novels once he finishes the final page. 7. Where does a unique writing voice come from? Three places: style, perspective, and a philosophy of life. 8. Be skeptical of conventional wisdom. There’s an encyclopedia of common wisdom in publishing. All of it is common and none of it is wise. You have to become aware of that, go your own way, and just stick with it because there are so many ways you can be sent wrong based on "that's the way we always do it." 9. The aesthetic plainness of contemporary writing (and culture at large) is crushing our souls. 10. Contemporary fiction is suffering from plainness in particular. It started when writers started imitating Hemingway (who stripped his prose down but kept the mystery and underlying strangeness of the world by implication). But the imitations that came later stripped the prose down while also removing the underlying depth that made Hemingway so great. 11. Koontz Law of Writing #1: Never go inside more than one character's mind in a scene. Each one should come from a singular viewpoint. 12. Koontz Law of Writing #2: Metaphors aren't meant to dazzle readers, but to seduce them into a more intimate relationship with the story. 13. Koontz Law of Writing #3: Metaphors and similes describe a scene more colorfully than a chain of adjectives — while reinforcing the mood. The point is that you can create depth by describing things metaphorically instead of using blunt adjectives. That’s what poetry does: it uses words to say more than the word itself says, which creates a mood. 14. Great prose doesn't come from piling on adjectives. It comes from finding the perfect metaphor that does triple duty: describes the scene, reinforces the mood, and reveals something about the character. 15. The goal is for metaphors not to pop out like showmanship, but to flow into the music of the language. 16. Develop an ear for the musicality of language. 17. A book can succeed with a mediocre plot if the characters are compelling. Character is the center of good fiction. If the characters work, the story works. 18. From the afterword of his book, Watchers: “We have within us the ability to change for the better and to find dignity as individuals rather than as drones in one mass movement or another. We have the ability to love, the need to be loved, and the willingness to put our own lives on the line to protect those we love, and it is in these aspects of ourselves that we can glimpse the face of God; and through the exercise of these qualities, we come closest to a Godlike state.” I've shared the full conversation with Dean Koontz below. The YouTube video link is in the replies, and so are the links to Apple and Spotify.

David Perell

74,956 Aufrufe • vor 1 Jahr

Spent two hours with Marc Andreessen, who gave me a masterclass on how to think, learn, read, research, and write. Here's what I learned: 1. Read, read, read... then read some more. 2. Many of your best ideas will emerge in fits of rage or frustration. Channel the fury. Smash the keyboard. Lean into the passion. Torch the page with your energy. 3. Marc doesn't have much of a formal writing process. He thinks and thinks, and when epiphany strikes, he hammers out an outline as fast as possible to get his ideas on paper. Then, he turns it into a full article. 4. Marc's motto for writing and thinking: "Strong views, weakly held." Put yourself out there, but stay on the hunt for dissenting opinions from smart and respectful people. 5. Online writing tolerates and even encourages stylistic idiosyncrasies that traditional publishing would not accommodate. Lean into them. 6. The world is awash in bad content. You need to punch through. Snappy one-liners and genuine conviction are two ways to do that. 7. Marc's been reading online for as long as anybody on the planet, and the biggest thing that's surprised him is how political the Internet's become. Something changed between ~2013-2015. The Internet was once an escape from political debates. Now it's a hotbed of them. 8. Writing software is halfway between writing a novel and building a bridge. 9. Play around with communication tools. Push the limits. Doesn't matter what the rules are. When Marc felt constrained by Twitter's 140-character limit, he started replying to his own tweets and invented the Twitter thread. 10. On the quest for good ideas, surround yourself with "lateral thinkers" who can't help but come up with variant perspectives on everything they see. They won't always be right, but they're always challenge your thinking. 11. Media formats are cyclical. Nietzsche wrote in aphorisms and Twitter is aphorisms-as-a-service. Hip-hop brought back poetry. Montaigne pioneered the essay format and blogs brought them back into vogue. 12. People should write more manifestos. 13. Marc's nomination for the best living American novelist: James Ellroy. 14. GPT has revealed how much writing is pure pablum. Bland, lifeless, uninsightful, unoffensive, and not worth the price of the ink it was printed with. 15. "With GPT, every writer now has a writing partner who can do an infinite amount of grunt work without complaining." 16. "ChatGPT plagiarism is a complete non-issue. If you can't out-write a machine, what are you doing writing?" 17. Marc writes from the heart. He doesn't do much editing and likes to provide reading recommendations instead of directly citing his sources. 18. The person who writes down the plan in an organization has tremendous power. If you want to find the up-and-comers at a tech company, look into who's writing the plan. Though they may not be coming up with all the ideas, you'll know they have the energy, motivation, and skills to organize and communicate ideas in a written form. 19. Marc uses a barbell approach to consume information. He focuses on what's happening right now while also reading a lot of things that were written 10+ years ago. The content is either timely or timeless, with almost nothing in between. I've shared the full conversation below. If you'd rather listen on YouTube, Spotify, or Apple, check out the replies below. If there was an Olympic category for most insights per minute, Marc Andreessen 🇺🇸 would be a guaranteed medalist.

David Perell

2,585,676 Aufrufe • vor 2 Jahren