Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Successfully deployed Deepseek R1 Distilled 70B (AWQ) across 8x NVIDIA RTX 3080 10G GPUs, achieving 60 tokens/s with full tensor parallelism via PCIe. Total hardware cost: $6,400 This demonstrates that consumer GPUs can deliver substantial ML inference capabilities at a fraction of the cost of datacenter hardware. For perspective,...

16,219 görüntüleme • 1 yıl önce •via X (Twitter)

6 Yorum

TensorBlock profil fotoğrafı
TensorBlock1 yıl önce

For more detailed discussion and setup configurations, check out our Reddit post here:

The Information profil fotoğrafı
The Information1 yıl önce

Elon Musk claims to have finished a 100,000-strong H100 cluster in four months. How likely is that?

Anti-Federalist Caesar profil fotoğrafı
Anti-Federalist Caesar1 yıl önce

@nvidia cc @__tinygrad__

אלף profil fotoğrafı
אלף1 yıl önce

@nvidia Awesome!

Novel Engineer profil fotoğrafı
Novel Engineer1 yıl önce

Why not 3090 - twice the vram

TensorBlock profil fotoğrafı
TensorBlock1 yıl önce

Thanks for the feedback! While 3090s with NVLink would offer higher bandwidth, this experiment focuses on validating distributed inference with consumer GPUs - exploring possibilities for widespread LLM deployment. Excited to share more findings soon.

Benzer Videolar

Groq is serving the fastest responses I've ever seen. We're talking almost 500 T/s! I did some research on how they're able to do it. Turns out they developed their own hardware that utilize LPUs instead of GPUs. Here's the skinny: Groq created a novel processing unit known as the Tensor Streaming Processor (TSP) which they categorize as a Linear Processor Unit (LPU). Unlike traditional GPUs that are parallel processors with hundreds of cores designed for graphics rendering, LPUs are architected to deliver deterministic performance for AI computations. The LPU's architecture is a departure from the SIMD (Single Instruction, Multiple Data) model used by GPUs and favor a more streamlined approach that eliminate the need for complex scheduling hardware. This design allows every clock cycle to be utilized effectively, ensuring consistent latency and throughput. For developers, this means that performance can be precisely predicted and optimized which is critical in real-time AI applications. Energy efficiency is another area where LPUs shine. By reducing the overhead of managing multiple threads and avoiding the underutilization of cores, LPUs can deliver more computations per watt. Groq's innovative chip design allows multiple TSPs to be linked together without the traditional bottlenecks found in GPU clusters making them extremely scalable. This enables linear scaling of performance as more LPUs are added simplifying the hardware requirements for large-scale AI models and making it easier for developers to scale their applications without rearchitecting their systems. So what does this all mean? LPUs could provide a massive improvement compared to GPUs for serving AI applications in the future! If anything it will be great to have alternative high performing hardware since A100s and H100s are so in demand

Jay Scambler

318,228 görüntüleme • 2 yıl önce

Today we announced our new Fairwater datacenter in Atlanta, connected with our first Fairwater site in Wisconsin and our broader Azure footprint to create the world’s first AI superfactory. Fairwater exemplifies our vision for a fungible fleet: infra that can serve any workload, anywhere, on fit-for-purpose accelerators and network paths, with maximum performance and efficiency. AI workloads have evolved beyond large-scale pre-training. Today, they encompass fine-tuning, reinforcement learning (RL), synthetic data generation, evaluation pipelines, and more. Fairwater is built to support this full lifecycle: Max density: Fairwater’s two-story design and liquid cooling system lets us place racks in three dimensions and pack them with GPUs as densely as possible, minimizing cable runs and improving latency and effective bandwidth. Fleet: Each Fairwater DC can integrate hundreds of thousands of the latest NVIDIA GPUs into a single coherent cluster. This provides flexible infra that can support the full spectrum of workloads, and ensure no GPU is left unnecessarily idle. And that’s on top of the more than 100,000 GB300s coming online this quarter alone for inference across the rest of our fleet. For us, it’s all about turning every gigawatt into the maximum number of useful tokens. Not every GW is created equal! Planet-scale: Every Fairwater DC will connect through our continent-spanning AI WAN to prior generations of AI supercomputers, forming a truly fungible pool of compute. This enables developers to scale beyond the capacity of a single site and dynamically land workloads on the right infra for their needs. Together, these innovations let us bring together different generations of silicon and AI systems across DCs and geos into a single elastic system that scales seamlessly across training and inference workloads And this elastic AI capacity is all available alongside all the other cloud services (compute, storage, databases, app services) that AI agents and workloads need. This is what we mean when we talk about building a fungible fleet – a single, unified platform that pushes the limits of performance per watt and per dollar. Read more:

Satya Nadella

908,065 görüntüleme • 9 ay önce

If intelligence is the log of compute… it starts with a lot of compute! And that’s why we’re scaling our GPU fleet faster than anyone else. Just last year, we added over 2 gigawatts of new capacity – roughly the output of 2 nuclear power plants. And today we’re going further, announcing the world's most powerful AI datacenter, located in southeastern Wisconsin. Fairwater is a seamless cluster of hundreds of thousands of NVIDIA GB200s, connected by enough fiber to circle the Earth 4.5 times. It will deliver 10x the performance of the world’s fastest supercomputer today, enabling AI training and inference workloads at a level never before seen. For AI training workloads, you need compute at exponential scale. That’s why we designed the datacenter, GPU fleet, and network together as one integrated system. This ensures a single job can run from day 1 at exponential scale across thousands of GPUs. Fairwater uses a liquid-cooled closed-loop system for cooling GPUs that requires zero water for operations after construction. And we’re matching all of the energy that is consumed with renewable sources. And of course, it is just one of several similar sites we’re lighting up across our 70+ regions. We have multiple identical Fairwater datacenters under construction in other locations across the US, in addition to our AI infrastructure already deployed in over 100 datacenters around the world, powering model training, test-time compute, RL tuning, and real-time inference at global scale. Too often during times like this, people go with the current and only later wonder, how did we get here? With Fairwater, we're charting a new path: doing the hard engineering work, bringing compute, network, and storage into one highly scaled cluster, and designing closed-loop energy systems to meet real-world computing needs. And partnering with local communities to ensure it's thoughtfully done in a way that is sustainable, creates new jobs, and expands opportunity. We are thrilled to see this take hold in Wisconsin, and we are just getting started.

Satya Nadella

2,024,356 görüntüleme • 11 ay önce

$AMD| $META is using $GOOGL to negotiate 🧵 The Ironwood pod is 5.1–10x more expensive annually ($148.3 million ÷ $14.87–$29.04 million) and 5.1–10x more expensive monthly ($12.36 million ÷ $1.24–$2.42 million) than renting 15 MI450 racks for equivalent compute. The rapidly evolving landscape of artificial intelligence infrastructure presents a complex interplay of technological innovation, market dynamics, and strategic maneuvering among major players. Recent leaked information suggesting that Meta Platforms ($META) might work with Google's Tensor Processing Unit (TPU) in 2027 has sparked speculation about its true intent. This leak is likely a strategic move by Meta to negotiate more favorable terms with AMD , leveraging the competitive dynamics of the AI hardware market to optimize its substantial investment in AI infrastructure. By examining the key elements of this scenario Meta's investment strategy, the comparative advantages of AMD's MI450 and Google's Ironwood TPU, and the broader market context; we can discern the potential beneficiaries and the strategic implications of this information. Meta's aggressive pursuit of AI capabilities is underscored by its planned expenditure of $66-72 billion on AI infrastructure in 2025, with expectations to escalate significantly in 2026. This investment is part of a broader strategy to build "titan clusters" like Prometheus, which are projected to reach 1 gigawatt of compute power by 2026. Such a scale of investment reflects Meta's recognition of the critical role that AI will play in its future growth, particularly in enhancing its social media platforms and developing new AI-driven applications. However, the financial burden of this infrastructure buildout necessitates a careful consideration of cost-effectiveness and scalability, which brings us to the leaked information about potential collaboration with Google's Ironwood TPU. Google's Ironwood TPU, introduced as the seventh-generation ASIC optimized for TensorFlow-based inference, represents a high-cost, cloud-locked solution priced at $445 million per pod (9,216 chips) over three years. This model, while offering significant performance gains and power efficiency, is tailored for pod-scale deployment and integrated with Google's cloud services, limiting flexibility and increasing costs for customers. In contrast, AMD's MI450 GPU, priced at $30,000–$40,000 per unit, provides a modular, open ROCm ecosystem that delivers comparable compute capacity at a fraction of the cost. Renting 15 MI450 racks could achieve similar 42+ exaFLOPS inference compute at 5–10x lower cost than renting a single Ironwood pod, underscoring AMD's competitive edge in terms of total cost of ownership (TCO). The leaked information about Meta's potential TPU deployment in 2027, therefore, can be interpreted as a negotiating tactic rather than a definitive shift in strategy. By signaling interest in Google's solution, Meta may be attempting to pressure AMD into offering more favorable terms/prices for 5-10GW. This tactic aligns with Meta's broader goal to finance most of its AI spend internally while exploring partnerships that can reduce costs and enhance flexibility. The post's emphasis on MI450's TCO advantage and its partnerships with major players like OpenAI, Microsoft, and Meta itself suggests that AMD is a critical component of Meta's AI infrastructure strategy. The threat of working with Google's TPU could prompt AMD to reassess its pricing, provide additional support, or offer incentives to retain Meta as a customer, thereby securing or expanding its market share. From a logical standpoint, Meta stands to benefit the most from this strategy. As a major buyer in a high-stakes market projected to surpass $1 trillion in annual spending by 2030, Meta's negotiating power is significant. The leaked information could lead to substantial cost savings on its $66-72 billion investment, enhancing its financial flexibility and allowing for further investment in AI capabilities. Moreover, this tactic reinforces Meta's position as a leader in the AI infrastructure race, potentially attracting more external financing for its data center projects and strengthening its competitive stance against other hyperscalers like Amazon and Microsoft. AMD could also benefit from this scenario. The negotiation pressure might lead to small short-term concessions, but it could also solidify long-term partnerships with Meta, ensuring continued demand for MI450 and other AI hardware solutions. Initially Meta's 42% allocation to AMD MI300X and its partnerships with Oracle, Dell, and HP indicates a deep integration of AMD's technology into Meta's infrastructure, which could be leveraged to maintain this relationship. For AMD, retaining Meta as a large key customer is crucial to capturing a larger share of the rapidly growing data center infrastructure market, driven by the insatiable demand for AI compute power. Google, on the other hand, faces a more limited benefit from this leaked information. While securing Meta as a customer would reinforce its position in the AI hardware market, the high cost and ecosystem lock-in of the Ironwood TPU might deter Meta from fully committing to this solution. The leaked information could prompt Google to reconsider its pricing or ecosystem strategy to remain competitive, but the immediate impact is likely to be minimal compared to the potential gains for Meta and AMD. Investors and market analysts also stand to benefit from this information, as it provides insights into the competitive dynamics of the AI hardware market. Adjustments in portfolios based on anticipated shifts in market share and profitability could lead to opportunities for those who correctly anticipate outcomes. The negotiation dynamic might introduce volatility, but it also highlights the strategic importance of cost-effective solutions in the AI infrastructure space. Lastly, the leaked information about Meta potentially working with Google's TPU in 2027 is likely a strategic move to negotiate with AMD, leveraging the competitive landscape to optimize its AI infrastructure investment. Meta, as the primary negotiator, stands to gain the most by securing better terms from AMD, reducing costs, and enhancing its financial flexibility. AMD, while initially at risk, could benefit from retaining a key customer and solidifying its market position. Google faces limited immediate benefits but may need to adapt its strategy to remain competitive. This scenario underscores the complex interplay of technology, market dynamics, and strategic maneuvering in the AI hardware market, where cost-effectiveness and scalability are paramount. As the data center infrastructure market continues to grow, the outcomes of such negotiations will shape the future of AI development and deployment.

Mike

182,273 görüntüleme • 9 ay önce