Loading video...

Video Failed to Load

Go Home

Argmax now runs on Google Tensor TPU, the first-ever SDK to harness this edge inference accelerator! Tensor TPU enabled us to deploy billion-scale transformers reliably on Pixel phones without impacting battery life or resource contention with traditional workloads.

27,205 views • 4 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Google Ironwood TPU Memory Hierarchy in 9 levels by hand ✍️ 1. Bit – The most basic unit of information, the on–off decision from which every number, tensor, and model state is ultimately constructed. 2. FP8 (1×8 → 8 bits) – Eight bits are grouped to form a floating-point value, typically used for inference, where reduced precision is a deliberate trade-off to maximize throughput and efficiency. 3. BF16 (×2 → 16 bits) – Two FP8-scale chunks are combined to gain more dynamic range and stability, while still staying friendly to high-throughput hardware. 4. Tensor tile (×1024 → 1K) – Data moves through the chip in blocks of 1024 values at a time, defining the granularity at which tensors are fetched and manipulated. 5. Matrix Multiplication Unit (MXU) (×64 → 64K) – A systolic array where matrix multiplication is not abstract but physical, with tensor tiles flowing through fixed hardware to achieve the highest possible throughput. 6. Vector Memory (VMEM) (×2048 → 128M) – On-chip working memory that holds activations, partial results, and intermediates, sized specifically to keep the systolic array busy without stalling. 7. Common Memory (CMEM) (×8 → 1 GB) – A small but critical shared memory sitting between VMEM and HBM, used for staging, accumulation, synchronization, and cross-lane coordination. 8. HBM (×96 → 96 GB) – Off-chip high-bandwidth memory where model weights and large states live, implemented as HBM3e with 16 stacks at 6 GB each, for a total of 96 GB. 9. Dual-Die (x2 → 192GB) – Two tightly coupled compute dies operate as a single logical accelerator, each with its own local HBM, effectively doubling memory capacity and bandwidth while allowing tensors and activations to stream seamlessly across dies as if they lived on one chip. I created this drawing for this week's seminar. I’ll take you through these 9 levels in a beginner-friendly way by hand ✍️. RSVP 👉

Tom Yeh

30,489 views • 7 months ago

On Friday, I hosted a Space with Jonathan Ross, the founder and CEO of Groq Inc - a company I invested in that is building custom chips for AI inference. Jonathan, a former high-school dropout, entered the chip industry while working on ad optimization at Google’s New York office. Jonathan overheard the speech recognition team complaining that they couldn't get enough compute. These were the early days of AI, and machine learning wasn’t really a thing yet. So he asked for some budget from Google and started putting together a chip-based machine learning accelerator for them. During the day, Jonathan would work in the normal ads part of the business, and at night, he would work with the accelerator team. After winning approval from Google, Jonathan and his team built a new chip called the Tensor Processing Unit, and began deploying it across Google’s data centers within a year. The TPU was a huge success within Google, eventually underpinning more than 50% of all of Google’s compute power. When the other hyper-scalers learned of this success, they tried to hire Jonathan to build custom chips for them too. During this process, it became increasingly clear to Jonathan that a gap would emerge between companies that had access to next-gen compute and companies that didn’t. So he founded Groq and set out to build a chip that would be available to everyone. I led Groq’s founding investment in 2016, and since then, Jonathan and his team have developed several types of AI hardware including the Language Processing Unit (LPU), a new type of silicon that is hyper-efficient at running inference for LLMs. In our conversation on Friday, we discussed the founding story of Groq, what you need for great AI hardware, large language models, and some of the implications for the key players in AI. It’s one of the most interesting conversations I’ve had on AI with a lot of learnings. You can listen to our conversation below:

Chamath Palihapitiya

326,752 views • 2 years ago

Today we announced our new Fairwater datacenter in Atlanta, connected with our first Fairwater site in Wisconsin and our broader Azure footprint to create the world’s first AI superfactory. Fairwater exemplifies our vision for a fungible fleet: infra that can serve any workload, anywhere, on fit-for-purpose accelerators and network paths, with maximum performance and efficiency. AI workloads have evolved beyond large-scale pre-training. Today, they encompass fine-tuning, reinforcement learning (RL), synthetic data generation, evaluation pipelines, and more. Fairwater is built to support this full lifecycle: Max density: Fairwater’s two-story design and liquid cooling system lets us place racks in three dimensions and pack them with GPUs as densely as possible, minimizing cable runs and improving latency and effective bandwidth. Fleet: Each Fairwater DC can integrate hundreds of thousands of the latest NVIDIA GPUs into a single coherent cluster. This provides flexible infra that can support the full spectrum of workloads, and ensure no GPU is left unnecessarily idle. And that’s on top of the more than 100,000 GB300s coming online this quarter alone for inference across the rest of our fleet. For us, it’s all about turning every gigawatt into the maximum number of useful tokens. Not every GW is created equal! Planet-scale: Every Fairwater DC will connect through our continent-spanning AI WAN to prior generations of AI supercomputers, forming a truly fungible pool of compute. This enables developers to scale beyond the capacity of a single site and dynamically land workloads on the right infra for their needs. Together, these innovations let us bring together different generations of silicon and AI systems across DCs and geos into a single elastic system that scales seamlessly across training and inference workloads And this elastic AI capacity is all available alongside all the other cloud services (compute, storage, databases, app services) that AI agents and workloads need. This is what we mean when we talk about building a fungible fleet – a single, unified platform that pushes the limits of performance per watt and per dollar. Read more:

Satya Nadella

908,065 views • 9 months ago