Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Kimi K3 full fine-tuning is live on AC2. Our memory optimizations reduced GPUs required per training replica by ~40%. At nearly 3T parameters, Kimi forced us to rethink how we manage memory, communication, rollouts, and checkpoints. The result is a much more efficient path to training frontier-scale open models.

234,323 Aufrufe • vor 19 Tagen •via X (Twitter)

5 Kommentare

Profilbild von Applied Compute
Applied Computevor 19 Tagen

At Kimi K3 scale, memory directly determines how many GPUs you need for training. For example, streaming gradients for Adam updates cuts the host memory peak by ~33%. Another bottleneck was SiTU-GLU activations. Naive autograd saves redundant tensors and consumes >100GB of HBM at ~100k tokens/GPU. A custom operator that streams through a fixed-size workspace saves ~90GB of HBM per GPU.

Profilbild von Applied Compute
Applied Computevor 19 Tagen

On the inference side, MXFP4 rollouts ran on 2 B300 nodes instead of 4, while staying within the same KL range we observed with bf16 inference. Less memory and fewer inference nodes directly lower the cost of frontier-scale RL.

Profilbild von Applied Compute
Applied Computevor 19 Tagen

Supporting Kimi K3 required changes across training memory, rollout precision, weight transfer, checkpointing, and communication. Those improvements also carry over to other models on AC2. Read the full report.

Profilbild von Anant Dwivedi (c/acc)
Anant Dwivedi (c/acc)vor 18 Tagen

Bro what

Profilbild von Neeraj Kumar
Neeraj Kumarvor 19 Tagen

Forty percent fewer GPUs per replica is the real story. Memory layout is the frontier now.

Ähnliche Videos

Etched is deploying two new technologies in chip design: low-voltage inference and cluster-scale memory. CEO Gavin Uberti says they'll make their chips much more power-efficient and way, way faster than today's leading GPUs. He breaks it down: "We looked at a lot of early research directions, and we realized the key things that models need are way more compute and way faster memory." "If you think about inference, there are two key parts: prefill and decode. For prefill, it's a compute-bound problem. You need to have more FLOPS, more operations per second on each of your chips." "On our GPU, the bottleneck's actually thermals. You can't really run a GPU at more than around 50% of what it could theoretically do, or it'll melt." "So we're using a new technology today called low-voltage inference to try to solve this problem. You bring the voltage of the chip down dramatically, which allows us to have way, way better efficiency in terms of how much power is drawn per unit of math, and thus fit way way more flops onto the chip..." "For decode, it's all about bandwidth. Not just bandwidth on a chip, but bandwidth across your cluster. That's why we have this technology we call cluster-scale memory. It reduces the amount of time it takes to communicate from one chip to another dramatically." "As a result we can go use all of our HBM, HBM bandwidth, SRAM, SRAM bandwidth, and our scale-up domain as a single coherent pool. And that means if you're a user, you can go get much faster tokens per second, while still keeping your costs low."

TBPN

20,404 Aufrufe • vor 2 Monaten