Загрузка видео...
Не удалось загрузить видео
1/ Can AI scaling continue through 2030? We examine whether constraints on power, chip manufacturing, training data, or data center latencies might hinder AI growth. Our analysis suggests that AI scaling can likely continue its current trend through 2030.
1,507,995 просмотров • 1 год назад •via X (Twitter)
Комментарии: 12

2/ Training state-of-the-art AI models requires a massive amount of computation, which is growing by 4x every year. If this trend continues, we will see training runs 10,000x larger than GPT-4 by the end of the decade.

3/ Achieving a 10,000x scale-up requires immense resources. We analyzed key bottlenecks: power, chips, data, and latencies. For each, we project potential scale using semiconductor foundries' expansion plans, electricity providers' forecasts, industry data, and our own research.

4/ Power. Meta's Llama 3.1 405B training used 16,000 H100 GPUs, consuming about 30MW. By 2030, the largest training runs could demand 5GW of power, after accounting for energy efficiency gains and increased training durations.

5/ A data center with a >1GW capacity would be unprecedented, but is in line with stated industry plans and supplier’s projections. Distributed training runs across US states could further surpass this, doubling or even 10x-ing the power a single campus could muster.

6/ Chip production. 16,000 H100s is far from the tens of millions of chips needed to scale 10,000x beyond GPT-4. While GPU production is constrained by advanced packaging and high-bandwidth memory, foundries like TSMC are on track to expand their capacity and meet this demand.

7/ Planned scale-ups and efficiency gains could enable 100M H100-equivalent GPUs to be dedicated to a 9e29 FLOP training run by 2030. This accounts for GPU distribution among labs and inference use. This could be much higher if most of TSMC’s top wafers went to AI.

8/ Training data. All indexed web text could be enough for several thousand-fold larger training runs today. By 2030, this stock of data may have grown enough for a 10,000x scale-up.

9/ Multimodal data (image, video, audio) could expand AI training scale by ~10x. Synthetic data shows promise in domains like coding/math, but risks model collapse. Synthetic data could enable multiple orders of magnitude more scaling, but with increased compute costs.

10/ Latency. As models grow, they need more sequential ops per example they are trained on, limiting the size of training runs. Increasing batch size helps, but this has diminishing returns.

11/ On modern hardware, these latency constraints would keep runs to ~1e32 FLOP. Exceeding this would require new network designs or lower-latency hardware.

12/ Despite these significant bottlenecks in AI training, our estimates suggest they won't significantly slow the growth rate of training runs. This suggests we could see another major scale-up—comparable to the jump from GPT-2 to GPT-4—by 2030.

13/ You can learn more about each of the bottlenecks and our assumptions in our report at the link below!



