Загрузка видео...
Не удалось загрузить видео
Motel Owner, Dylan Patel, talks about GPT6's architecture which will continue the trend of increasing MoE sparsity. This will require even wider expert parallelism for MoE dispatch & MoE combine collectives during decode phase.
37,640 просмотров • 6 месяцев назад •via X (Twitter)
Комментарии: 0
Нет доступных комментариев
Здесь появятся комментарии из оригинального поста
