Загрузка видео...
Не удалось загрузить видео
🚨 YOUR GPU IS PROBABLY WASTING MORE THAN YOU THINK. vLLM is built to squeeze far more useful work out of your GPU when serving LLMs. Running an LLM at scale isn’t just about having a powerful GPU. The real problem is how efficiently you use its memory and... show more
14,138 просмотров • 3 дней назад •via X (Twitter)
Комментарии: 0
Нет доступных комментариев
Здесь появятся комментарии из оригинального поста
Похожие видео
1:27
Sensitive content
768GB of VRAM in one 4U server. That number sounds ridiculous until you remember it’s not one 768GB GPU. This Gigabyte G494-ZB4 packs: 8× RTX PRO 6000 Blackwell 768GB VRAM 2× EPYC 9554 1TB DDR5 ECC And that’s the part people often miss with multi-GPU systems. You can have almost 1TB of GPU memory, but the model still has to be split across 8 separate memory pools. So I think the next AI hardware race isn't just about how much memory you can put in a server. It's about how close that memory can behave to one giant pool. That’s a much harder problem to solve.
KyzoroX
86,494 просмотров • 3 дней назад
