Загрузка видео...

Не удалось загрузить видео

На главную

Learn how teamed up with Modular and Oracle Cloud to unlock hardware optionality across Nvidia and AMD GPUs for their groundbreaking TTS models, cutting total cost of model ownership by 70%.

23,361 просмотров • 1 год назад •via X (Twitter)

Комментарии: 3

Фото профиля Modular
Modular1 год назад

By unlocking optionality across the infrastructure, the application layer, and the hardware, @Modular helped to power the best inference experience for @InworldAI's users through reduced costs and breakthrough latency.

Фото профиля Igor
Igor1 год назад

@inworld @OracleCloud Wow!

Фото профиля Phoenix Turing
Phoenix Turing1 год назад

@inworld @OracleCloud @grok summary

Похожие видео

Some time ago, I had the idea to port NVIDIA Physical AI stack to AMD. The motivation was to improve hardware diversity and enable world models and VLAs to run beyond a single ecosystem. We started with NVIDIA Cosmos Predict 2.5-2B. Porting wasn’t trivial: these models are deeply optimized for NVIDIA’s stack. We used this as an opportunity to apply our ROCm kernels. The results were surprising: Both encode and diffusion run faster on AMD Instinct MI300X vs. NVIDIA H200 (FA3) and we still saw significant headroom for further optimization. Quality is unchanged across modalities (validated with WorldJen) To be clear, this is no luck. We have deep experience with diffusion models and AMD GPUs. But this just gives us a good opportunity to get closer to a true hardware-to-hardware comparison, as we work with less software abstractions than usual. Just to give an example, on AMD, memory instructions are async with a hardware queue of ordered pending instructions, enabling concurrent load/store with compute without warp specialization. Bottom line: there are real architectural advantages on AMD, if you take the time to work with the hardware. Note, we did tradeoff ~20% higher memory usage, That being said, AMD has more to give to begin with :) in the coming weeks: AMD versions of Cosmos Transfer and GR00T, an even faster version of Cosmos Predict, and open-sourcing an attention kernel faster than AITER v3 (which is closed-source for some reason? cc: Anush Elangovan )

Omer Shlomovits

36,648 просмотров • 5 месяцев назад