Загрузка видео...

Не удалось загрузить видео

На главную

Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇

12,875 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 5

Фото профиля Sanchit Gandhi
Sanchit Gandhi2 лет назад

Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Фото профиля Mobius Labs
Mobius Labs2 лет назад

Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Фото профиля Praveen kumar
Praveen kumar2 лет назад

Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Фото профиля Sentient
Sentient2 лет назад

If u have Nvidia card ? Or regardless on phones as well?

Фото профиля Rekt Stoic
Rekt Stoic2 лет назад

Can something like this be applied to any transformer? Like flan-t5?

Похожие видео