Загрузка видео...
Не удалось загрузить видео
Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇
12,875 просмотров • 2 лет назад •via X (Twitter)
Комментарии: 5

Sanchit Gandhi2 лет назад
Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Mobius Labs2 лет назад
Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Praveen kumar2 лет назад
Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Sentient2 лет назад
If u have Nvidia card ? Or regardless on phones as well?

Rekt Stoic2 лет назад
Can something like this be applied to any transformer? Like flan-t5?
