Video yükleniyor...
Video Yüklenemedi
Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇
12,875 görüntüleme • 2 yıl önce •via X (Twitter)
5 Yorum

Sanchit Gandhi2 yıl önce
Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Mobius Labs2 yıl önce
Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Praveen kumar2 yıl önce
Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Sentient2 yıl önce
If u have Nvidia card ? Or regardless on phones as well?

Rekt Stoic2 yıl önce
Can something like this be applied to any transformer? Like flan-t5?
