Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇

12,875 görüntüleme • 2 yıl önce •via X (Twitter)

5 Yorum

Sanchit Gandhi profil fotoğrafı
Sanchit Gandhi2 yıl önce

Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Mobius Labs profil fotoğrafı
Mobius Labs2 yıl önce

Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Praveen kumar profil fotoğrafı
Praveen kumar2 yıl önce

Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Sentient profil fotoğrafı
Sentient2 yıl önce

If u have Nvidia card ? Or regardless on phones as well?

Rekt Stoic profil fotoğrafı
Rekt Stoic2 yıl önce

Can something like this be applied to any transformer? Like flan-t5?

Benzer Videolar