Loading video...
Video Failed to Load
Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇
12,875 views • 2 years ago •via X (Twitter)
5 Comments

Sanchit Gandhi2 years ago
Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Mobius Labs2 years ago
Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Praveen kumar2 years ago
Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Sentient2 years ago
If u have Nvidia card ? Or regardless on phones as well?

Rekt Stoic2 years ago
Can something like this be applied to any transformer? Like flan-t5?
