Video wird geladen...
Video konnte nicht geladen werden
Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇
12,875 Aufrufe • vor 2 Jahren •via X (Twitter)
5 Kommentare

Sanchit Gandhivor 2 Jahren
Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Mobius Labsvor 2 Jahren
Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Praveen kumarvor 2 Jahren
Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Sentientvor 2 Jahren
If u have Nvidia card ? Or regardless on phones as well?

Rekt Stoicvor 2 Jahren
Can something like this be applied to any transformer? Like flan-t5?
