Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Whisper 🤝 Torch Compile Speed-up inference by a factor of 4x with just 2 additional lines of code No accuracy degradation. No additional training. Just kernels 😍 Code example in thread 👇

12,875 Aufrufe • vor 2 Jahren •via X (Twitter)

5 Kommentare

Profilbild von Sanchit Gandhi
Sanchit Gandhivor 2 Jahren

Google Colab example: Enabling compile requires 2 lines of additional code: 1. Apply the `torch.compile` transformation 2. Set the static kv cache

Profilbild von Mobius Labs
Mobius Labsvor 2 Jahren

Great to see it live. Linking to that we shared a few months ago, which has detailed benchmarks and the impact of quantization. Hope it was useful.

Profilbild von Praveen kumar
Praveen kumarvor 2 Jahren

Love it. The fact that it can run without Ampere Architecture makes it more worthwhile.

Profilbild von Sentient
Sentientvor 2 Jahren

If u have Nvidia card ? Or regardless on phones as well?

Profilbild von Rekt Stoic
Rekt Stoicvor 2 Jahren

Can something like this be applied to any transformer? Like flan-t5?

Ähnliche Videos