Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Full F16 precision 34B Code Llama at >20 t/s on M2 Ultra

1,158,151 görüntüleme • 3 yıl önce •via X (Twitter)

10 Yorum

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

"Wait, Georgi, how is this even possible?" you might ask. After all, the M2 Ultra only has 800GB/s bandwidth. Other people normally need 4 high-end GPUs to do this The answer is: Speculative Sampling

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

In this example we demonstrate unbiased F16 34B sampling with the help of a Q4 7B quantum "draft" model (Code Llama 7B) Individually, the speed of these models are: - F16 34B: ~10 t/s - Q4 7B: ~80 t/s However, in combination with speculative sampling we achieve ~20 t/s

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

The speed of course can vary depending on the content that is generated. But the approach seems to work quite well for code generation as most of the tokens are correctly guessed by the draft model Use cases with grammar sampling might also benefit significantly from this

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

Here is what a classic F16 sampling looks like without the speculative help

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

Here are a couple of more examples of speculative sampling

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

Meta should have release a couple of (1B and 3B) drafter models with the Code Llama release. Is it too late for them to train them or we have to wait for v2 🤔

Dan Siroker profil fotoğrafı
Dan Siroker3 yıl önce

Well done, Georgi! Is that the GPT-5 source code on your other tab? 😂

Georgi Gerganov profil fotoğrafı
Georgi Gerganov3 yıl önce

it's top secret 😉

Alex Volkov (Thursd/AI) profil fotoğrafı
Alex Volkov (Thursd/AI)3 yıl önce

This is incredible, things are happening so fast! I wonder if this is at all usable on an M1 🤔 Will mention on @thursdai_pod 👏

Alex Skryl profil fotoğrafı
Alex Skryl3 yıl önce

@dylan522p look at the GPU-Poor go!

Benzer Videolar