Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Yeah. I have ChatGPT at home. Not a silly 7b model. A full-on 65B model that runs on my pi cluster, watch how the model gets loaded across the cluster with mmap and does round-robin inferencing 🫡 (10 seconds/token) (sped up 16x)

453,068 görüntüleme • 3 yıl önce •via X (Twitter)

10 Yorum

Loki (cute/acc) profil fotoğrafı
Loki (cute/acc)3 yıl önce

for your passive perusal, docs, tutorial gist:

kache profil fotoğrafı
kache3 yıl önce

lol. was following the gh issue. didn't realize it was you. good work

Loki (cute/acc) profil fotoğrafı
Loki (cute/acc)3 yıl önce

couldn't have found the issue without your help <3

Teknium (e/λ) profil fotoğrafı
Teknium (e/λ)3 yıl önce

Would adding more rpi's speed it up or is the maximum value in only adding enough to fit the model in ram

Loki (cute/acc) profil fotoğrafı
Loki (cute/acc)3 yıl önce

The latter, fortunately That means as bigger models come by you're looking at approximately a 100$ peripheral add on for 8 gigs Also larger context windows if the model supports it I have a 3090 and still can't run a 65B without offloading

ponzi enjoyer (z/acc) profil fotoğrafı
ponzi enjoyer (z/acc)3 yıl önce

Yooo, how many nodes for this? I was thinking of picking up some 16GB OrangePi’s for a slow cluster

Loki (cute/acc) profil fotoğrafı
Loki (cute/acc)3 yıl önce

6 X 8 gig raspberry pis 🫡 It's memory bandwidth thats the limiter, cpu and disk seem to be easier/cheaper Go for good mhz ram in the soc 🤍

TDM (e/λ) (vibe coder 💫) profil fotoğrafı
TDM (e/λ) (vibe coder 💫)3 yıl önce

Using the new llama.cpp changes?

Loki (cute/acc) profil fotoğrafı
Loki (cute/acc)3 yıl önce

Yup! Distributed Mpi 🫡

andy profil fotoğrafı
andy3 yıl önce

@yacineMTB idk way but the aesthetic of this is perfect, like i feel like im coding with gasoline or something

Benzer Videolar