Video yükleniyor...
Video Yüklenemedi
Yeah. I have ChatGPT at home. Not a silly 7b model. A full-on 65B model that runs on my pi cluster, watch how the model gets loaded across the cluster with mmap and does round-robin inferencing 🫡 (10 seconds/token) (sped up 16x)
453,068 görüntüleme • 3 yıl önce •via X (Twitter)
10 Yorum

for your passive perusal, docs, tutorial gist:

lol. was following the gh issue. didn't realize it was you. good work

couldn't have found the issue without your help <3

Would adding more rpi's speed it up or is the maximum value in only adding enough to fit the model in ram

The latter, fortunately That means as bigger models come by you're looking at approximately a 100$ peripheral add on for 8 gigs Also larger context windows if the model supports it I have a 3090 and still can't run a 65B without offloading

Yooo, how many nodes for this? I was thinking of picking up some 16GB OrangePi’s for a slow cluster

6 X 8 gig raspberry pis 🫡 It's memory bandwidth thats the limiter, cpu and disk seem to be easier/cheaper Go for good mhz ram in the soc 🤍

Using the new llama.cpp changes?

Yup! Distributed Mpi 🫡

@yacineMTB idk way but the aesthetic of this is perfect, like i feel like im coding with gasoline or something
