Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I spent my summer building TinyTPU : An open source ML inference and training chip. it can do end to end inference + training ENTIRELY on chip. here's how I did it👇:

356,700 görüntüleme • 1 yıl önce •via X (Twitter)

48 Yorum

surya profil fotoğrafı
surya1 yıl önce

I worked on this with @evanliin, @XanderChin, and @kennykgguo — incredibly smart people! check out our article to see how our chip works and how we went about developing it: you can also find the code here and play with it yourself:

surya profil fotoğrafı
surya1 yıl önce

our first step was to decide the scale of this project. we decided to target the simplest possible neural network — the XOR problem. however, we still wanted to make this scalable so a core design philosophy for us was to ensure all of our mechanisms could scale to larger networks.

surya profil fotoğrafı
surya1 yıl önce

before we got to designing any hardware, we started off with properly understanding the math behind MLPs. we worked out the math by hand for inference and training of our network.

surya profil fotoğrafı
surya1 yıl önce

then we build the heart of any TPU: the systolic array! each processing element (PE) in the systolic array performs a multiply-accumulate in one clock cycle. when connected in a grid, multiple output matrix elements compute simultaneously. this allows us to very efficiently perform matrix multiplication, the most compute-heavy operation in neural networks!

surya profil fotoğrafı
surya1 yıl önce

the next operation is adding the bias, for which we made a bias module. since this is an element wise operation, we placed 2 bias modules right under the systolic array (one for each column). we structured the activation module similarly — we chose the leaky ReLU and placed one module under each column of the systolic array.

surya profil fotoğrafı
surya1 yıl önce

however, we had a big problem with our systolic array — computation at the end of every layer to load in the weights of the new layer (this was negligible with a 2x2 systolic array, but it would very inefficient if we scaled it) to solve this, we introduced double buffering to the systolic array. Each PE in the systolic array would have two buffers — an active and inactive buffer. while the outputs of the previous layer are being computed, we can load in the weights of the next layer in the inactive buffers. once the previous layer is finished, we can instantly start computation on the next layer. this nearly DOUBLES the speed of our TPU when we scaled to a larger systolic array size.

surya profil fotoğrafı
surya1 yıl önce

now we can move on to training! this was easily the most challenging part since we couldn’t find any resource where someone has done this before. a key insight we had while doing the math for training was that the long chain in the computational graph of the backdrop was identical to the forward pass computation graph. this meant we could calculate the long chain first, cache the gradients and then compute the individual weight and bias gradients!

surya profil fotoğrafı
surya1 yıl önce

the only problem was...we didn't have any on-chip to store the gradients (LOL) so we decided to make a unified buffer. as we developed this, we realized the unified buffer could replace our accumulators as well, making our design more elegant!

surya profil fotoğrafı
surya1 yıl önce

the VPU was the next big change we made. all those modules UNDERNEATH the systolic array (bias, activation, loss, derivatives) process column vectors element-wise. we unified them into one scalable unit with pathway bits to enable/skip operations, which is a lot more elegant than interfacing with N number of individual modules (N scales with systolic array size).

surya profil fotoğrafı
surya1 yıl önce

and finally, here's our 94-bit VLIW instruction set:

surya profil fotoğrafı
surya1 yıl önce

this was our final TPU architecture:

surya profil fotoğrafı
surya1 yıl önce

having no hardware knowledge or experience at all until just 6 months ago, this was a very ambitious project to work on. I had no idea how difficult this would be or if I could even complete it without the "prerequisites". but throughout the last 4 months, I developed a style of thinking and solving problems that I can carry forward with me. it can be encompassed by these two philosophies: - always try the dumb ideas first - DRAW EVERYTHING OUT to TRULY understand a concept you can read more about our background and thought process at and with that here's our final waveform that shows the outputs of inference and training:

Lloret profil fotoğrafı
Lloret1 yıl önce

Bro’s bout to receive a job offer from an AI lab

evan profil fotoğrafı
evan1 yıl önce

@karpathy

Xander Chin profil fotoğrafı
Xander Chin1 yıl önce

greatest thing ive ever helped create

Thread Reader App profil fotoğrafı
Thread Reader App1 yıl önce

Your thread is going viral! #TopUnroll 🙏🏼@ain3sh for 🥇unroll

krupa profil fotoğrafı
krupa1 yıl önce

@karpathy ‼️

Noah Vandal profil fotoğrafı
Noah Vandal1 yıl önce

wow this is very nice. were you at all tempted to use verilog and fpgas?

surya profil fotoğrafı
surya1 yıl önce

yes we did use verilog to build this if you look at we did everything in simulation so far, but plan to deploy on an fpga very soon

Noah Vandal profil fotoğrafı
Noah Vandal1 yıl önce

wow that is cool

saksham profil fotoğrafı
saksham1 yıl önce

@karpathy

sλrthak profil fotoğrafı
sλrthak1 yıl önce

this guy is CRACKED af

evan profil fotoğrafı
evan1 yıl önce

best teammates i could ask for 🫶

Priyav K Kaneria profil fotoğrafı
Priyav K Kaneria1 yıl önce

love it when randomly some cracked team shows up with something cooked up hard

Peaboff profil fotoğrafı
Peaboff1 yıl önce

Nice, I did a project like this too. Implemented it for the DE1-SoC + TensorFlow compatible quantization. The hardest part is wrapping your head around the quantization and wrangling all the hardware bits. I remember I was having so much difficulty getting the DMA reading to work

mikael haji profil fotoğrafı
mikael haji1 yıl önce

🔥

richa 👩‍💻 profil fotoğrafı
richa 👩‍💻1 yıl önce

@devpatelio KILLED IT

yash karthik profil fotoğrafı
yash karthik1 yıl önce

Cool!

Satvik Garimella profil fotoğrafı
Satvik Garimella1 yıl önce

This was a great read!

Little Architect profil fotoğrafı
Little Architect1 yıl önce

very nice work. you should try parametrizing your buffer sizes, pe sizes, and of course the systolic array num elements and do some sweeps on those to get some insights on the scaling you can get on lat and throughput, the balance needed for bw/compute and hw costs.

saksham profil fotoğrafı
saksham1 yıl önce

inspirational bro

braeden hall profil fotoğrafı
braeden hall1 yıl önce

i dont understand it, but its cool asf

dhruvr_43 profil fotoğrafı
dhruvr_431 yıl önce

Yoooo bro went viral. Congrats tho inspirational fr

Andy profil fotoğrafı
Andy1 yıl önce

this is crazy!!!

Satvik Garimella profil fotoğrafı
Satvik Garimella1 yıl önce

Very informative

omkaar profil fotoğrafı
omkaar1 yıl önce

i think @Si_Boehm's diagrams were in the spirit of this article

ταῦ profil fotoğrafı
ταῦ1 yıl önce

Dope! 🔥

chi profil fotoğrafı
chi1 yıl önce

incredible

Milind Kumar profil fotoğrafı
Milind Kumar1 yıl önce

Jeez big man, let’s gooo! 💪

Satvik Garimella profil fotoğrafı
Satvik Garimella1 yıl önce

Wowwww. Greattttt

tornikeo profil fotoğrafı
tornikeo1 yıl önce

This is literally you guys 😅 this is some actually impressive stuff. Also the blog reads very well! Nice work!

Lingstr profil fotoğrafı
Lingstr1 yıl önce

Great post, I see lot of research is done behind this. Thanks for sharing this

cumalot daddy profil fotoğrafı
cumalot daddy10 ay önce

Thanks for sharing your work. I have been looking for something like this.

Aayush profil fotoğrafı
Aayush1 yıl önce

goat 🐐

Hazel Bains profil fotoğrafı
Hazel Bains1 yıl önce

🔥super impressive growth in 6 months

Atif Saleem profil fotoğrafı
Atif Saleem1 yıl önce

Are you launching your inference cloud with competitive pricing without compromising on performance? Or is it just open-source? Need to read more about this

unathi 🇿🇦 profil fotoğrafı
unathi 🇿🇦1 yıl önce

fire!

grasgor profil fotoğrafı
grasgor10 ay önce

Insane

Benzer Videolar