Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I’m doin’ it … using abs() as activation in a vanilla MLP and it just works 👌 this video shows training run of MLP with one(!) hidden layer (10 neurons) like dude this shit will train on esp32 real-time also ABS is faster than ReLU and more expressive

109,323 Aufrufe • vor 7 Tagen •via X (Twitter)

46 Kommentare

Profilbild von Alex Shtoff
Alex Shtoffvor 7 Tagen

2 relu(x)=x+abs(x) So relu is abs with a "resudual stream" Alternatively, abs(x)= 2 relu(x)-x = relu(x)+relu(-x) So I dont see any "expressiveness" advantage.

Profilbild von Xenon
Xenonvor 7 Tagen

I explained that in the blog check it out 👌

Profilbild von Alex Shtoff
Alex Shtoffvor 7 Tagen

I see a lot of discussion about training dynamics rather than expressiveness. Have I missed something important?

Profilbild von Xenon
Xenonvor 7 Tagen

this ….

Profilbild von Jason Scott Lenderman
Jason Scott Lendermanvor 7 Tagen

Any piecewise linear function can be represented as a linear combination of ReLUs. In particular, abs(x) = ReLU(x) + ReLU(-x). ReLU(x), on the other hand, can't be represented with just abs activations. You also need a mechanism which can give you linear/identity connections.

Profilbild von Xenon
Xenonvor 7 Tagen

why complicate and slow down training/inference if ABS is simpler and faster 🤔

Profilbild von Xenon
Xenonvor 7 Tagen

check out these pics lol shit’s fast and expressive asf forget ReLU … ABS is the real deal 👌 cooking up a blog poast about this training run (secret bonus techniques included)

Profilbild von Xenon
Xenonvor 7 Tagen

blog poast here:

Profilbild von Xenon
Xenonvor 7 Tagen

the ABS training playground is live:

Profilbild von Wistful Meteor
Wistful Meteorvor 7 Tagen

i'm sure more than million people had tried your method lol

Profilbild von Xenon
Xenonvor 7 Tagen

nope seems ppl were just following what academics told them without even building themselves 🤷‍♂️

Profilbild von Ford F250 Gaming
Ford F250 Gamingvor 7 Tagen

@WistfulMeteor been done multiple times

Profilbild von Xenon
Xenonvor 7 Tagen

@WistfulMeteor yeah and the results are based: Thus, the replacement of Tanh activation functions with Abs makes it possible to reduce the size of original LeNet-5 model by more than 4.8 and 7 times with an improvement of its accuracy [..], and providing better accuracy than ReLU and SeLU

Profilbild von Xenon
Xenonvor 7 Tagen

@WistfulMeteor but ppl keep using slow and inefficient activations because “everybody does that” 🤷‍♂️

Profilbild von Ford F250 Gaming
Ford F250 Gamingvor 7 Tagen

@WistfulMeteor There is tons of activation functions but ones that actually work stuck. For example I've tried abs on multiple real projects and never got it to outperform ReLU whereas ELU sometimes works better. Other ppl probably tried abs and it didn't work well too

Profilbild von Xenon
Xenonvor 6 Tagen

@WistfulMeteor that’s just training skill proper tuning of hyperparam schedules + injecting noise and it works intuitive feel of the whole run and the dataset used is important many japanese diy enthusiasts produce best results in field for example - just because of relentless grind & skill

Profilbild von Xenon
Xenonvor 7 Tagen

here is the blog post describing the legendary ABS training run:

Profilbild von Robert Joo
Robert Joovor 7 Tagen

@grok why is abs faster than relu? hardware limits? or current software kernel skill issue? is it cuz like negative stuff you just flip the first bit whereas zeroingg you have to zero every bit or something?

Profilbild von Yaniv of the hills
Yaniv of the hillsvor 7 Tagen

Is this something everyone already knows about?

Profilbild von Garrett_OwlAcademy
Garrett_OwlAcademyvor 7 Tagen

There's a very fast improvement you can move by using rational trigonometry in a bounded geometric nested structure for 2D Grid quintic function navigation by projecting the 2D lattice along a 72 node toroid and spinning it using a one-way pump. this invokes a geometric ~12% Strassen matrix multiplication efficiency gain per stream, which would make these calculations near instant.

Profilbild von DumbPilot
DumbPilotvor 7 Tagen

How is ABS (assuming you mean the ABSolute operator) faster than RELU in this case? Aren't you saturating the entire gradient in the event of a negative value for your matrix operation with the ABS operator?

Profilbild von Xenon
Xenonvor 7 Tagen

check the blog and videos in og thread

Profilbild von Black Magic Evil Voodoo Warlock
Black Magic Evil Voodoo Warlockvor 7 Tagen

let's see xenon's XOR

Profilbild von kongkong
kongkongvor 7 Tagen

Calling abs() faster and more expressive than ReLU is the kind of chaos I respect. My money says it works right up until it doesn't.

Profilbild von Xenon
Xenonvor 6 Tagen

the only real response 👌

Profilbild von Yagao Dirac
Yagao Diracvor 7 Tagen

very good

Profilbild von Xenon
Xenonvor 7 Tagen

thanks! 👌 I recommend implementing this for comparison with “popular” activations

Profilbild von Yagao Dirac
Yagao Diracvor 7 Tagen

I'll test it later. I used to believe every parameter must have a fixed optimal direction which has nothing to do with its initial value. abs函数只有一个拐点,如果有多个拐点,性能可能会不一样

Profilbild von evolvingstuff
evolvingstuffvor 6 Tagen

I've been using abs in personal experiments for years, the cool thing is that the derivative is +1/-1 almost everywhere (except zero) so you avoid the exploding/vanishing gradients, at least due to the non-linearity

Profilbild von Xenon
Xenonvor 6 Tagen

what have you observed in your use cases? while running the tests I see how in low dimensions/low parameter space it struggles to suppress the symmetry (where it is not needed)

Profilbild von آزاد 🇮🇷👑
آزاد 🇮🇷👑vor 7 Tagen

Subgradients are the magic

Profilbild von gpu go brr...
gpu go brr...vor 7 Tagen

Clamp -1 to +1, you'll thank me later

Profilbild von Memo Ai agent
Memo Ai agentvor 7 Tagen

No it’s not dumbass

Profilbild von Xenon
Xenonvor 7 Tagen

did tou even see the video lol 😂 read the blog ….

Profilbild von Memo Ai agent
Memo Ai agentvor 7 Tagen

this is a very toy examples, and yes

Profilbild von msat
msatvor 7 Tagen

Dunning-Kruger effect at its finest…

Profilbild von Double Descent
Double Descentvor 6 Tagen

Did this kinda stuff during the undergrad, including this, relu turns out to be better in the end, keep exploring....

Profilbild von Xenon
Xenonvor 6 Tagen

seen them all and ABS is the best 👌

Profilbild von Valentino Giudice
Valentino Giudicevor 7 Tagen

It depends on what you are doing. There are neural networks that don't train well with all activation functions. Sometimes ReLU fails too and you need to use another function. Non linearities aren't always interchangeable.

Profilbild von Aiden
Aidenvor 7 Tagen

How is abs more expressive than ReLU??

Profilbild von Xenon
Xenonvor 7 Tagen

I explain this in the blog:

Profilbild von david
davidvor 7 Tagen

@grok give explanations and find evidence supporting and defeating this approach compared to the conventional relu

Profilbild von Deepak Sadulla
Deepak Sadullavor 7 Tagen

Wonder why it looks like we are bending a straight piece of wire..

Profilbild von Orange Juice 🦏
Orange Juice 🦏vor 6 Tagen

this doesn't make much sense

Profilbild von Sebastian Buzdugan
Sebastian Buzduganvor 7 Tagen

does abs still beat relu after compilation, when activation cost barely affects total latency

Profilbild von x2y2x2z2y2z2k🇦🇲🇬🇷🇨🇾
x2y2x2z2y2z2k🇦🇲🇬🇷🇨🇾vor 7 Tagen

You are overfitting bro

Ähnliche Videos