Video wird geladen...
Video konnte nicht geladen werden
I’m doin’ it … using abs() as activation in a vanilla MLP and it just works 👌 this video shows training run of MLP with one(!) hidden layer (10 neurons) like dude this shit will train on esp32 real-time also ABS is faster than ReLU and more expressive
109,323 Aufrufe • vor 7 Tagen •via X (Twitter)
46 Kommentare

2 relu(x)=x+abs(x) So relu is abs with a "resudual stream" Alternatively, abs(x)= 2 relu(x)-x = relu(x)+relu(-x) So I dont see any "expressiveness" advantage.

I explained that in the blog check it out 👌

I see a lot of discussion about training dynamics rather than expressiveness. Have I missed something important?

this ….

Any piecewise linear function can be represented as a linear combination of ReLUs. In particular, abs(x) = ReLU(x) + ReLU(-x). ReLU(x), on the other hand, can't be represented with just abs activations. You also need a mechanism which can give you linear/identity connections.

why complicate and slow down training/inference if ABS is simpler and faster 🤔

check out these pics lol shit’s fast and expressive asf forget ReLU … ABS is the real deal 👌 cooking up a blog poast about this training run (secret bonus techniques included)

blog poast here:

the ABS training playground is live:

i'm sure more than million people had tried your method lol

nope seems ppl were just following what academics told them without even building themselves 🤷♂️

@WistfulMeteor been done multiple times

@WistfulMeteor yeah and the results are based: Thus, the replacement of Tanh activation functions with Abs makes it possible to reduce the size of original LeNet-5 model by more than 4.8 and 7 times with an improvement of its accuracy [..], and providing better accuracy than ReLU and SeLU

@WistfulMeteor but ppl keep using slow and inefficient activations because “everybody does that” 🤷♂️

@WistfulMeteor There is tons of activation functions but ones that actually work stuck. For example I've tried abs on multiple real projects and never got it to outperform ReLU whereas ELU sometimes works better. Other ppl probably tried abs and it didn't work well too

@WistfulMeteor that’s just training skill proper tuning of hyperparam schedules + injecting noise and it works intuitive feel of the whole run and the dataset used is important many japanese diy enthusiasts produce best results in field for example - just because of relentless grind & skill

here is the blog post describing the legendary ABS training run:

@grok why is abs faster than relu? hardware limits? or current software kernel skill issue? is it cuz like negative stuff you just flip the first bit whereas zeroingg you have to zero every bit or something?

Is this something everyone already knows about?

There's a very fast improvement you can move by using rational trigonometry in a bounded geometric nested structure for 2D Grid quintic function navigation by projecting the 2D lattice along a 72 node toroid and spinning it using a one-way pump. this invokes a geometric ~12% Strassen matrix multiplication efficiency gain per stream, which would make these calculations near instant.

How is ABS (assuming you mean the ABSolute operator) faster than RELU in this case? Aren't you saturating the entire gradient in the event of a negative value for your matrix operation with the ABS operator?

check the blog and videos in og thread

let's see xenon's XOR

Calling abs() faster and more expressive than ReLU is the kind of chaos I respect. My money says it works right up until it doesn't.

the only real response 👌

very good

thanks! 👌 I recommend implementing this for comparison with “popular” activations

I'll test it later. I used to believe every parameter must have a fixed optimal direction which has nothing to do with its initial value. abs函数只有一个拐点,如果有多个拐点,性能可能会不一样

I've been using abs in personal experiments for years, the cool thing is that the derivative is +1/-1 almost everywhere (except zero) so you avoid the exploding/vanishing gradients, at least due to the non-linearity

what have you observed in your use cases? while running the tests I see how in low dimensions/low parameter space it struggles to suppress the symmetry (where it is not needed)

Subgradients are the magic

Clamp -1 to +1, you'll thank me later

No it’s not dumbass

did tou even see the video lol 😂 read the blog ….

this is a very toy examples, and yes

Dunning-Kruger effect at its finest…

Did this kinda stuff during the undergrad, including this, relu turns out to be better in the end, keep exploring....

seen them all and ABS is the best 👌

It depends on what you are doing. There are neural networks that don't train well with all activation functions. Sometimes ReLU fails too and you need to use another function. Non linearities aren't always interchangeable.

How is abs more expressive than ReLU??

I explain this in the blog:

@grok give explanations and find evidence supporting and defeating this approach compared to the conventional relu

Wonder why it looks like we are bending a straight piece of wire..

this doesn't make much sense

does abs still beat relu after compilation, when activation cost barely affects total latency

You are overfitting bro
