Загрузка видео...

Не удалось загрузить видео

На главную

I’m thrilled to present the KAN Convolutional Layer, a promising neural architecture for image processing. Last week KANs came out, as an alternative to the MLP. We extended this idea to Convolutional Layers, creating the KAN Convolution. Join me in this thread to more about it🧵

201,550 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 46

Фото профиля Xingchao Liu
Xingchao Liu2 лет назад

how about KANvolution?😂

Фото профиля Alex Bodner
Alex Bodner2 лет назад

HAHAHAHAHAH

Фото профиля Alex Bodner
Alex Bodner2 лет назад

All the code is available in the following repository: Co Authored with @JackSpolski @SantiPourteau @antotepsich Lets dive in what we've created 👇

Фото профиля Alex Bodner
Alex Bodner2 лет назад

What is a KAN? A KAN connection applies a function defined as a learned B Spline, plus a Residual activation function b(x), and all this times a learnable parameter w. The whole node value is the sum of the different ’s that enter the node.

Фото профиля Alex Bodner
Alex Bodner2 лет назад

What is a KAN Convolution? KAN Convolutions are very similar to convolutions, we apply a Learnable Non Linear function in each edge, and then add them up. The kernel of the KAN Convolution is equivalent to a KAN Linear Layer of 4 inputs and 1 output neuron.

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Results In our first evaluations with MNIST Dataset, we got very tight results between KAN vs Classic convolutions with a MLP at the end, with the last one getting 0.06 more accuracy. But the KAN Convs with a KAN got only 0.04 less accuracy with almost half the parameters!

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Work in progress In the coming days and weeks we will be thoroughly tuning the hyperparameters of the models we use in the comparison. In addition, we will evaluate on different and more complex datasets to extract definitive conclussions.

Фото профиля Alex Bodner
Alex Bodner2 лет назад

And that's all for now, if you liked it give the repo a star⭐, help spreading the work with a RT and follow me for more AI content!

Фото профиля Diego Porres
Diego Porres2 лет назад

@Ethan_smith_20 a bit of what was discussed last time, but I think the jump needs to be better directed, i.e., I don't think splines are best suited here (but I am of course willing to shut up and let them cook)

Фото профиля Valeriy M., PhD, MBA, CQF
Valeriy M., PhD, MBA, CQF2 лет назад

Amazing!

Фото профиля Patricio Clark
Patricio Clark2 лет назад

Muy buen trabajo!

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Muchas gracias Pato!

Фото профиля Angus
Angus2 лет назад

Felicitaciones Alex y co!

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Muchas gracias!

Фото профиля Lichu Acuña
Lichu Acuña1 год назад

based af

Фото профиля lodestone-rock
lodestone-rock2 лет назад

training speed? is it parallelizable ?

Фото профиля Alex Bodner
Alex Bodner2 лет назад

@LodestoneE621 We will update in the following days with training speed evaluations, but if you checkout the notebook in the repo you should get a clue. The current problem with KANs is that they depend on non linear functions that you cant parallelize in GPU ATM

Фото профиля Dan Shlamovitz
Dan Shlamovitz2 лет назад

Amazing!!

Фото профиля rakesh
rakesh2 лет назад

@TechRadar9

Фото профиля Satyabrat Singh
Satyabrat Singh2 лет назад

Nice :)

Фото профиля Rùa
Rùa2 лет назад

It sounds like a prophecy but I'm looking forward to KAN proving it will replace MLP in all Vision encoder architectures and Language encoders (or it proves itself to be just an illusion from expectations :)))) I wonder if you want to experiment on transformer architectures?

Фото профиля HDP
HDP2 лет назад

@DeeperThrill Code in next tweet

Фото профиля Sugar
Sugar2 лет назад

@readwise save

Фото профиля JPL
JPL2 лет назад

Nice job! What about the training time of the KAN convolutions vs the classic ones?

Фото профиля Mike McCarthy
Mike McCarthy2 лет назад

@threadreaderapp unroll

Фото профиля Ken aka Frosty
Ken aka Frosty2 лет назад

!!! this is SO COOL to already see KANs getting put to work Cant wait to see what else comes 🍿

Фото профиля Marcel B.
Marcel B.1 год назад

This KAN Convolution idea is fascinating! Given KANs' attention to function complexity, I'm curious about how you're quantifying the effective receptive field or spatial feature hierarchy compared to traditional CNNs. 🤔

Фото профиля Kevin
Kevin2 лет назад

Inference speed?

Фото профиля Mohammed Benaissa
Mohammed Benaissa2 лет назад

Is it possible to adapt it to the following ccn model .

Фото профиля Jack Vial
Jack Vial2 лет назад

Very interesting work! Does some version of the convolution theorem still apply to KAN convolutions? As in, is there a dual of KAN convolution in the Fourier domain?

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Not sure about it, my gut tells me that the theorem as is doesnt stand, but is work to do!

Фото профиля Jack Vial
Jack Vial2 лет назад

Yeah I'm thinking the non-linearity means there is no Fourier transform dual but maybe there is a Volterra/Wiener series transform

Фото профиля Thread Reader App
Thread Reader App2 лет назад

Your thread is gaining traction! #TopUnroll 🙏🏼@QTMDots for 🥇unroll

Фото профиля Pistis
Pistis2 лет назад

Have you thought about the possibility to do equivariant convolutions as well? Efficiency is especially important in that field.

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Sorry, but, what do you mean when you say Equivariant Convolutions?

Фото профиля Pistis
Pistis2 лет назад

I mean as in equivariant and geometric deep learning. Standard convolutions are equivariant to translations, but in some domains (eg medical image analysis) the objective is also for example rotation invariant.

Фото профиля Pistis
Pistis2 лет назад

Here’s an example of such work:

Фото профиля Émile Amajar
Émile Amajar2 лет назад

What's the point ? Novel architectures should be easier to compute on current machines This is just another alternate formulation of a parameterized network, it can't bear any advantage in and of itself.

Фото профиля Benedict
Benedict2 лет назад

That is so freakin awesome. Thanks for working on this and in the open

Фото профиля notro
notro2 лет назад

Really cool, think this could scale reasonably to more layers?

Фото профиля Anima FX
Anima FX2 лет назад

honestly i’m struggling to see how KANs are in any way superior to classical MLPs. Just the non-parallelized architecture should indicate that it isn’t likely to overtake MLPs for ML anytime soon.

Фото профиля Swaroop
Swaroop2 лет назад

Great stuff! I wonder if the spline enables better performance in generative tasks, due to the introduction of non linearity in the kernel.

Фото профиля 😮‍💨😆🤳
😮‍💨😆🤳2 лет назад

Why do you use a SiLU activation on the input x before passing it through the residual connection? Referring to

Фото профиля AtanOldman
AtanOldman2 лет назад

Awsome! What about explainability? Is there any added value on this regard?

Фото профиля Alex Bodner
Alex Bodner2 лет назад

Thanks, one could visualize the function that each pixel of the kernel learns, but IMO it is very hard to conclude something out of that. But we are working on it. In one of my figures you can see the outputs of conv layers and it Seems that they learn the shape.

Фото профиля AtanOldman
AtanOldman2 лет назад

Lovely! All the best

Похожие видео

Fukushima's video (1986) shows a CNN that recognises handwritten digits [3], three years before LeCun's video (1989). CNN timeline taken from [5]: ★ 1969: Kunihiko Fukushima published rectified linear units or ReLUs [1] which are now extensively used in CNNs. ★ 1979: Fukushima published the basic CNN architecture with convolution layers and downsampling layers [2]. He called it neocognitron. It was trained by unsupervised learning rules. Compute was 100 times more expensive than in 1989, and a billion times more expensive than today. ★ 1986: Fukushima's video on recognising hand-written digits [3]. ★ 1988: Wei Zhang et al had the first "modern" 2-dimensional CNN trained by backpropagation, and also applied it to character recognition [4]. Compute was about 10 million times more expensive than today. ★ 1989-: later work by others [5]. REFERENCES (more in [5]) [1] K. Fukushima (1969). Visual feature extraction by a multilayered network of analog threshold elements. IEEE Transactions on Systems Science and Cybernetics. 5 (4): 322-333. This work introduced rectified linear units or ReLUs, now widely used in CNNs and other neural nets. [2] K. Fukushima (1979). Neural network model for a mechanism of pattern recognition unaffected by shift in position—Neocognitron. Trans. IECE, vol. J62-A, no. 10, pp. 658-665, 1979. The first deep convolutional neural network architecture, with alternating convolutional layers and downsampling layers. In Japanese. English version: 1980. [3] Movie produced by K. Fukushima, S. Miyake and T. Ito (NHK Science and Technical Research Laboratories), in 1986. YouTube: [4] W. Zhang, J. Tanida, K. Itoh, Y. Ichioka. Shift-invariant pattern recognition neural network and its optical architecture. Proc. Annual Conference of the Japan Society of Applied Physics, 1988. First "modern" backpropagation-trained 2-dimensional CNN, applied to character recognition. [5] J. Schmidhuber (AI Blog, 2025). Who invented convolutional neural networks?

Jürgen Schmidhuber

742,172 просмотров • 9 месяцев назад