正在加载视频...

视频加载失败

I’m thrilled to present the KAN Convolutional Layer, a promising neural architecture for image processing. Last week KANs came out, as an alternative to the MLP. We extended this idea to Convolutional Layers, creating the KAN Convolution. Join me in this thread to more about it🧵

201,550 次观看 • 2 年前 •via X (Twitter)

46 条评论

Xingchao Liu 的头像
Xingchao Liu2 年前

how about KANvolution?😂

Alex Bodner 的头像
Alex Bodner2 年前

HAHAHAHAHAH

Alex Bodner 的头像
Alex Bodner2 年前

All the code is available in the following repository: Co Authored with @JackSpolski @SantiPourteau @antotepsich Lets dive in what we've created 👇

Alex Bodner 的头像
Alex Bodner2 年前

What is a KAN? A KAN connection applies a function defined as a learned B Spline, plus a Residual activation function b(x), and all this times a learnable parameter w. The whole node value is the sum of the different ’s that enter the node.

Alex Bodner 的头像
Alex Bodner2 年前

What is a KAN Convolution? KAN Convolutions are very similar to convolutions, we apply a Learnable Non Linear function in each edge, and then add them up. The kernel of the KAN Convolution is equivalent to a KAN Linear Layer of 4 inputs and 1 output neuron.

Alex Bodner 的头像
Alex Bodner2 年前

Results In our first evaluations with MNIST Dataset, we got very tight results between KAN vs Classic convolutions with a MLP at the end, with the last one getting 0.06 more accuracy. But the KAN Convs with a KAN got only 0.04 less accuracy with almost half the parameters!

Alex Bodner 的头像
Alex Bodner2 年前

Work in progress In the coming days and weeks we will be thoroughly tuning the hyperparameters of the models we use in the comparison. In addition, we will evaluate on different and more complex datasets to extract definitive conclussions.

Alex Bodner 的头像
Alex Bodner2 年前

And that's all for now, if you liked it give the repo a star⭐, help spreading the work with a RT and follow me for more AI content!

Diego Porres 的头像
Diego Porres2 年前

@Ethan_smith_20 a bit of what was discussed last time, but I think the jump needs to be better directed, i.e., I don't think splines are best suited here (but I am of course willing to shut up and let them cook)

Valeriy M., PhD, MBA, CQF 的头像
Valeriy M., PhD, MBA, CQF2 年前

Amazing!

Patricio Clark 的头像
Patricio Clark2 年前

Muy buen trabajo!

Alex Bodner 的头像
Alex Bodner2 年前

Muchas gracias Pato!

Angus 的头像
Angus2 年前

Felicitaciones Alex y co!

Alex Bodner 的头像
Alex Bodner2 年前

Muchas gracias!

Lichu Acuña 的头像
Lichu Acuña1 年前

based af

lodestone-rock 的头像
lodestone-rock2 年前

training speed? is it parallelizable ?

Alex Bodner 的头像
Alex Bodner2 年前

@LodestoneE621 We will update in the following days with training speed evaluations, but if you checkout the notebook in the repo you should get a clue. The current problem with KANs is that they depend on non linear functions that you cant parallelize in GPU ATM

Dan Shlamovitz 的头像
Dan Shlamovitz2 年前

Amazing!!

rakesh 的头像
rakesh2 年前

@TechRadar9

Satyabrat Singh 的头像
Satyabrat Singh2 年前

Nice :)

Rùa 的头像
Rùa2 年前

It sounds like a prophecy but I'm looking forward to KAN proving it will replace MLP in all Vision encoder architectures and Language encoders (or it proves itself to be just an illusion from expectations :)))) I wonder if you want to experiment on transformer architectures?

HDP 的头像
HDP2 年前

@DeeperThrill Code in next tweet

Sugar 的头像
Sugar2 年前

@readwise save

JPL 的头像
JPL2 年前

Nice job! What about the training time of the KAN convolutions vs the classic ones?

Mike McCarthy 的头像
Mike McCarthy2 年前

@threadreaderapp unroll

Ken aka Frosty 的头像
Ken aka Frosty2 年前

!!! this is SO COOL to already see KANs getting put to work Cant wait to see what else comes 🍿

Marcel B. 的头像
Marcel B.1 年前

This KAN Convolution idea is fascinating! Given KANs' attention to function complexity, I'm curious about how you're quantifying the effective receptive field or spatial feature hierarchy compared to traditional CNNs. 🤔

Kevin 的头像
Kevin2 年前

Inference speed?

Mohammed Benaissa 的头像
Mohammed Benaissa2 年前

Is it possible to adapt it to the following ccn model .

Jack Vial 的头像
Jack Vial2 年前

Very interesting work! Does some version of the convolution theorem still apply to KAN convolutions? As in, is there a dual of KAN convolution in the Fourier domain?

Alex Bodner 的头像
Alex Bodner2 年前

Not sure about it, my gut tells me that the theorem as is doesnt stand, but is work to do!

Jack Vial 的头像
Jack Vial2 年前

Yeah I'm thinking the non-linearity means there is no Fourier transform dual but maybe there is a Volterra/Wiener series transform

Thread Reader App 的头像
Thread Reader App2 年前

Your thread is gaining traction! #TopUnroll 🙏🏼@QTMDots for 🥇unroll

Pistis 的头像
Pistis2 年前

Have you thought about the possibility to do equivariant convolutions as well? Efficiency is especially important in that field.

Alex Bodner 的头像
Alex Bodner2 年前

Sorry, but, what do you mean when you say Equivariant Convolutions?

Pistis 的头像
Pistis2 年前

I mean as in equivariant and geometric deep learning. Standard convolutions are equivariant to translations, but in some domains (eg medical image analysis) the objective is also for example rotation invariant.

Pistis 的头像
Pistis2 年前

Here’s an example of such work:

Émile Amajar 的头像
Émile Amajar2 年前

What's the point ? Novel architectures should be easier to compute on current machines This is just another alternate formulation of a parameterized network, it can't bear any advantage in and of itself.

Benedict 的头像
Benedict2 年前

That is so freakin awesome. Thanks for working on this and in the open

notro 的头像
notro2 年前

Really cool, think this could scale reasonably to more layers?

Anima FX 的头像
Anima FX2 年前

honestly i’m struggling to see how KANs are in any way superior to classical MLPs. Just the non-parallelized architecture should indicate that it isn’t likely to overtake MLPs for ML anytime soon.

Swaroop 的头像
Swaroop2 年前

Great stuff! I wonder if the spline enables better performance in generative tasks, due to the introduction of non linearity in the kernel.

😮‍💨😆🤳 的头像
😮‍💨😆🤳2 年前

Why do you use a SiLU activation on the input x before passing it through the residual connection? Referring to

AtanOldman 的头像
AtanOldman2 年前

Awsome! What about explainability? Is there any added value on this regard?

Alex Bodner 的头像
Alex Bodner2 年前

Thanks, one could visualize the function that each pixel of the kernel learns, but IMO it is very hard to conclude something out of that. But we are working on it. In one of my figures you can see the outputs of conv layers and it Seems that they learn the shape.

AtanOldman 的头像
AtanOldman2 年前

Lovely! All the best

相关视频

Fukushima's video (1986) shows a CNN that recognises handwritten digits [3], three years before LeCun's video (1989). CNN timeline taken from [5]: ★ 1969: Kunihiko Fukushima published rectified linear units or ReLUs [1] which are now extensively used in CNNs. ★ 1979: Fukushima published the basic CNN architecture with convolution layers and downsampling layers [2]. He called it neocognitron. It was trained by unsupervised learning rules. Compute was 100 times more expensive than in 1989, and a billion times more expensive than today. ★ 1986: Fukushima's video on recognising hand-written digits [3]. ★ 1988: Wei Zhang et al had the first "modern" 2-dimensional CNN trained by backpropagation, and also applied it to character recognition [4]. Compute was about 10 million times more expensive than today. ★ 1989-: later work by others [5]. REFERENCES (more in [5]) [1] K. Fukushima (1969). Visual feature extraction by a multilayered network of analog threshold elements. IEEE Transactions on Systems Science and Cybernetics. 5 (4): 322-333. This work introduced rectified linear units or ReLUs, now widely used in CNNs and other neural nets. [2] K. Fukushima (1979). Neural network model for a mechanism of pattern recognition unaffected by shift in position—Neocognitron. Trans. IECE, vol. J62-A, no. 10, pp. 658-665, 1979. The first deep convolutional neural network architecture, with alternating convolutional layers and downsampling layers. In Japanese. English version: 1980. [3] Movie produced by K. Fukushima, S. Miyake and T. Ito (NHK Science and Technical Research Laboratories), in 1986. YouTube: [4] W. Zhang, J. Tanida, K. Itoh, Y. Ichioka. Shift-invariant pattern recognition neural network and its optical architecture. Proc. Annual Conference of the Japan Society of Applied Physics, 1988. First "modern" backpropagation-trained 2-dimensional CNN, applied to character recognition. [5] J. Schmidhuber (AI Blog, 2025). Who invented convolutional neural networks?

Jürgen Schmidhuber

742,172 次观看 • 9 个月前