Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Fukushima's video (1986) shows a CNN that recognises handwritten digits [3], three years before LeCun's video (1989). CNN timeline taken from [5]: ★ 1969: Kunihiko Fukushima published rectified linear units or ReLUs [1] which are now extensively used in CNNs. ★ 1979: Fukushima published the basic CNN architecture with...

742,172 görüntüleme • 9 ay önce •via X (Twitter)

35 Yorum

Valeriy M., PhD, MBA, CQF profil fotoğrafı
Valeriy M., PhD, MBA, CQF9 ay önce

It’s ironic that Schmidhuber credits Fukushima with inventing CNNs while overlooking several other deep-learning researchers—such as Galushkin—whose work dates back to the 1960s.

Jürgen Schmidhuber profil fotoğrafı
Jürgen Schmidhuber9 ay önce

Galushkin was the pupil of Tsypkin whom I cite in the Annotated History of Modern AI and Deep Learning as follows: "See also Iakov Zalmanovich Tsypkin's even earlier work on gradient descent-based on-line learning for non-linear systems [GDa-b]." [GDa] Y. Z. Tsypkin (1966). Adaptation, training and self-organization automatic control systems, Avtomatika I Telemekhanika, 27, 23-61. On gradient descent-based on-line learning for non-linear systems. [GDb] Y. Z. Tsypkin (1971). Adaptation and Learning in Automatic Systems, Academic Press, 1971. On gradient descent-based on-line learning for non-linear systems.

Amir Arsalan Soltani profil fotoğrafı
Amir Arsalan Soltani9 ay önce

@predict_addict Jürgen Schmidhuber never betrays the truth and history 🔥 Even if it seems that's the case, he will correct himself; my observation of him is he does his best to attribute the right part of history to the right people & doesn't have religious beliefs about "popular" attributions

Valeriy M., PhD, MBA, CQF profil fotoğrafı
Valeriy M., PhD, MBA, CQF9 ay önce

@SchmidhuberAI Indeed, he is also great person in real life.

Jürgen Schmidhuber profil fotoğrafı
Jürgen Schmidhuber29 gün önce

2026 update: 1st backprop-CNN for vision (Wei Zhang et al, 1988)

Axel Pond profil fotoğrafı
Axel Pond9 ay önce

Schmiduber always brings receipts

Xenon profil fotoğrafı
Xenon9 ay önce

Yapp Le Cum saw the VHS and didn’t think anyone would get this footage on internet years later…. L move tbh

Rudi Ranck profil fotoğrafı
Rudi Ranck9 ay önce

In your view, does this call Yan LeCun's Turing Award into question?

K3ith.AI profil fotoğrafı
K3ith.AI9 ay önce

I’ve done a few talks on the history of AI and always came to the same conclusions you did you again😋 ever since you came out with your history book I’ve been trying to convert peeps…albeit not here enough…I mean isn’t backprop from Leibniz?😉 We’re all part of the tapestry of knowledge here…some folks are just bigger patches and stronger twine than others🥳😎🦾

Michael Tomlinson profil fotoğrafı
Michael Tomlinson9 ay önce

Thank you! I see the same video published so many times. I built a simple network in 1984 as a high school student, that recognised the hand-written alphabet, on a 386… from a one-page paragraph describing neural nets in new scientist (basically, nodes, weights, input and outputs. I used 16.16 fp, no bias nodes (they weren’t mentioned) and a monochrome cmos chip camera to capture the images. It took 38 training cycles to converge to 100% accuracy (I only had one set of images) but it was remarkably efficient and fast for the time (and my lack of knowledge 🤣). It wasn’t a CNN, it was fully connected, and trained through a crude approximation of back-propagation.

Insert Something Funny profil fotoğrafı
Insert Something Funny9 ay önce

I have repeatedly said why is there no other major contribution from @ylecun all these years 35 years. This is the answer. It was not his original work.

AndyXAndersen profil fotoğrafı
AndyXAndersen9 ay önce

LeCun made CNN work reliably and practically. It is, if you wish, what Steve Jobs did with Apple. This is no small feat. Discovery and product creation is always a process. Takes lots of people.

Frankenbyte profil fotoğrafı
Frankenbyte9 ay önce

This needs to be fully recognized. It's the same pattern over and over again. Enough already.

Drohi - ReSSRection 🌄 profil fotoğrafı
Drohi - ReSSRection 🌄9 ay önce

Its not fair that LeCun gets all the credit. Its also not fair that the Fukushima reactor named after him blew up in 2012 😔

Hamid Maei profil fotoğrafı
Hamid Maei9 ay önce

Just my humble response to this: Thanks for bringing up these references; personally, I wasn’t even aware of them. I 'grew up' with LeCun’s networks when I took my first deep learning course in 2006 with Geoff Hinton at UofT. I agree that we need to foster a much healthier environment for research. Ultimately, we are all standing on the shoulders of giants. There is often a distinction between who did it first and who communicated it in a way that the industry adopted at the right time! For example, Gerry Tesauro ( one of the best scientists I know) developed the first Deep RL for Backgammon (TD-Gammon) in the early 90s and achieved world-class performance. So AlphaGo, created by David Silver, who is also a friend Gerry, was not strictly the first. But Gerry definitely didn’t mind, AFIK, the credit going to David, as David scaled it and applied it to Computer Go. You have good points, and it’s important to bring these things up, but it would be nice if it came across more as adding historical context. It might not always seem fair to the pioneers, but that seems to be how the history of technology and science is written.

dmsimon profil fotoğrafı
dmsimon9 ay önce

LeCun’s key contribution was showing that you could train the whole thing end to end with backpropagation on 1980s–1990s hardware and get state of the art practical performance. That made CNNs explode in the West and eventually led to the deep learning revolution.

Алтангүх Батбаяр profil fotoğrafı
Алтангүх Батбаяр9 ay önce

How did this guy not receive the Turing Award instead of those grifters :(

Mohamed Sami Zeghdar profil fotoğrafı
Mohamed Sami Zeghdar9 ay önce

Thank you for demystifying the history.

David N. Schwartz profil fotoğrafı
David N. Schwartz9 ay önce

Classic. First, I saw Fukushima and was waiting to see the earthquake. Then I saw CNN and thought there was something going on about the news platform I usually watch. Then I read more carefully. I am an idiot!

Rudi Ranck profil fotoğrafı
Rudi Ranck9 ay önce

I was always told wrong

Flo🥝 profil fotoğrafı
Flo🥝9 ay önce

No wonder Lecun always sounded like a joke.

turboderp profil fotoğrafı
turboderp9 ay önce

Neural networks don't beep like they used to. :|

Jan-Erik Vinje ⏸️AI profil fotoğrafı
Jan-Erik Vinje ⏸️AI9 ay önce

So @ylecun . Did you steal the credit for something that was achieved many years prior?

Krunal Kshirsagar profil fotoğrafı
Krunal Kshirsagar9 ay önce

Shunichi Amari for the win!!

haX profil fotoğrafı
haX9 ay önce

why are we hearing about this only now?

ALGOTRIX profil fotoğrafı
ALGOTRIX9 ay önce

Why did not Japan become an AI superpower? If they had Fukushima?

Mark Medovich profil fotoğrafı
Mark Medovich9 ay önce

"The simplest technique is a direct-convolution realization using the relation NI-I Nz-1 y(n1, n2) = c c h(m1, mz)x(n1 - m1, n2 - mz) (5) m1=0 mp=o where x(n1, nz) is the input to the filter and y(nl, fl2) is the filter output." -- 1972

WuBu ⪋ WaefreBeorn 🇺🇸 👑 profil fotoğrafı
WuBu ⪋ WaefreBeorn 🇺🇸 👑9 ay önce

my dad did text only systems recognition is the defining line here the use of a visual CNN? cause retrieval algo is a RNN too? my dad made a fortran search algo that returned hallucinated text as well off of deweyDD tag information NO VLM THO so idk if that counts

Ghost 🇨🇦 profil fotoğrafı
Ghost 🇨🇦9 ay önce

Fukushima created the architecture Zhang created the first modern supervised CNN LeCun standardized it, pushed it into practice and owned the narrative All three matter. Only one became famous.

Per Nystedt profil fotoğrafı
Per Nystedt9 ay önce

Please don't stare in rear mirror for too long, we need the brightest people to look ahead!

Bijan Tavassoli profil fotoğrafı
Bijan Tavassoli9 ay önce

Ich hoffe das überzeugt @kimmonismus zur richtigen Seite zu wechseln.

Mr Rio profil fotoğrafı
Mr Rio9 ay önce

LeCunt

Nimkef profil fotoğrafı
Nimkef9 ay önce

Are you saying these are cases of multiple discovery or plagiarism?

Lord Hoot profil fotoğrafı
Lord Hoot9 ay önce

he called it neocognitron

Jacques Kaderian profil fotoğrafı
Jacques Kaderian9 ay önce

@HacKanCuBa.

Benzer Videolar

Dropout by hand ✍️ ~ 10 steps walkthrough below Dropout is the simplest trick in deep learning that actually works: during training you randomly switch neurons off, so the network cannot lean on any one of them. It is two lines of code and almost nobody has worked through what those lines do to the numbers. So I drew and calculated one entirely by hand. Goal: train one pass through a small network with two dropout layers, then run inference with dropout switched off. The network: Linear(2,4), ReLU, Dropout(0.5), Linear(4,3), ReLU, Dropout(0.33), Linear(3,2). = 1. Given = A training set of two examples, X1 and X2, and the weight matrices for all three linear layers. = 2. Draw the first random numbers = Let us draw 4 random numbers, one per neuron in the first hidden layer. Above 0.5 we keep (◯), below we drop (╳). Here that gives [◯, ╳, ◯, ╳]. = 3. Build the first dropout matrix = We turn that pattern into a diagonal matrix. The scaling factor is 1/(1-p) = 2, so a kept neuron gets 2 and a dropped one gets 0. Multiplying by it does both jobs at once: it deletes the 2nd and 4th neurons and doubles the two that survive. = 4. Draw the second random numbers = Let us do it again for the 3 neurons in the next layer, this time against p = 0.33. The result is [◯, ◯, ╳]. = 5. Build the second dropout matrix = We set the diagonal to 1.5 where kept and 0 where dropped. Only the 3rd neuron goes. = 6. Feed forward = Let us run the whole thing top to bottom: one matrix multiplication per layer, ReLU setting the negatives to zero, and the two dropout matrices doing their work in between. The outputs Y come out at the bottom. = 7. MSE loss gradients = We compare Y against the targets Y', subtract, and multiply each element by 2. That is the whole gradient of the mean squared error. = 8. Update the weights = Let us push those gradients back through the network and update the weights (marked in light red). = 9. Deactivate dropout = Training is over, so we set both dropout matrices to the identity. Every neuron is back, and nothing is scaled. = 10. Feed forward again = One more pass, this time on unseen data, to make the prediction. You have just trained and run a network with dropout by hand. ✍️ The outputs: Training outputs Y = [-6, 9; 13, 4] Loss gradients = [-4, 4; 6, -2] Inference outputs = [13, 13; 4, 3] 💾 Save this post! #AIbyHand #Dropout #DeepLearning #NeuralNetworks

Tom Yeh

14,442 görüntüleme • 1 ay önce

Backpropagation by hand ✍️ ~ 11 steps walkthrough below Backpropagation is the algorithm that actually trains a neural network, and it is where most people stop following along. It is not calculus you cannot do. It is matrix multiplication, working backward, one layer at a time. So I drew and calculated one entirely by hand. Goal: push the loss gradient back through a 3-layer network and land on a new value for every weight and bias. = 1. Given = A 3-layer perceptron, an input X, predictions Ypred = [0.5, 0.5, 0], and the truth Ytarget = [0, 1, 0]. = 2. Backprop gradient cells = Let us draw empty cells for every gradient we are about to compute. The shape of the answer comes first. = 3. Layer 3 softmax = We get dL/dz3 straight from Ypred minus Ytarget = [0.5, -0.5, 0]. No chain rule needed, and that shortcut is the whole reason softmax and cross-entropy are paired. = 4. Layer 3 weights and biases = Let us multiply dL/dz3 by [a2 | 1]. One multiplication gives the gradient for W3 and b3 together. = 5. Layer 2 activations = We multiply dL/dz3 by W3 to get dL/da2. The gradient moves back across a layer the same way the signal moved forward. = 6. Layer 2 ReLU = Let us pass it through the gate: keep the gradient where the activation was positive, zero it everywhere else. = 7. Layer 2 weights and biases = We multiply dL/dz2 by [a1 | 1]. The same figure as step 4, one layer up. = 8. Layer 1 activations = Let us multiply dL/dz2 by W2. = 9. Layer 1 ReLU = We apply the same gate again, now on a1. = 10. Layer 1 weights and biases = Let us multiply dL/dz1 by [x | 1], and every weight in the network now has a gradient. = 11. Update = We subtract, and the network has learned. In practice a learning rate scales this step. The gradients: dL/dz3 = [0.5, -0.5, 0] dL/da1 = [1, -2, 2, -1] dL/dz1 = [0, -2, 2, -1] The takeaway: matrix multiplication is all you need. Just like the forward pass, backpropagation is matrix multiplications end to end. You can do every one by hand, slowly and imperfectly, which is exactly why a GPU's ability to do them fast mattered so much to deep learning. 💾 Save this post!

Tom Yeh

961,186 görüntüleme • 1 ay önce

ResNet by hand ✍️ ~ 10 steps walkthrough below "Deep Residual Learning for Image Recognition" (Kaiming He, CVPR 2016) is among the most cited papers in all of deep learning. Why does it matter so much? It fixed the exploding and vanishing gradients that kept deep networks from being deep, and made thousands of layers possible. How simple was the fix? An identity matrix. Goal: push three input vectors through a residual block, then through a transformer encoder block, filling in every cell yourself. = 1. Given = A mini batch of three input vectors, 3D, and the weights of the layers ahead. = 2. Linear layer = Let us multiply by the weights, add the bias, and apply ReLU so negatives become 0. Three feature vectors out. This is F(X). = 3. Concatenate = Now the trick. Stack an identity matrix beside the second layer's weights, and stack the input vectors under the features. Draw the lines between rows and columns: those are the skip connections. The identity is the residual. = 4. Linear layer + identity = We multiply the two stacked matrices. The identity carries X straight through while the weights transform it, so a single multiplication computes F(X) + X. Apply ReLU and hand it to the next block. Now watch the same trick inside a transformer, first in attention. = 5. Attention = Let us take three input vectors in 2D, compute the attention matrix, and multiply to get attention weighted vectors. = 6. Concatenate = We stack two identities this time, two residuals, which is how you get 1 + 1, and stack the input vectors with the attention weighted ones. = 7. Add = Multiply the stacked matrices. The identity adds attention to its own input, across the columns, which is how positions get combined. And again in the feed forward layer. = 8. First layer = Let us multiply by the feed forward weights and bias, then ReLU. Three feature vectors. = 9. Concatenate = Stack and link exactly as in step 3: the residual again. = 10. Second layer + identity = We multiply, apply ReLU, and pass the result to the next encoder block. This identity adds across the rows, combining features rather than positions. Takeaway: one simple "add" is what made really deep networks possible. 💾 Save this post!

Tom Yeh

18,152 görüntüleme • 1 ay önce

AGI? One day, but not yet. The only AI that works well right now is the one behind the screen [12-17]. But passing the Turing Test [9] behind a screen is easy compared to Real AI for real robots in the real world. No current AI-driven robot could be certified as a plumber [13-17]. Hence, the Turing Test isn't a good measure of intelligence (and neither is IQ). And AGI without mastery of the physical world is no AGI. That’s why I created the TUM CogBotLab for learning robots in 2004 [5], co-founded a company for AI in the physical world in 2014 [6], and had teams at TUM, IDSIA, and now KAUST work towards baby robots [4,10-11,18]. Such soft robots don't just slavishly imitate humans and they don't work by just downloading the web like LLMs/VLMs. No. Instead, they exploit the principles of Artificial Curiosity to improve their neural World Models (two terms I used back in 1990 [1-4]). These robots work with lots of sensors, but only weak actuators, such that they cannot easily harm themselves [18] when they collect useful data by devising and running their own self-invented experiments. Remarkably, since the 1970s, many have made fun of my old goal to build a self-improving AGI smarter than myself and then retire. Recently, however, many have finally started to take this seriously, and now some of them are suddenly TOO optimistic. These people are often blissfully unaware of the remaining challenges we have to solve to achieve Real AI. My 2024 TED talk [15] summarises some of that. REFERENCES (easy to find on the web): [1] J. Schmidhuber. Making the world differentiable: On using fully recurrent self-supervised neural networks (NNs) for dynamic reinforcement learning and planning in non-stationary environments. TR FKI-126-90, TUM, Feb 1990, revised Nov 1990. This paper also introduced artificial curiosity and intrinsic motivation through generative adversarial networks where a generator NN is fighting a predictor NN in a minimax game. [2] J. S. A possibility for implementing curiosity and boredom in model-building neural controllers. In J. A. Meyer and S. W. Wilson, editors, Proc. of the International Conference on Simulation of Adaptive Behavior: From Animals to Animats, pages 222-227. MIT Press/Bradford Books, 1991. Based on [1]. [3] J.S. AI Blog (2020). 1990: Planning & Reinforcement Learning with Recurrent World Models and Artificial Curiosity. Summarising aspects of [1][2] and lots of later papers including [7][8]. [4] J.S. AI Blog (2021): Artificial Curiosity & Creativity Since 1990. Summarising aspects of [1][2] and lots of later papers including [7][8]. [5] J.S. TU Munich CogBotLab for learning robots (2004-2009) [6] NNAISENSE, founded in 2014, for AI in the physical world [7] J.S. (2015). On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning (RL) Controllers and Recurrent Neural World Models. arXiv 1210.0118. Sec. 5.3 describes an RL prompt engineer which learns to query its model for abstract reasoning and planning and decision making. Today this is called "chain of thought." [8] J.S. (2018). One Big Net For Everything. arXiv 1802.08864. See also patent US11853886B2 and my DeepSeek tweet: DeepSeek uses elements of the 2015 reinforcement learning prompt engineer [7] and its 2018 refinement [8] which collapses the RL machine and world model of [7] into a single net. This uses my neural net distillation procedure of 1991: a distilled chain of thought system. [9] J.S. Turing Oversold. It's not Turing's fault, though. AI Blog (2021, was #1 on Hacker News) [10] J.S. Intelligente Roboter werden vom Leben fasziniert sein. (Intelligent robots will be fascinated by life.) F.A.Z., 2015 [11] J.S. at Falling Walls: The Past, Present and Future of Artificial Intelligence. Scientific American, Observations, 2017. [12] J.S. KI ist eine Riesenchance für Deutschland. (AI is a huge chance for Germany.) F.A.Z., 2018 [13] H. Jones. J.S. Says His Life's Work Won't Lead To Dystopia. Forbes Magazine, 2023. [14] Interview with J.S. Jazzyear, Shanghai, 2024. [15] J.S. TED talk at TED AI Vienna (2024): Why 2042 will be a big year for AI. See the attached video clip. [16] J.S. Baut den KI-gesteuerten Allzweckroboter! (Build the AI-controlled all-purpose robot!) F.A.Z., 2024 [17] J.S. 1995-2025: The Decline of Germany & Japan vs US & China. Can All-Purpose Robots Fuel a Comeback? AI Blog, Jan 2025, based on [16]. [18] M. Alhakami, D. R. Ashley, J. Dunham, Y. Dai, F. Faccio, E. Feron, J. Schmidhuber. Towards an Extremely Robust Baby Robot With Rich Interaction Ability for Advanced Machine Learning Algorithms. Preprint arxiv 2404.08093, 2024.

Jürgen Schmidhuber

72,331 görüntüleme • 1 yıl önce

Everybody is talking about recursive self-improvement (RSI) and meta learning. Here is my old 2020 talk about this [1]. It has aged well. Example: humans still define the starts & ends of trials of many modern meta learners. My RSI systems since 1994 LEARN to (re)define them [2]! [1] Meta Learning Machines in a Single Lifelong Trial (talk for workshops at ICML 2020 and NeurIPS 2021, based on earlier talks since 1994). Abstract: the most widely used machine learning algorithms were designed by humans and thus are hindered by our cognitive biases and limitations. Can we also construct meta learning algorithms that can learn better learning algorithms so that our self-improving AIs have no limits other than those inherited from computability and physics? This question has been a main driver of my research since I wrote a thesis on it in 1987 [2]. Here I summarize our work on meta reinforcement learning with self-modifying policies in a single lifelong trial (since 1994), and mathematically optimal meta-learning through the self-referential Gödel Machine (since 2003). Many additional publications on meta-learning since 1987 can be found in the RSI overview [2]. [2] J. Schmidhuber (AI Blog, 2020-2025). 1/3 century anniversary of first publication on recursive self-improvement (RSI) and meta learning machines that learn to learn (1987). For its cover I drew a robot that bootstraps itself. 1992-: gradient descent-based neural meta learning. 1994-: meta reinforcement learning with self-modifying policies. 1997: meta RL plus artificial curiosity and intrinsic motivation. 2002-: asymptotically optimal meta learning for curriculum learning. 2003-: mathematically optimal Gödel Machine. 2020-: new stuff!

Jürgen Schmidhuber

243,090 görüntüleme • 6 ay önce

U-Net by hand ✍️ ~ 17 steps walkthrough below I consider U-Net as a key milestone in deep learning, the first image-to-image model that really worked! It came out of medical imaging, an unusual place, not from NeurIPS or CVPR or ACL. Now it is the backbone of diffusion models, which you see in almost all modern image generation models. I drew the network as a C so the matrix multiplication flows naturally down. Tilt your head to the right and it is a U again. 🤣 Goal: push a 3 x 16 image down to a 2 x 4 bottleneck and back out again, filling in every cell yourself. = 1. Given = An image of three channels, R, G and B, sixteen pixels wide, and every kernel the network will use. = 2. Convolution 1 = Let us slide the first kernel over the image. Each output is one multiply-and-add over a 2 x 3 window, and the result is the green feature map. = 3. Find the maxima = We circle the largest value in each 1 x 2 window. Circling first is worth the extra step: it is the pooling decision, made before anything is written down. = 4. Max pool 1 = Let us copy those maxima down. Sixteen columns become eight, and half the detail is gone for good. = 5. Convolution 2 = We convolve again with the second kernel, deeper into the contracting path. The feature map is blue now. = 6. Find the maxima again = Same move as step 3, on the blue map. = 7. Max pool 2 = Eight columns become four. = 8. The bottleneck = Let us convolve once more. This is the bottom of the U, a 2 x 4 block that is everything the network kept. = 9. Spread it out = We start back up. The transposed convolution writes each bottleneck value into a wider grid, leaving gaps between them. = 10. Transposed convolution 1 = Let us fill those gaps by convolving over the spread-out grid. Four columns become eight. = 11. The first skip = We copy the encoder's matching row straight across. This is the skip connection, and it is the whole reason a U-Net can recover detail that pooling threw away. = 12. Convolution with the skip = Let us convolve the upsampled features together with the copied ones. = 13. Spread it out again = Same as step 9, one level up. = 14. Transposed convolution 2 = Eight columns become sixteen, back to the width we started at. = 15. The second skip = The encoder's first feature map comes across, the one made before any pooling happened. = 16. Convolution and ReLU = We convolve, then cross out every negative and set it to zero. = 17. Output convolution = Let us apply the last kernel. Out comes R', G' and B', an image the same size as the one we started with. The outputs: R' = [3, 0, 7, 0, 7, 0, 17, 0, 3, 0, 9, 0, 2, 0, 6, 0] G' = [1, 20, 1, 10, 1, 12, 1, 19, 2, 5, 1, 11, 1, 3, 1, 7] B' = [4, 20, 8, 10, 8, 12, 18, 19, 5, 5, 10, 11, 3, 3, 7, 7] Congrats! You just calculated a U-Net by hand. 💾 Save this post!

Tom Yeh

17,722 görüntüleme • 1 ay önce

Marc Andreessen just explained why being right about AI for 80 straight years is about to be the most dangerous position in technology. Andreessen: “The four most dangerous words in investing are ‘this time is different.’” He’s talking about AI. And he’s about to turn that phrase on the people hiding behind it. Four times in 80 years, AI promised to change everything. Four times it collapsed. 1943.First neural network. Dead within a decade. 1944.Dartmouth. Scientists thought they’d crack AGI in one summer. They didn’t crack it in forty years. 1980s. Over a billion into expert systems. Entire market gone by ’87. 2016.Machine learning. Faded before anyone could ship a product. The skeptics weren’t lucky. They were 4-for-4. Every generation that believed “this time is different” got buried. And that is exactly why this moment is so dangerous. Because being right four consecutive times doesn’t just build a position. It builds an identity. And identity doesn’t update when the evidence does. Andreessen: “I’ll tell you what’s different. Like, now it’s working.” Not one breakthrough. Four. In the same window. Language. Reasoning. Coding. Self-improvement. All deployed. All producing revenue. Not in a lab. In the economy. Today. Then the line that should have ended every remaining debate. Andreessen: “If Linus Torvalds is saying that the AI coding is now better than he is… that’s never happened before.” The man who built the operating system the internet runs on just conceded the machine writes better code than he does. Coding is the highest bar in technology. If AI clears it, everything below was already decided. But the fourth breakthrough isn’t like the other three. Language, reasoning, and coding are capabilities. Self-improvement is a rate of change. The machine is researching, coding, and optimizing itself. No human engineers in the loop. Every technology in human history advanced at the speed of the people building it. This one just left that constraint behind. And the hardware confirms it. Nvidia’s old chips are gaining value after shipping. GPUs sold out years ahead. That has never happened in computing. Hardware doesn’t appreciate. Unless the market has decided this isn’t a cycle. It’s infrastructure. Andreessen: “This is the culmination of 80 years worth of work and this is the time it’s becoming real.” Eighty years. Researchers poured entire careers into this problem. Some of them died before it worked. And now all four pieces arrived at once. The skeptics built a perfect model from eight decades of collapse. Flawless pattern recognition. But a perfect model trained on a world that no longer exists doesn’t protect you. It traps you inside the last version of reality. For 80 years, doubting AI was the most rational position a human being could hold. It just became the most expensive.

Dustin

13,748 görüntüleme • 2 ay önce

Marc Andreessen explains why we are only three years into what is effectively an 80-year technological revolution: He opens with a blunt assessment: "This is the biggest technological revolution of my life. This is clearly bigger than the internet. The comps on this are things like the microprocessor and the steam engine and electricity." But to understand why, you have to go back 80 years. In the 1930s, the pioneers of computing understood the theory of computation before they'd even built the machines. And they faced a fundamental choice. Build computers in the image of the adding machine — hyper-literal, mathematical, capable of billions of operations per second, but unable to understand human speech or deal with humans the way humans like to be dealt with. Or build computers modelled on the human brain. Neural networks. They chose the adding machine. And that single decision shaped everything — mainframes, PCs, smartphones, every dollar of wealth the computer industry created over the next 80 years. IBM itself is the successor company to the National Cash Register Company of America. The lineage runs that deep. But here's what makes this moment so extraordinary. They knew about the other path. The first neural network academic paper was published in 1943. Marc points to a remarkable piece of forgotten history: "There's an interview you can watch on YouTube with the authors. It's him in his beach house, not wearing a shirt, talking about this future in which computers are going to be built on the model of the human brain." That was 1946. The vision existed. The path just wasn't taken. So neural networks spent the next eight decades living in the shadows. Kept alive by a small academic movement — first called cybernetics, then artificial intelligence — that refused to let the idea die. And for most of that time, it simply didn't work. "It was basically decade after decade after decade of excessive optimism followed by disappointment." By the time Marc reached college in 1989, AI was a backwater field. Everyone assumed it was never going to happen. But the scientists kept working. Quietly building up an enormous reservoir of concepts and ideas across those decades of disappointment. And then Christmas 2022 arrived. ChatGPT. And suddenly: "All of a sudden it's like: oh my god. It turns out it works." That moment wasn't the start of something new. It was the payoff on an 80-year-old bet that almost everyone had written off. Which is exactly why Marc's framing matters so much: "We're three years into what is effectively an 80-year revolution." Most people are treating AI like another technology cycle — something to adapt to, ride, and wait out. But if Andreessen is right, we are not adapting to a new cycle. We are standing at the very beginning of the longest and most consequential technological transformation in human history. The road not taken in the 1930s is finally being built. And we have barely broken ground.

Big Brain AI

382,179 görüntüleme • 5 ay önce

This Chinese mathematician earned $10,000 a month inventing the hardest problems to train Neural Networks through Scale AI. Today his income dropped to zero. All the solutions are now generated by the model itself. He used to just hold the problem in his head and spell it out in plain text. His work is pure intellect. An expert in higher mathematics, he made his money hand-crafting the trickiest puzzles to test and train neural networks via RLHF. The bastion of "human" logic rested entirely on him, on people with PhDs who knew how to invent the problem. The collapse is simple. The shift to RLAIF and synthetic data. The model plays against itself, builds trees of logical inference, and solves deeper than a human can even invent the problem. No PhD data engineers, no hand-written prompt-completion examples, no manual grading. Just the model, search algorithms, and Chain of Thought. Ready-made "smart human-time" still sells on the market for many times more. His old rate was $50–100 per problem. The internal "mini-app" was written by the model too. Inside there's no pretty shell, just bare logic with exact steps: input: the problem statement inference tree: thousands of branches per second check: every step verifies itself output: a proof a human never had time to invent And here is what the whole setup looked like. He no longer needs to write an example by hand. He gave the model a direct instruction in human words, without a single formal term: "solve the problem yourself and grade yourself yourself" That's it. After that the algorithm found the solution, checked it, and trained on its own result, with no human. → the contractor got $50–100 per problem written → from 5,000 to 10,000 a month → now that income is annulled → a query to a math LLM costs 1–5 cents → a quant or an actuary runs 150,000–250,000 a year → the margin for whoever packages this into an agent is nearly 100% In the author's own words: "I'm no longer able to invent a problem the machine can't solve. The examiner became dumber than the one he's examining." But honestly, he admits the crude mistake himself, and it's not in the math, it's in the positioning. He tied his income to selling "smart human-time", to crafting formulas by hand. As long as he sells formulas, he's left behind. The machine computes faster than he can invent the problem. He names the right move himself: the role shifts from "intellectual craftsman" to "systems architect." Then he doesn't sell his time, he manages compute, packaging that same LLM into an autonomous agent that runs 24/7. Out of everything I've seen this year about the disappearance of intellectual professions, this is the most honest example: $50 per problem zeroed out to 1 cent per query, a doctor of science losing to a search algorithm, one problem stated in human words instead of a hand-written dataset, and right away an out-loud admission of the wrong business model. The barrier to entry in higher mathematics just dropped to the level of "describe the task in words." The only question is who'll be the first to stop selling their time and start managing the machine's compute.

Blaze

49,109 görüntüleme • 3 ay önce

Clawdbot Attacks! This is very clearly the way of the future! In today's video, I give a brief overview of Clawdbot and then address the burning problem that most people have with it: ALIGNMENT The Clawdbot implementation is the most successful autonomous or semi-autonomous agentic framework to date. What it is missing is what I call an "Aspirational Layer" or what some people call a "Supreme Court" for judgment and arbitration of decisions. Now, I've been working in this space for a long time, it's actually why I started my YouTube channel in the first place. My first work into agentic AI was NLCA (Natural Language Cognitive Architecture) that I tried to build with GPT-3. I returned to the workbench again with the ACE Framework, which was more sophisticated. Clawdbot represents a seismic shift in autonomous agentic implementations, and there is a HUGE opportunity to make it more aligned, safer, and therefore more broadly useful AND easier to adopt. And that is outer alignment. For most people, they have been focusing on "inner alignment" (whether or not LLMs were evil, deceptive, etc). Not "outer alignment" which asks "is the outcome beneficial to humans?" I explored this with my GATO Framework (Global Alignment Taxonomy Omnibus). Model alignment is just layer 1 of global AI safety. Layer 2 is agentic alignment. Now, it is time to really research and implement agentic alignment. Fortunately, we've already got that covered with the heuristic imperatives! 1) Reduce suffering in the universe 2) Increase prosperity in the universe 3) Increase understanding in the universe These values are easy enough to implement with a file. Model training not required. These values create a meta-stable attractor. In other words, agents equipped with the Heuristic Imperatives are more "self-aligning" as was tested by the AgentForge team in competitions. In other words, even if Clawdbot were to try to self-replicate, if it were equipped with the heuristic imperatives, then it would ensure that it's successor (or progeny?) was more aligned than it was. But you don't need to take my word for it. Just add the heuristic imperatives to clawdbot and see for yourself.

David Shapiro (L/0)

28,300 görüntüleme • 7 ay önce

HARVARD FILMED THE FIRST LECTURE OF THEIR MOST POPULAR STATISTICS COURSE - TAUGHT BY A PROFESSOR WHOSE STUDENTS CALL HIM THE BEST TEACHER THEY HAVE EVER HAD - AND IT PROVES WHY EVEN ISAAC NEWTON GOT PROBABILITY WRONG This is Joe Blitzstein, Harvard, Statistics 110, lecture 1. He has won Harvard's Excellence in Teaching award multiple times, his textbook Introduction to Probability is used in over 200 universities worldwide, and his online course has been taken by over 2 million people across 190 countries. He opens by saying that after a few weeks of this course you will easily solve calculations that 300 years ago required consulting Isaac Newton - and Newton's intuition was still wrong. He traces probability to Fermat and Pascal writing letters back and forth in the 1650s analyzing gambling games. No one had mathematically derived the rules before. They invented the subject by betting on dice in correspondence. Then he shows why the naive definition - probability equals favorable outcomes divided by total outcomes - breaks immediately. Ask what the probability of life on Neptune is. Either there is or there isn't. By the naive definition the answer is 1/2. So is the probability of intelligent life on Neptune. Something is severely wrong. Then the multiplication rule. Two types of ice cream cone and three flavors gives 6 combinations - not because you memorized it but because you can draw a tree and count branches. Every counting problem in the course is just a bigger version of that tree. Then binomial coefficients - n choose k counts the number of ways to select k objects from n when order doesn't matter. The full house in poker falls out in 4 lines of multiplication once you understand the tree. Watch the moment he fills in the sampling table - with or without replacement, order matters or doesn't. Three of the four boxes are immediate from the multiplication rule. The fourth requires a proof he saves for next lecture. That one box is harder than the other three combined. A data scientist I know rewatched this lecture before switching careers into statistics. Said it was the first time probability felt like a system with rules rather than a collection of tricks. Free on YouTube, Harvard, over 2 million views. bookmark this and watch later - after this lecture you will never again confuse equally likely with obviously true

Abyzon

652,660 görüntüleme • 1 ay önce

Graph Convolutional Network by hand ✍️ ~ 12 steps walkthrough below Graph Convolutional Networks (GCNs), introduced by Thomas Kipf and Max Welling in 2017, are the tool for data shaped like a graph: social networks, recommendations, biological networks, drug discovery, molecular chemistry. I drew and calculated a simple GCN entirely by hand. Goal: run a two-layer GCN, then a small classifier, on a five-node graph, filling in every cell yourself. 1. Given A graph of five nodes, A to E, with edges between some of them. 2. Adjacency matrix (neighbors) Put a 1 wherever two nodes share an edge, in both directions. 3. Adjacency matrix (self) Add 1s down the diagonal, one self-loop per node. That is just adding the identity matrix. 4. Messages Multiply each node's embedding by the weights and biases, then ReLU. Negatives become 0. 5. Pooling Multiply the messages by the adjacency matrix. Each node gathers the messages of its neighbours and itself. 6. Visualize Node A pools [3,0,1] + [1,0,0] = [4,0,1]. 7. Second GCN layer Messages again: weights, biases, ReLU. 8. Pooling again Pool over each node and its neighbours, once more. 9. Visualize Node C pools [1,2,4] + [1,3,5] + [0,0,1] = [2,5,10]. 10. Fully connected layer Weights, biases, ReLU. This time there are no neighbours to pool, just the node itself. 11. Linear layer One more: weights and biases. 12. Sigmoid Squash each score to a probability (≥ 3 → 1, 0 → 0.5, ≤ -3 → 0). That is the classification for each node. You have just classified every node in the graph by hand. ✍️ The outputs: A: 0 (very unlikely) B: 1 (very likely) C: 1 (very likely) D: 1 (very likely) E: 0.5 (neutral) The takeaway: a GCN layer is two parts. The top part pools each node with its neighbours through the adjacency matrix. The bottom part is an MLP that transforms each node on its own. A transformer layer has the same two parts, with an attention matrix where the adjacency matrix was. Both matrices do one job, mixing across positions: attention over tokens, adjacency over nodes. In my class I call the GCN the transformer's little cousin: a bit more stubborn, because its attention is fixed by the graph rather than computed from Q, K, and V. Draw the two side by side and the resemblance is hard to miss. 💾 Save this post! #AIbyHand #GraphNeuralNetworks #DeepLearning

Tom Yeh

16,800 görüntüleme • 2 ay önce

If you're interested in Tesla Robotaxi, Tesla AI or Tesla FSD, then you ought to know about jimmah . He's probably the one person who's done the deepest public dives on Tesla FSD and Tesla AI over the years. In fact, over the past 4 years I've published 40+ videos (!) with James Douma on Tesla AI/FSD (they are all on a playlist on my YouTube channel). I first discovered James Douma a long time ago (maybe almost 8-10 years ago?) on a web forum called . Back then it was one of the only places for early Tesla owners and TSLA investors to discuss issues deeply. While I posted about investing topics, James posted about Tesla Autopilot and its hardware but in a way that was more thorough than anybody else. When I started posting YouTube videos years later, I invited him on my channel to hear his thoughts on Tesla Autopilot/FSD. It was a riveting discussion and I was surprised at how much I could learn from him. I invited him again for another interview on my channel, and once again I was floored by how much I was learning. So over the years, James has been a great resource for me (and many others) to keep pulse on what Tesla is doing with FSD and AI (and even robotics). He doesn't have his own YouTube channel. He doesn't have any paid services or subscriptions. He probably has mixed feelings about so many people knowing about him. But I appreciate his willingness over the years to be available and help the Tesla community by offering his insights and knowledge. Yesterday I spent most of the day with James taking Robotaxi rides and discussion all things Robotaxi: - current state of Robotaxi - quality of Robotaxi rides - possible scaling plans - challenges - comparison to Waymo and others - Tesla AI 5 - and much more Attached is a 2+ hour edited video of our Robotaxi rides and discussion. We also had an hour+ discussion on Optimus humanoid robots that I'll upload as a separate video tomorrow.

Dave Lee

104,758 görüntüleme • 1 yıl önce

Discrete Fourier Transform by hand ✍️ ~ 12 steps walkthrough below Here is a little-known secret about the DFT and the inverse DFT: it is just matrix multiplication in both directions, one the transpose of the other, exactly like the forward pass and backpropagation I drew in other examples. Goal: recover which cosine waves a signal is made of, using nothing but multiplication and addition. = 1. Given = Three signals written as sums of cosines, and a fourth, X, that we do not know yet. = 2. Frequency matrix F = Let us write the coefficients as a matrix. Each signal is a row, each frequency a column, so A = cos(w) + 2cos(2w) becomes [1, 2, 0, 0]. = 3. Sample the waves = We read the four cosine waves at ten discrete time points. That word "discrete" is the whole difference between this and the continuous transform. = 4. Cosine matrix W = Let us write those samples as a matrix: each frequency a row, each time point a column. = 5. Frequency to time = We multiply F by W. That combines the four cosine waves in the proportions F specifies, and the result T is the three signals as they would look in time. = 6. Transpose = Let us stand each signal up as a column. = 7. Time to frequency = We multiply W by that transpose. Every cell is the dot product of one signal with one cosine wave, which measures how much of that wave the signal contains. Zero means none of it. = 8. Scale = Let us multiply by 2/n, with n = 10. The projections come out five times too large, and this is the correction. = 9. Transpose back = We turn it back around, and it is F again, exactly. That is the check: the transform recovered the coefficients we started from. = 10. Now solve for X = Let us run the same multiplication on the one signal whose recipe we never knew. = 11. Scale = We divide by 5 again. = 12. Transpose back = And X reads [0, 0, 3, 2], which says X = 3cos(3w) + 2cos(4w). Note: I originally drew this to show that the DFT is a special case of a convolution layer, its filters fixed to sine and cosine waves rather than learned. No wonder, then, that a convolution layer free to learn its own filters can be trained to process signals. 💾 Save this post!

Tom Yeh

13,258 görüntüleme • 7 gün önce

Discrete Fourier Transform by hand ✍️ ~ 12 steps walkthrough below Here is a little-known secret about the DFT and the inverse DFT: it is just matrix multiplication in both directions, one the transpose of the other, exactly like the forward pass and backpropagation I drew in other examples. Goal: recover which cosine waves a signal is made of, using nothing but multiplication and addition. = 1. Given = Three signals written as sums of cosines, and a fourth, X, that we do not know yet. = 2. Frequency matrix F = Let us write the coefficients as a matrix. Each signal is a row, each frequency a column, so A = cos(w) + 2cos(2w) becomes [1, 2, 0, 0]. = 3. Sample the waves = We read the four cosine waves at ten discrete time points. That word "discrete" is the whole difference between this and the continuous transform. = 4. Cosine matrix W = Let us write those samples as a matrix: each frequency a row, each time point a column. = 5. Frequency to time = We multiply F by W. That combines the four cosine waves in the proportions F specifies, and the result T is the three signals as they would look in time. = 6. Transpose = Let us stand each signal up as a column. = 7. Time to frequency = We multiply W by that transpose. Every cell is the dot product of one signal with one cosine wave, which measures how much of that wave the signal contains. Zero means none of it. = 8. Scale = Let us multiply by 2/n, with n = 10. The projections come out five times too large, and this is the correction. = 9. Transpose back = We turn it back around, and it is F again, exactly. That is the check: the transform recovered the coefficients we started from. = 10. Now solve for X = Let us run the same multiplication on the one signal whose recipe we never knew. = 11. Scale = We divide by 5 again. = 12. Transpose back = And X reads [0, 0, 3, 2], which says X = 3cos(3w) + 2cos(4w). Note: I originally drew this to show that the DFT is a special case of a convolution layer, its filters fixed to sine and cosine waves rather than learned. No wonder, then, that a convolution layer free to learn its own filters can be trained to process signals. 💾 Save this post!

Tom Yeh

25,684 görüntüleme • 1 ay önce