Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

📁 Terence Tao, Fields Medalist mathematician, explains that neural networks do not understand concepts, but they uncover hidden patterns. In knot theory, a classical neural network discovered correlations between mathematical invariants no one had anticipated. It started as a black box, then became a real clue to deep structures...

204,985 Aufrufe • vor 7 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

AI's Secret Pattern: The Surprising Role of Fractals in Neural Networks In the realm of artificial intelligence (AI), a groundbreaking discovery has emerged, challenging our conventional understanding of neural network training and optimization. This revelation centers around the identification of fractal patterns at the boundary between trainable and untrainable neural network hyperparameters, presenting a series of profound implications and avenues for further research. Fractals, known for their intricate, self-similar patterns that recur at every scale, have long fascinated mathematicians and scientists alike. Typically associated with simple, one-dimensional iterative functions, the appearance of fractals within the complex, multivariate domain of neural network training introduces a striking contrast. The organic and asymmetric nature of these fractals, as derived from the training processes, suggests a deeper, unexplored connection between the mathematical properties of fractals and the functional dynamics of neural networks. The study’s focus on two-dimensional slices of hyperparameter space barely scratches the surface of the complexity inherent in neural networks, which are characterized by a vast array of hyperparameters. The existence of fractals in this context hints at an underlying high-dimensional structure, a concept that challenges our current capabilities and understanding. Extending fractal analysis to these higher dimensions represents a significant, yet exciting, challenge that could illuminate new aspects of neural network behavior and learning capabilities. An unexpected finding from the research is the persistence of clean fractal patterns even in the presence of stochastic elements introduced during minibatch training. This resilience suggests a parallel to Lyapunov fractals, where the iterative process involves randomly changing functions. This phenomenon prompts a reevaluation of how stochastic and deterministic processes influence fractal formation within neural networks, potentially offering new insights into the fundamental mechanisms of learning and adaptation. From a practical standpoint, the fractal nature of the boundary between trainable and untrainable hyperparameters has significant implications for the field of metalearning. The chaotic behavior of the meta-loss landscape, attributed to its extreme sensitivity, presents a formidable challenge for algorithms designed to optimize hyperparameters. Understanding the fractal characteristics of this landscape could provide valuable guidance for navigating its complexities, ultimately improving the efficiency and effectiveness of metalearning strategies. Beyond the technical and theoretical implications, the discovery also reveals an unexpected aesthetic dimension to neural network fractals. The visual beauty and meditative qualities of these patterns offer a unique opportunity to engage with the material in a deeply personal and contemplative manner. This aspect suggests potential psychological and physiological benefits from exposure to the intricate designs of neural network fractals, opening up novel intersections between technology, art, and well-being. In conclusion, the identification of fractal patterns within neural network hyperparameter spaces unveils a fascinating new frontier at the intersection of fractal geometry and deep learning. This discovery not only challenges existing paradigms but also opens up myriad possibilities for mathematical characterization, algorithmic development, and even subjective exploration. As researchers continue to delve into this rich vein of inquiry, the promise of uncovering new knowledge and advancing our understanding of neural networks and their training processes remains as compelling as ever.

Carlos E. Perez

133,529 Aufrufe • vor 2 Jahren

Mathematician Terence Tao offers a counterintuitive take: AI doesn't look intelligent because our definition of intelligence was wrong all along. He argues that the entire history of AI has followed a predictable pattern: "The history of AI has been here's a task that only humans can do, like maybe it is read natural language or win at chess or solve a math problem, and then one by one someone finds some AI algorithm that also does that." But every time a machine cracks one of these "uniquely human" tasks, we move the goalposts. The solution never feels like real thinking: "You look at how it's done and it doesn't feel like intelligence. It's, oh, it was some trick. You just cobbled together these neural networks and you ran some algorithm, and we were looking for some elusive intelligent way of thinking, and we don't see it in the tools that actually solve our goals." Tao then flips the problem on its head. What if the issue isn't with the machines, but with us? "But maybe it's actually because intelligence is not what we think it is." He points to large language models as the clearest case. What they do sounds almost embarrassingly simple: "Large language models in particular become very successful, and a lot of what they're doing is just predicting the next token, clicking the next word in a sentence. And that doesn't sound like something which is intelligent." To show why this feels wrong, Tao draws a comparison to how we'd judge a human doing the same thing: "If you ask someone to improvise a speech and they have no preparation, and at every moment they're just saying the next word that comes to their mind, you don't think that this could actually work." And yet it works for LLMs. Which forces an uncomfortable possibility: "Maybe that's actually a lot of what humans do as well."

Big Brain AI

69,490 Aufrufe • vor 3 Monaten

Sergey Brin on AI and Jeff Dean back in December when Gemini peaked: “I would say in some ways we for sure messed up in that we underinvested and sort of didn't take it as seriously as we should have, say, eight years ago when we published the Transformer paper. We actually didn't take it all that seriously and didn't necessarily invest in scaling the compute. And also we were too scared to bring it to people because chatbots say dumb things. And, you know, OpenAI ran with it, which is good for them. It was a super smart insight, and it was also our people like Ilya who went there to do that. But I do think we still have benefited from that long history. So we had a lot of the research and development of neural networks kind of going back to Google Brain. That was also kind of lucky. It wasn't luck that we hired Jeff Dean. I mean, we were lucky to get him. But we were in this sort of mindset that deep technical things matter. And so we hired him. We hired a lot of people from DEC, honestly, because they had the top research labs at the time. But he [Jeff Dean] was passionate about neural networks. And it stemmed, I think, from actually his college experiment. I don't know. He was like, whatever, curing third world disease and figuring out neural networks when he was like 16. He's done crazy things. But he was passionate about it. He built up a whole effort. And actually in my division at the time in Google X. We had him but I didn't I was like okay Jeff you do whatever you want. He's like oh we can tell cats from dogs. I'm like oh okay cool. But you know, you also trust your technical people and soon enough they were developing all these algorithms these neural nets that were doing some of our search. And then, you know, Noam came up with a transformer and we were able to do more and more. But yeah, I mean, so we had the underpinnings, we had the R&D. We did underinvest for a number of years and didn't take it as seriously as we should have.” Both Jeff Dean and Noam Shazeer left Google in the past two months.

tae kim

247,281 Aufrufe • vor 12 Tagen

This video, created by my dear coauthor Mahdi E Kahou for our teaching and papers, shows how overparameterized neural networks produce smooth function approximations even in the context of the Runge phenomenon. Some background. Imagine you want to approximate the Runge function using polynomial interpolation at equally spaced points. It is well known that, despite targeting an infinitely differentiable function, such a polynomial approximation produces oscillatory behavior that worsens with the degree of the polynomial. In other words, higher-degree polynomial approximations might not improve accuracy. Instead, approximate the Runge function with a neural network (here, two layers are just to make the example concrete; nothing fundamental depends on it). As you increase the number of parameters well above the 11 training points (in our example, a two-layer neural network with 128 nodes each), you nicely converge to the target, without wild oscillations. Yes, this has much to do with double descent and benign overparameterization, but the main punchline of this post is that neural networks are really very different types of animals than polynomial approximations. And yes, Chebyshev nodes and splines exist, and in this case, they will prevent the oscillations. But that's not the point. Chebyshev nodes and splines still confront Faber’s theorem, which states that for any system of polynomial interpolation nodes, there exists a continuous function whose sequence of interpolating polynomials diverges as the number of nodes grows to infinity. Faber’s theorem does not apply to neural networks because they are not polynomials. The notebook, if you want to check the details, is here: Stay tuned for more on this 👀

Jesús Fernández-Villaverde

47,212 Aufrufe • vor 3 Monaten

Terence Tao has an IQ above 200. Youngest gold medalist in Math Olympiad history. Fields Medal winner. The greatest living mathematician by nearly any measure. And he just said something most people aren’t ready for. Tao: “This whole era of AI is teaching us that our idea of what intelligence is, is not really accurate.” We spent centuries building civilization on one assumption. That intelligence was sacred. Irreducible. Uniquely ours. The one thing that made the entire human story make sense. Then AI started solving things we swore only we could. Chess. Language. Vision. Math. And every time, we reached for the same defense. That’s not real intelligence. It’s just tricks. Just pattern matching. Just an algorithm. Tao: “You look at how it’s done and it doesn’t feel like intelligence.” So we moved the line. Again. And again. And again. Because intelligence was supposed to feel like something. Something deep. Something we could point to and say… this is what separates us from everything else. But AI kept solving the problems. And that feeling never arrived. Tao: “We were looking for some elusive, intelligent way of thinking and we don’t see it in the tools that actually solve our goals.” Here’s what makes it worse. Large language models work by predicting the next word. One word at a time. No grand architecture. No deep understanding. Just probability. And it works. Tao: “Maybe that’s actually a lot of what humans do as well.” The greatest living mathematician just told you human thought might run on the same machinery. Not some transcendent spark. Pattern recognition. Prediction. One thought, one decision, one word at a time. We built religion around intelligence. Philosophy around it. An entire species identity around it. And a machine running probability just held up a mirror. We didn’t lose intelligence to AI. We just finally saw what it always was. What haunts us isn’t that machines learned to think. It’s that thinking was never what we needed it to be.

Dustin

564,499 Aufrufe • vor 3 Monaten

How were humans able to recognize that Newton's laws of motion govern both the flight of a bird and the motion of a pendulum? This ability to identify the same mathematical patterns across vastly different contexts lies at the heart of scientific discovery—whether studying the aerodynamics of bird wings or designing the blades of a wind turbine. Yet, AI systems often struggle to discern these deep structural similarities. 💡The key may lie in mathematical isomorphisms—patterns that preserve their relationships regardless of context. For example, the same principles of fluid dynamics apply to blood flowing through arteries and air streaming over an airplane wing, or the motion of a molecule. This raises a fundamental question in artificial intelligence: how can we enable machines to understand the world through these invariant structures rather than surface features? 🚀Our work introduces Graph-Aware Isomorphic Attention, improving how Transformers recognize patterns across domains. Drawing from category theory, models can learn unifying structural principles that describe phenomena as diverse as the hierarchical assembly of spider silk proteins and the compositional patterns in music. By making these deep similarities explicit, Isomorphic Attention enables AI to reason more like humans do—seeing past surface differences to grasp fundamental patterns that unite seemingly disparate fields. Through this lens, AI systems can learn and generalize, moving beyond superficial pattern matching to true structural understanding. The implications span from scientific discovery to engineering design, offering a new approach to artificial intelligence that mirrors how humans grasp the underlying unity of natural phenomena. Some key insights include: 1️⃣ Graph Isomorphism Neural Networks (GINs): GIN-style aggregation ensures structurally distinct graphs map to distinct embeddings, improving generalization and avoiding relational pattern collapse. 2️⃣ Category Theory Perspective: Transformers as functors preserve structural relationships. Sparse-GIN refines attention into sparse adjacency matrices, unifying domain knowledge across tasks. 3️⃣ Information Bottleneck & Sparsification: Sparsity reduces overfitting by filtering irrelevant edges, aligning with natural systems. Sparse-GIN outperforms dense attention by focusing on crucial connections. 4️⃣ Hierarchical Representation Learning: GIN-Attention captures multiscale patterns, mirroring structures like spider silk. Nested GINs model local and global dependencies across fields. 5️⃣ Practical Impact: Sparse-GIN enables domain-specific fine-tuning atop pre-trained Transformer foundation models, reducing the need for full retraining. Other impacts: ✅Real-World Relevance: Whether we are looking at protein structures, designing new materials, or working on social network analytics, graph-aware Transformers can capture subtle relational patterns traditional architectures may miss. ✅The juncture of graph isomorphism theory, category theory, and sparsification, these GIN-Transformers step beyond sequential modeling to tackle the relational nature of complex data. #Transformers #GraphNeuralNetworks #AI #MachineLearning #Isomorphism #CategoryTheory #ArtificialIntelligence #DeepLearning Link to paper & code in response ⤵️

Markus J. Buehler

19,416 Aufrufe • vor 1 Jahr