Loading video...

Video Failed to Load

Go Home

1/n Tweeprint 📢📢📢 on neuro-symbolic AI: Are basic neural networks (2-layer MLP) able to learn simple algorithms (addition of real numbers)? Prior theory says they should be and yet, during training, we see a highly complex error surface evolving

77,397 views • 3 years ago •via X (Twitter)

10 Comments

David Klindt's profile picture
David Klindt3 years ago

2/n Plotting the prediction error in input space (2D), we see MLPs learn a ridge of good solutions through the training data with increasing complexity for wider networks. Intriguingly, we see a similar behavior in Gaussian processes with small (RBF kernel) length scale

David Klindt's profile picture
David Klindt3 years ago

3/n The pattern adapts to the input domain (e.g., annulus) and becomes more and more intricate for wider networks and larger training datasets

David Klindt's profile picture
David Klindt3 years ago

4/n Following the construction in Neal 1996 ( we can control the smoothness of the function learned by the network – analogous to a GP

David Klindt's profile picture
David Klindt3 years ago

5/n Going back to addition, we can see the effect of making the networks more smooth – becoming more similar to a GP with well tuned length scale

David Klindt's profile picture
David Klindt3 years ago

6/n This has implications for neural network robustness / out-of-distribution generalization (see paper). However, it does not solve the larger problem that MLPs cannot make the (inductive) inference step from a (complex) look-up table to an underlying algorithm.

David Klindt's profile picture
David Klindt3 years ago

7/n This was the most minimal setting I could think of to study that problem. I hope that future work will look into other models (e.g., transformers, GNNs) and try to understand this phenomenon of the evolving loss surfaces in MLPs.

David Klindt's profile picture
David Klindt3 years ago

8/n This work was inspired by @MLStreetTalk and @PetarV_93's work on neural algorithmic reasoning:

David Klindt's profile picture
David Klindt3 years ago

9/n I am still extremely excited about these ideas, but I think that it needs something more in the mix. Maybe not just new models, but statistical theory about minimal failures of most basic networks, and epistemological theory about conceptual challenges of inductive inference.

David Klindt's profile picture
David Klindt3 years ago

Lastly, I want to thank @schott_lukas @Jingyang_zhou @TonyZador, anonymous reviewers @TmlrOrg & others for excellent feedback 🙏 This was my first, and frankly, last solo project. It was an interesting experience, but honestly, science is 10x more fun and productive in a team 😁

Sara A Solla's profile picture
Sara A Solla3 years ago

Prior theory says that a two-layer network with a sufficiently wide hidden layer of nonlinear units should be able to implement such a function. It does not say that it should be able to learn it from examples.

Related Videos

What if #AI became as decentralized as #Bitcoin? We sat down with our new friend 3700 from Bitcoin Virtual Machine to hear what their incredible team of anons are working on - "Truly Open AI." Full interview here:👇 1: What positive impact will Layer 2s have on Bitcoin? Layer 2s on Bitcoin open up opportunities for innovation, allowing developers to build dApps and smart contracts on top of Bitcoin, expanding its utility and use cases. By submitting transactions for final settlement on the Bitcoin network, Bitcoin Layer 2 networks claim to achieve the same (or close to) level of security and decentralization as the Bitcoin blockchain. Building a separate execution layer allows them the freedom to employ several technologies (such as rollups). Layer 2 can significantly improve Bitcoin's scalability by processing transactions off-chain, reducing congestion on the main blockchain. Overall, Layer 2s on Bitcoin have the potential to address some of Bitcoin's key limitations, making it more efficient, accessible, and versatile in the long run. 2: What does the ETF approval mean for Layer 2 on Bitcoin? The approval of ETF could potentially have several implications for Layer 2 on Bitcoin: Innovation and Development: With a growing interest in Bitcoin spurred by ETF approval, there could be a surge in research and development efforts focused on enhancing Layer 2. Developers and projects may be incentivized to create new and improved Layer 2 protocols to meet the evolving needs of the expanding Bitcoin ecosystem. An ETF approval could boost mainstream Bitcoin adoption and liquidity. This influx of users may also drive interest in Layer 2 on Bitcoin as a means to enhance the scalability and functionality of Bitcoin. 3: What are the primary challenges facing L2s on Bitcoin? The interoperability of different Layer 2s and their compatibility with Bitcoin's main blockchain can be a challenge. Ensuring seamless interaction between various Layer 2 networks and the Bitcoin blockchain is essential for a cohesive and efficient ecosystem. Some Layer 2s may introduce centralization risks if they rely heavily on centralized entities or trusted intermediaries. Maintaining decentralization and censorship resistance, which are core tenets of Bitcoin, while scaling with Layer 2s is a challenge. 4: What aspects of Layer 2 solutions for Bitcoin are you most enthusiastic about? AI represents one of the cornerstones of our modern era. However, achieving a decentralized AI infrastructure, owned and managed by users, has posed significant challenges. The primary obstacle has been the limited capacity to store and execute AI models due to size and computational limitations. To address this challenge, we propose a new blockchain architecture enabling developers to deploy their own Bitcoin Layer 2 solutions tailored specifically for AI tasks, called Truly Open AI. These Layer 2 blockchains are optimized to handle computationally intensive tasks, such as matrix multiplication, directly on-chain. These Bitcoin Layer 2 solutions offer exceptional throughput, minimal latency, and cost-effectiveness. AI dApps are programmed as Solidity smart contracts, ensuring they operate precisely as intended, free from interference or manipulation. Our BVM AI Contracts Library simplifies the integration of neural networks into dApps, empowering developers to embed AI seamlessly. In summary, I'm particularly enthusiastic about the potential of Layer 2 solutions for Bitcoin to revolutionize decentralized AI by providing scalability, security, and accessibility. 5: How is your Layer 2 different from others being built? BVM distinguishes itself as a Modular infrastructure that empowers thousands of distinct Bitcoin Layer 2 networks, spanning Gaming, DeFi, Social, and AI applications. We're continuously enriching the BVM Module Store with new modules to enhance its capabilities. With each new module, builders gain access to a wider array of tools to explore different use cases on the Bitcoin network. Recent additions include the Filecoin module for affordable storage and the AI Contracts Library for constructing AI-powered Bitcoin Layer 2 chains. We're also gearing up to release a ZK roll-up module in the coming weeks to offer an alternative to the standard optimistic roll-up. We aim to simplify the process of launching a Bitcoin Layer 2 network customized to specific requirements. Think of it as a SaaS offering with predefined best practices. Whether it's a DeFi Bitcoin Layer 2 or a GameFi Bitcoin Layer 2, we provide default solutions tailored to each use case. We're dedicated to expanding the BVM ecosystem by incentivizing more builders to join the Bitcoin network. Through various programs and grants, we support builders in covering their operational costs for Bitcoin Layer 2. Additionally, we offer rewards akin to 'L2 mining' to those who contribute to expanding the user base and total value locked on the network. In summary, BVM stands out with its modular infrastructure, tailored solutions, and efforts to grow the Bitcoin ecosystem.

Supra

83,548 views • 2 years ago

"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "Second, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)

Rohan Paul

35,485 views • 19 days ago

Experiments in progress. The one on the right has been learning for ~3 hours, the one in the middle for ~1 hour, and the one on the left just started a few minutes ago. The initial motivation for making the physical Atari was just to commit ourselves to a subset of algorithms that can make progress in this setup. This commitment rules out algorithms that require billions of samples to learn (or worse, require multiple environments running in parallel). Atari games are simple enough that we should be able to show learning on them in a short amount of time with no prior knowledge. Since then, I've realized that this setup is also a good way to compare different paradigms in robotics in a principled way. These paradigms are sim2real, learning from tele-operated data, and learning directly on the robots. So far, I have observed that getting sim2real to work reliably is hard. It requires tweaks that don't scale. Policies that can play perfectly in simulation fall apart because of latencies and the messiness of the real world. These aspects could be modeled to improve the simulation, but not without sinking significant human engineering hours. I have higher hopes for learning from tele-operated data, but that requires a human to learn the task first. These experiments are on my to-do list. I have to learn to play some of the games well through the robot. I’m half-decent at playing Pong and Ms Pacman now. Learning directly on robots is looking like the most promising approach. This approach takes away pesky distribution shifts and makes it possible to have algorithms that continually improve with more data and time without any human intervention. It feels great to let experiments run overnight and wake up to find improved policies. With learning on robots, I should, in principle, be able to go on a long vacation and come back to find better policies for complex tasks beyond Atari games. Whether that is possible with current learning algorithms is a different question.

Khurram Javed

52,110 views • 8 months ago

AI's Secret Pattern: The Surprising Role of Fractals in Neural Networks In the realm of artificial intelligence (AI), a groundbreaking discovery has emerged, challenging our conventional understanding of neural network training and optimization. This revelation centers around the identification of fractal patterns at the boundary between trainable and untrainable neural network hyperparameters, presenting a series of profound implications and avenues for further research. Fractals, known for their intricate, self-similar patterns that recur at every scale, have long fascinated mathematicians and scientists alike. Typically associated with simple, one-dimensional iterative functions, the appearance of fractals within the complex, multivariate domain of neural network training introduces a striking contrast. The organic and asymmetric nature of these fractals, as derived from the training processes, suggests a deeper, unexplored connection between the mathematical properties of fractals and the functional dynamics of neural networks. The study’s focus on two-dimensional slices of hyperparameter space barely scratches the surface of the complexity inherent in neural networks, which are characterized by a vast array of hyperparameters. The existence of fractals in this context hints at an underlying high-dimensional structure, a concept that challenges our current capabilities and understanding. Extending fractal analysis to these higher dimensions represents a significant, yet exciting, challenge that could illuminate new aspects of neural network behavior and learning capabilities. An unexpected finding from the research is the persistence of clean fractal patterns even in the presence of stochastic elements introduced during minibatch training. This resilience suggests a parallel to Lyapunov fractals, where the iterative process involves randomly changing functions. This phenomenon prompts a reevaluation of how stochastic and deterministic processes influence fractal formation within neural networks, potentially offering new insights into the fundamental mechanisms of learning and adaptation. From a practical standpoint, the fractal nature of the boundary between trainable and untrainable hyperparameters has significant implications for the field of metalearning. The chaotic behavior of the meta-loss landscape, attributed to its extreme sensitivity, presents a formidable challenge for algorithms designed to optimize hyperparameters. Understanding the fractal characteristics of this landscape could provide valuable guidance for navigating its complexities, ultimately improving the efficiency and effectiveness of metalearning strategies. Beyond the technical and theoretical implications, the discovery also reveals an unexpected aesthetic dimension to neural network fractals. The visual beauty and meditative qualities of these patterns offer a unique opportunity to engage with the material in a deeply personal and contemplative manner. This aspect suggests potential psychological and physiological benefits from exposure to the intricate designs of neural network fractals, opening up novel intersections between technology, art, and well-being. In conclusion, the identification of fractal patterns within neural network hyperparameter spaces unveils a fascinating new frontier at the intersection of fractal geometry and deep learning. This discovery not only challenges existing paradigms but also opens up myriad possibilities for mathematical characterization, algorithmic development, and even subjective exploration. As researchers continue to delve into this rich vein of inquiry, the promise of uncovering new knowledge and advancing our understanding of neural networks and their training processes remains as compelling as ever.

Carlos E. Perez

133,529 views • 2 years ago

📢📢 𝐀𝐯𝐚𝐭𝟑𝐫 📢📢 Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,763 views • 1 year ago