Loading video...

Video Failed to Load

Go Home

Introducing PC-ALM, a local-learning alternative to backpropagation. Our method trains 1000-layer neural nets using only local dynamics, and without backprop. Blog: Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit...

432,013 views • 4 days ago •via X (Twitter)

8 Comments

kongtou's profile picture
kongtou4 days ago

So how effective is it?

decipherx's profile picture
decipherx4 days ago

1000 layers is the headline. the interesting bit is the dual variables: each layer becomes a local PI controller, so credit doesn't have to diffuse end-to-end. if that maps efficiently to neuromorphic hardware, backprop stops being the only scalable story

Alex Miller's profile picture
Alex Miller4 days ago

@hardmaru Is this good or bad?

saietta's profile picture
saietta4 days ago

the local-dynamics story only pays off once neuromorphic hardware ships at volume, until then this is training an alternative to backprop for compute that barely exists outside a few research chips

Drew Hawkswood✨'s profile picture
Drew Hawkswood✨4 days ago

Nothingburger explanation = trashed

Rish's profile picture
Rish4 days ago

@MarcoLinSight

Unjuno's profile picture
Unjuno4 days ago

@hardmaru @grok このポストの内容を確認して、 何が革新的で何が問題になりそうなの? わかりやすくレポートをまとめて

Constance Ardiles-Lee's profile picture
Constance Ardiles-Lee4 days ago

That’s awesome ♾️

Related Videos

*New Paper on AI & Democracy* Imagine two approaches to democracy. The one we have today, where citizens choose a professional politician to represent them and others. Or an augmented form of democracy, where each citizen controls a personalized AI that helps them participate in thousands of nuanced decisions. This second approach is the idea of Augmented Democracy I introduced six years ago at TED. In our latest paper we explore a simplified version of Augmented Democracy by combining off-the-shelf LLMs, such as ChatGPT, with data collected using a collaborative government program builder. This was an online game where people build a personalized government program using proposals extracted from the programs of the candidates of the 2022 presidential election in Brazil. So how accurate are these augmented forms of democracy? Imagine a user who gave us 40 answers. We can use the first 20 to fine-tune a model that we can test using the 20 answers the model didn’t see. We can then compare the accuracy of these predictions with the ones obtained by a “bundle” rule, which assumes that users that self-reported to be from the left or right always chose the proposals from the candidate that shares their political identity. This showed us that LLMs were more accurate at predicting policy preferences than the bundle rule, meaning that the preferences captured in the participation data were more nuanced than a left-right axis, and that the LLMs can capture some of that nuance. Also, the LLMs can choose among policies coming from the same candidate, which is something that we cannot do using a bundle rule. But can these LLMs help us complete the aggregate preferences of the population? Direct or unbundled forms of participation can result in incomplete data when people answer only a fraction of all questions. In our paper, we simulate this incompleteness by sampling the full dataset. We ask how close we can get to the full dataset by using a random sample, or a random sample augmented by predictions made by these LLMs. Overall, we find that LLM-augmented data gets much closer to the full dataset than a pure random sample. These results do not mean that augmented democracy technology is ready, but they means we are in a much better place to continue exploring this idea than six years ago. This paper was a collaborative effort with Jairo Gudino, PhD student at CCL at the University of Toulouse Capitole and Umberto Grandi from IRIT also at the University of Toulouse Capitole. We hope you find these results insightful!

César A. Hidalgo

26,915 views • 1 year ago

📢 PERORMANCE V4 IS LIVE We've spent over 10 years at the Top of Performance Improvement companies, earning our place as the world’s #1 E-Sports PC Optimization Specialists. From elite players to top-tier orgs and hardware giants, our mission has always been clear: unlock every ounce of power your PC holds. Today, that mission reaches everyone. Whether you're a competitive gamer or managing high-level operations, tuned performance and low system latency matters. That’s why we’re proud to unveil Performance V4: a completely free utility app crafted and designed by my team and I as the first glimpse into increasing PC Performance for entirely free. Performance V4 is the beginning stages of the upcoming Paragon Tweak Utility (PTU): a revolutionary full-suite optimization platform, soon available through our website, and eventually to the Epic Games Store and Microsoft Store. Our current business strategy has two massive scale issues— the human resources required to optimize each customer's PC, and time to execution with appointment setting & correspondence. We believe that the next step is to create software that replaces that work, and in turn makes PC optimizations more accessible and more common for all, which is why we are launching on Believe. With more access to expendable cash, we can create our vision faster. The Performance V4 is LIVE , alongside an exclusive first look at PTU later next week. If you want to believe in something, believe in us 🫡

Paragon│Boost Gaming PC Performance

55,054 views • 1 year ago

Physicist: Consciousness DOES NOT Come From The BRAIN The prevailing materialist paradigm asserts that consciousness is a byproduct of neural activity, a mere epiphenomenon of biochemical interactions in the brain. However, this reductionist view crumbles under deeper scrutiny, as it fails to account for the vast spectrum of consciousness, from transcendent mystical states to near-death experiences and non-local awareness. Consciousness is not confined within the brain; rather, the brain is a transceiver, a finely tuned instrument that receives and modulates the vast ocean of awareness permeating the cosmos. Just as a radio does not generate the music it plays but instead decodes signals from an unseen field, the brain is an interface between the physical realm and the infinite, omnipresent field of consciousness. Mystic science, in alignment with ancient wisdom and cutting-edge quantum research, reveals that consciousness is fundamental-an organizing principle of reality itself. Walter Russell's work echoes this truth, demonstrating that mind is primary and matter is a consequence of its rhythmic pulsations. The brain, much like a crystalline matrix, is structured to interpret and shape consciousness into coherent experience, but it does not generate it. In this light, consciousness is not local, nor is it constrained by the physical form. It is the unseen architect behind the rhythms of existence, the hidden intelligence orchestrating the grand cosmic symphony to believe that the brain creates consciousness is akin to believing that the eye creates light or that a mirror generates the image it reflects. It is not the origin but the instrument. As mystic scientists, we recognize thay true awakening lies in shifting our perception from brain-centered awareness to the realization that we are conduits of an eternal intelligence, woven into the very fabric of existence. Consciousness is not inside us—we are inside it. ✨🙌🏾💫 © Dr. Jason Yuan

🧬Maxpein🧬

45,371 views • 11 months ago

Check out our latest work, "Actor-Critic Model Predictive Control: Differentiable Optimization meets Reinforcement Learning for Agile Flight," published in the IEEE Transactions on Robotics, where we reconcile #OptimalControl and #ReinforcementLearning, achieving the same super-human performance, but with superior generalizability, as our previous model-free deep RL! Code released! PDF: Code: Full Video: Model-free #ReinforcementLearning (RL) is known for its strong task performance and flexibility in optimizing general reward formulations. On the other hand, #ModelPredictiveControl (MPC) provides robustness, constraint handling, and powerful online replanning capabilities. In this work, we extend our previous AC-MPC paper (Romero, ICRA'24) by taking a deeper look at how both approaches can be unified. We introduce and extend Actor-Critic Model Predictive Control (AC-MPC), a framework that embeds a differentiable MPC inside an Actor-Critic RL architecture. This integration allows the MPC-based actor to perform short-term predictive optimization, while the critic facilitates long-horizon learning and exploration. We conduct a comprehensive study that highlights AC-MPC’s key advantages: - Better out-of-distribution generalization, both against unknown disturbances and changes in the quadrotor dynamics - Improved sample efficiency - A novel empirical analysis uncovering a relationship between the critic’s value function and the MPC cost function, providing deeper insight into their interplay. We validate our method in simulation and the real world on a quadcopter flying at superhuman speeds of up to 21 m/s, matching state-of-the-art model-free RL performance, and retaining the predictive structure of MPC for more reliable out-of-distribution behavior. Reference: Actor-Critic Model Predictive Control: Differentiable Optimization meets Reinforcement Learning for Agile Flight IEEE Transactions on Robotics (T-RO), 2025 PDF: Full Video: Code: Kudos to Ángel Romero, Elie Aljalbout, Yunlong Song! University of Zurich UZH Science UZH Space Hub AUTOASSESS European Research Council (ERC) UZHai

Davide Scaramuzza

27,279 views • 8 months ago

New episode with Dr. Konrad Kording (Kording Lab 🦖), professor of bioengineering and neuroscience at the University of Pennsylvania (Penn) and co-director of CIFAR's Learning in Machines & Brains program (CIFAR). Konrad works at the intersection of causality, machine learning, and neuroscience, building rigorous methods for causal reasoning when experiments aren't possible — and challenging how researchers interpret neural data and build AI. Konrad argues the most promising path to understanding how the brain works is to read the brain’s wiring directly, down to the molecular detail of each connection, and to build compilers and simulations to understand the brain’s computation directly. In this episode we go deep into how neurons work, how neurons wire together, and how organic and artificial neural networks differ. We discuss why organic neurons are doing much more; how a model of a single organic neuron can solve MNIST — computing more like a 3-layer artificial neural network; how the brain might learn by solving credit assignment with only local signals; how to approximate backprop without a global algorithm; why AI and humans are intelligent along different dimensions; why Konrad isn’t very worried about AI replacing us; economic models of intelligence and physical work; and much more. Konrad is a brilliant, contrarian thinker who explains complex concepts very intuitively. It is a solid computational neuroscience primer. I hope you enjoy this conversation as much as I did! Other links to this episode and references below. Chapters 00:00:00 Introduction 00:01:01 How organic neurons work 00:24:13 How the brain learns: circuits and credit assignment 00:45:29 Recording the brain 00:52:47 Why simulating brains is hard 01:05:00 A new approach: connectomes and compilers 01:21:00 Why simulate brains? 01:29:50 How AI and human intelligence differ 01:41:04 Evolution, intelligence and AI risk 01:52:42 Robotics, causality, and the roots of intelligence 02:05:53 AI for science and scientific rigor 02:13:05 The economics of intelligence 02:27:50 A hopeful future

Juan Benet

50,121 views • 2 months ago

New Paper: Continuous Thought Machines 🧠 Neurons in brains use timing and synchronization in the way that they compute, but this is largely ignored in modern neural nets. We believe neural timing is key for the flexibility and adaptability of biological intelligence. We propose a new neural architecture, “Continuous Thought Machines” (CTMs), which is built from the ground up to use neural dynamics as a core representation for intelligence. By using neural dynamics as a first-class representational citizen, CTMs naturally perform adaptive computation. Many emergent, interesting behaviors arise as a result: CTMs solve mazes by observing a raw maze image and producing step-by-step instructions directly from its neural dynamics. When tasked with image recognition, the CTM naturally takes multiple steps to examine different parts of the image before making its decision. This step-by-step approach not only makes its behavior more interpretable but also improves accuracy: the longer it “thinks,” the more accurate its answers become. We also found that this allows the CTM to decide to spend less time thinking on simpler images, thus saving energy. When identifying a gorilla, for example, the CTM’s attention moves from eyes to nose to mouth in a pattern remarkably similar to human visual attention. I think this work underscores an important, yet often lost, synergy between neuroscience and AI. While modern AI is ostensibly brain-inspired, the two fields often operate in surprising isolation. By starting with such inspiration and iteratively following the emergent, interesting behaviors, we developed a model with unexpected capabilities, such as its surprisingly strong calibration in classification tasks, a feature that was not explicitly designed for. When we initially asked, “why do this research?”, we hoped the journey of the CTM would provide compelling answers. By embracing light biological inspiration and pursuing the novel behaviors observed, we have arrived at a model with emergent capabilities that exceeded our initial designs. We are committed to continuing this exploration, borrowing further concepts to discover what new and exciting behaviors will emerge, pushing the boundaries of what AI can achieve.

hardmaru

257,599 views • 1 year ago

Introducing The AI CUDA Engineer: An agentic AI system that automates the production of highly optimized CUDA kernels. The AI CUDA Engineer can produce highly optimized CUDA kernels, reaching 10-100x speedup over common machine learning operations in PyTorch. Our system is also able to produce highly optimized CUDA kernels that are much faster than existing CUDA kernels commonly used in production. We believe that fundamentally, AI systems can and should be as resource-efficient as the human brain, and that the best path to achieve this efficiency is to use AI to make AI more efficient! We are excited to publish our paper, The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization and Composition. We also release a dataset of over 17,000 verified CUDA kernels produced by The AI CUDA Engineer. Paper: Kernel Archive Webpage: HuggingFace Dataset: The AI CUDA Engineer utilizes evolutionary LLM-driven code optimization to autonomously improve the runtime of machine learning operations. Our system is not only able to convert PyTorch code into CUDA kernels, but through the use of evolution, it can also optimize the runtime performance of CUDA kernels, fuse multiple operations, and even discover novel solutions for writing efficient CUDA operations by learning from past innovations! We believe The AI CUDA Engineer opens a new era of AI-driven acceleration of AI and automated inference time optimization. We (Robert Lange, Aaditya Prasad 🇺🇸, sssss, Maxence Faldor, Yujin Tang, hardmaru) are excited to continue Sakana AI's mission of leveraging AI to improve AI.

Sakana AI

1,160,729 views • 1 year ago

A transformer can learn not just the outcomes of dynamics, but the operator that executes the rules. To show this we trained a transformer on roughly 0.04% of a discrete rule space - 100 of 262,144 possible rules - and it learned to apply unseen rules from the same rule class. The model does not simply memorize specific rules. It learns the operator that maps a supplied rule plus an initial state, including unseen rules from this class, to the correct next state. This is relevant because it is a shift from “neural networks approximate dynamics” to “neural networks can learn to execute symbolic programs within a defined rule class”. The rule itself is supplied at inference time, as data, and the network has internalized how rules act, not which rules to apply. On previously unseen rules, the model achieves 98.5% perfect one-step forecasts and reconstructs governing rules with up to 96% functional accuracy. Two results make this hold up under scrutiny. First, inductive bias decay. As we scaled training rule diversity, the correlation between functional inference accuracy and distance-from-nearest-training-rule collapsed to R² = 0.00. At the largest tested training-rule diversity, the model’s performance on a new rule shows no measurable dependence on how similar that rule is to anything it was trained on. The bias toward training data (the thing we worry most about in compositional generalization claims) is something we can measure decaying, and we find that at scale it is gone. Second, an identifiability theory. We derive a closed-form expression for the number of rules consistent with a single observation. This reframes the inverse problem: failure to recover ground truth is not necessarily a model defect, but can be correct behavior when the data underdetermine the rule. The model is sampling the equivalence class; and identifiability is governed by coverage, not capacity. The methodological move underneath both results is amortization. Classical work on rule inference (e.g. the Santa Fe EVCA program, evolutionary search over CA rule space) was per-instance: search the rule space for each new system. We replace that with a single forward pass of a transformer trained across many instantiations of the rule class. That is what makes symbolic rule inference scalable as a research direction rather than a curiosity. We show that this works in a tightly constrained domain: binary, deterministic, local cellular automata on small grids. The locality-break experiment shows the model fails sharply when target systems violate its structural priors (which is itself a useful diagnostic, but it bounds the operator class). We don't yet know how this scales to multistate, higher-dimensional, or stochastic CA, or whether it transfers cleanly to non-CA systems whose coarse-grained dynamics admit local surrogates. The identifiability framework - what can be inferred from observation, given a hypothesis class - should transfer wherever finite local rules meet sparse data. The amortization argument transfers wherever per-instance symbolic search has been the bottleneck. Those are the pieces I expect to outlive the cellular automata setting. Led by Jaime Berkovich with Noah David, at LAMM@MIT. Out now in Advanced Science Advanced Portfolio (link to paper & code below).

Markus J. Buehler

39,019 views • 4 months ago

Elon Musk: We had to print out 25,000 pages of paper to start Tesla's factory in Germany. Alice Weidel: “We need to free our firms, our companies, and the individuals of these obnoxious bureaucratic conditions here [in Germany]. Do you know how long it takes, how many days it takes to get a business permit in Germany?” Elon Musk: “As it turns out, I do, because we built a gigantic car factory just near Berlin. We had many, many challenges. To be clear, we actually had a lot of support, a lot of local support, a lot of local support from the local government, from the national government. And despite all that support, just the sheer number of rules that the people in the government are required to follow is completely crazy. I think our permit was 25,000 pages, and it had to be all printed on paper. I think maybe more than that in the end. There has to be many, many copies made. It literally was a truck of paper. We were like, surely we can make this electronic? Isn't that better for everyone? And they said, no, it has to be paper. This is crazy. This was only a few years ago. It's not the distant past. We're a quarter of the way through the 21st century. It's like, guys, 25,000 actual printed pages ... I believe every page needed to be stamped with a physical stamp. Honestly, it's going to really tire somebody out to do so much stamping. They're going to get some sort of repetitive stress injury. They will end up in the hospital, I mean, that's too much. But I'm not trying to blame the individuals who are doing this because they are just following the rules. So, we have to change the rules.” Conversation with Alice Weidel, January 9, 2025

ELON CLIPS

5,576,525 views • 1 year ago