Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🚨 [New Paper] The Adam optimizer is a zombie algorithm... It senses and adapts the learning rate, sure. But the update rule itself? Fixed, frozen. Decided before even the training starts. It works in some regions of the loss landscape and fails in others. What if the optimizer itself...

18,589 Aufrufe • vor 4 Monaten •via X (Twitter)

16 Kommentare

Profilbild von Sattam
Sattamvor 4 Monaten

PILOT receives a signal at every step: Does the current gradient agree with my past ones? It's basically the cosine similarity between consecutive gradients, smoothed into a running agreement score we call 'rho'. If rho is high, then gradients agree, and the landscape is smooth, PILOT should go aggressive. If rho is low or negative, then the landscape is noisy and chaotic; PILOT should be cautious.

Profilbild von Sattam
Sattamvor 4 Monaten

Since rho is our signal, it gets fed into three learned polynomials that learn three control knobs that define PILOT's policy: 1- p_m: How much should we trust the momentum against the new raw gradient? 2- p_v: How much should we normalize the second moment, mimic Adam? skip normalization? or something in between? 3- p_s: Should we use the full magnitude? or compress to pure sign updates Each of these values is learned with weights, biases, and sigmoid activation. If degree 'd' is 2 (quadratic polynomial), then the entire policy is learned by 9 parameters (3 values * 2 coefficients + 3 biases)

Profilbild von Sattam
Sattamvor 4 Monaten

The policy reshapes the update rule at every step... For example, since p_m controls trust in momentum vs raw gradient, then the blended direction 'n' at step 't' is: n_t = p_m * momentum_t + (1 - p_m) * gradient_t The rest of the update rule is shown in the equation attached. Notice how if we set p_m = 1, p_v = 0.5, and p_s = 0, then we recover Adam's update rule exactly! But if we set p_v = 0, and p_s = 1, then it's a pure sign update like Lion (which outperforms Adam in many cases) PILOT interpolates through many known and novel update rules freely, and online during the model training, it learns and updates its policy with the same loss as your model.

Profilbild von Sattam
Sattamvor 4 Monaten

We trained ResNet-18 and a small CNN on benchmark datasets, CIFAR-10 and FashionMNIST. We compared the training of various optimizers that are closest to PILOT (Adam and AdamW, Lion, Sophia, and AdaBelief) for each combination. PILOT outperformed all the other optimizers in terms of accuracy, Loss, and on CIFAR-10, lowest Loss Variance (training stability). Please refer to the paper for the full benchmarking report and discussion.

Profilbild von Sattam
Sattamvor 4 Monaten

Here, PILOT follows a distinct trajectory through the loss landscape and converges to a lower-loss region compared to Adam, AdamW, Lion, and Sophia.

Profilbild von Sattam
Sattamvor 4 Monaten

The wildest finding, though, is that PILOT's learned policy can generalize over unseen datasets! We trained PILOT on CIFAR-10, froze the policy, and dropped it on FashionMNIST with a fresh model. The frozen PILOT beat AdamW on Accuracy and F1-Score. Also, it converged faster, reaching 90% in only 4 epochs vs 7 epochs of AdamW! The frozen PILOT (pretrained on CIFAR-10) also showed competitive performance with that of a new, fresh PILOT variant trained online on FashionMNIST. This shows that the policy encodes broadly applicable optimization dynamics rather than dataset-specific patterns.

Profilbild von Sattam
Sattamvor 4 Monaten

Huge thanks to my co-author @Lama_s1, and to supervisors Dr. Muhammad Mubashar & @ProfNaeemKhan 🫡 Special thanks to @KAUST_Academy & @Dr_S_Albarakati for their support. For the full paper, installation, & code you can find it all here: 🔗:

Profilbild von F H
F Hvor 4 Monaten

Wallahi elegant approach using the directional agreement (cosine similarity of consecutive gradients) as a proxy for landscape stability is a clever way to bypass expensive second-order or Hessian approximations I do have a question regarding the polynomial policy… did you notice sensitivity or stability issues when scaling the degree d past 3 - 4?

Profilbild von Sattam
Sattamvor 4 Monaten

Thanks 🙏 And for your question, from our runs and experiments the polynomial degree did not cause any stability or sensitivity issues. The deviation is consistent across degrees 1-15 (std 0.05-0.07) We found setting degree to 1-2 would yield slightly better results.

Profilbild von F H
F Hvor 4 Monaten

Okayyy staying that stable all the way up to degree 15 is good . It really proves how robust the formulation is against overfitting. Thanks for the solid explanation Sattam great work on PILOT

Profilbild von Ice ❄️ | مُوسَى ابْنُ رَاشِد
Ice ❄️ | مُوسَى ابْنُ رَاشِدvor 4 Monaten

قوووووة 🤍🤍

Profilbild von Jude Al Jafsher Alqhtani
Jude Al Jafsher Alqhtanivor 4 Monaten

Super impressive

Profilbild von Sattam
Sattamvor 4 Monaten

Thanks Jude!

Profilbild von Abdullah Albadri
Abdullah Albadrivor 4 Monaten

Im so amazed by the approach you did its just outstanding 👏👏

Profilbild von Sattam
Sattamvor 4 Monaten

Thank you Abdullah, im glad

Profilbild von Coach Retx43
Coach Retx43vor 4 Monaten

Goat 🐐

Ähnliche Videos

What happened to Jeffrey Sachs is not an argument. It is a ritual, the auto-da-fé of a dying civilization. Europe no longer debates. It excommunicates. Every time truth crosses its borders, it is denounced as heresy, burned in the square of public opinion, and buried under the flags of "values" it no longer lives by. The Italian senator who called Professor Sachs a liar wasn’t defending Ukraine. He was defending the psychological architecture of European dependence. He cannot admit the truth because his career, his ideology, and his identity all collapse if he does. Europe is no longer a continent. It is a colony that thinks itself free. Washington writes the script, Brussels recites it, and the people pay for the performance in cold homes and silent factories. They call it "solidarity." But solidarity with your own jailer is not virtue. It is pathology. Europe kneels before America and mistakes the floor for high ground. It sanctions Russia and bankrupts itself. It sacrifices its own citizens to fund a war it cannot win. It destroys its own energy, its own diplomacy, its own industry, all to prove its loyalty to a master who despises it. Jeffrey Sachs did not embarrass Europe. He revealed it. A continent that once produced Beethoven, Goethe, and Marx now worships at the altar of CNN and NATO press releases. It has traded reason for narrative and memory for submission. The tragedy of Europe is not that it was conquered. It is that it volunteered. It begged for occupation, and now calls vassalage "values." When Professor Sachs spoke, the Italian senator did not hear an argument. He heard a mirror, and mirrors terrify those who live by illusions. Europe isn’t being silenced by America. It is silencing itself out of fear of remembering what it used to be. A civilization that once claimed to civilize the world can no longer govern itself. It outsourced its sovereignty, privatized its conscience, and mortgaged its dignity for access to Washington’s approval. The slave masters of history have become servants in suits, begging their overseer for scraps of relevance. Europe is no longer a continent. It is an accent in America’s voice. And even that accent is fading.

Sony Thăng

363,325 Aufrufe • vor 11 Monaten

The robot flipped a pancake nobody taught it! 🥞 Skild AI team assumed pancake flipping had to be somewhere in the training data. So they searched. Millions of hours of pre-training data. Nothing. S1 inferred the whole task from a single human demonstration. That's their new general robot model, built as an in-context learner from the ground up. Every new robot task today starts with days of teleoperation and a fine-tuning run on a specialist policy. S1 skips all of it. Much like a language model, it never updates its weights to learn a new task. The demonstration enters the context window, and the policy uses it to decide what to do next. → Ten-minute tasks it was never trained on, composed from primitives learned in pre-training: a new style of coffee, potting a plant, frying pancakes. → Soil and pots arrived at their office at 8:54 PM. The robot was running the task autonomously by 9:27 PM. → Slide objects away mid-reach, swap them, change the lighting, it still finishes. → The prompt waters a plant with a watering can, but only a cup is available. It uses the cup. It doesn't rigidly replay what it saw, but it recovers from its own errors, and sometimes executes with more precision than the demonstrator, when the human fumbles an egg and makes a mess, S1 performs the same step cleanly. The demonstration is a specification of the goal, and not a trajectory to copy. On unseen tasks after 100K hours of pre-training: language-prompted VLAs reach 9%. Their new model reaches 66%. It's already deploying with industrial partners, with a wider rollout over the coming months. Congrats Deepak Pathak and team behind this! 😮‍💨 🔗 Link to their latest blog: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

11,406 Aufrufe • vor 1 Monat

Former Meta Chief AI Scientist Yann LeCun on the three paradigms of machine learning — and why the third is what made ChatGPT possible: Here's each one, and where it breaks. First, supervised learning. You tell the machine the answer. "You show it a picture, let's say of a table, and you tell it this is a table. So it's supervised because you tell it what the correct answer is." Get it wrong, and the machine rewrites itself: "The system computes its output, and if it says something else than table, then it's going to adjust its parameters, its internal structure, so that the output it produces gets closer to the output you want." Repeat at scale and something more than memorisation appears: "Eventually the system will find a way to recognize every image you trained it on, but also images it's never seen that are similar to the one you train it on. This is called a generalization ability." The limit: a human has to supply every single answer. That doesn't scale to the size of the internet. Second, reinforcement learning. You don't give the answer, only a verdict. "You don't tell the system what the correct answer is. You only tell it whether the answer it produced was good or bad." Learning to ride a bike, essentially: "You try to ride a bike and you don't know how to ride the bike and after a while you fall. So you know you did something bad and so you change your strategy a little bit. And eventually you learn how to ride a bike." For years the field assumed this was the closest thing to how animals actually learn. Yann LeCun's verdict: "Now it turns out reinforcement learning is extremely inefficient." It dominates wherever failure is free: "It works really well if you want to train a system to play chess or play go or poker, because you can have the system play millions and millions of games against itself and basically fine-tune itself. But it doesn't really work in the real world." The limit, in one image: "If you want to train a car to drive itself, you're not going to do it with reinforcement learning. It's going to crash thousands of times." On robotics he's careful rather than dismissive: "Reinforcement learning can be part of the solution, but it's not the complete answer. It's not sufficient." Third, self-supervised learning. You tell the machine nothing at all. "And this is what has enabled the recent progress in natural language understanding and chatbots." The strange part is that you stop asking for a task: "You don't train the system to accomplish any particular task. You just train it to basically capture the structure..." The method is deliberate sabotage: "You take a piece of text, you corrupt it in some way, by for example removing some words, and then you train a big neural net to predict the words that are missing." And one narrow version of that trick runs every chatbot on Earth: "A special case of this is that you take a piece of text and the last word in that text is not visible, and so you train the system to predict the last word in that text — and this is the way large language models are trained on." So why did the third one win? Supervised learning needs a human. Reinforcement learning needs a crash. Self-supervised learning needs neither — because the missing word and the correct answer are the same thing. The data grades itself.

Big Brain AI

49,293 Aufrufe • vor 1 Monat

Introducing Headlong, an open source microharness for persistent agents: self-guided agents that think continuously. Most agent harnesses are reactive: you send a task, the agent completes it, and then it sits frozen until the next request. Cron jobs and heartbeats wake it up to run a checklist and put it back to sleep. A Headlong agent is never asleep. It keeps generating thoughts about whatever it decides is interesting, in a self-guided loop inspired by human inner monologue. Your message doesn't start a session. It's one more observation that lands in the agent's thought stream, and the agent decides if and when to reply. Headlong is built on the idea of persistent agency: continuous inner thought generation between external interactions. The agent sets its own interests and priorities, comes up with its own projects, and sometimes pings you unprompted with progress. To keep our prototype as simple and small as possible, we implemented Headlong as a microharness: a complete agent harness in under 10K lines of Bash, organized as a handful of small executables. It includes a loop that generates the next thought, shellm (a recursive language model written in Bash), a trajectory stored as a DAG of jsonl files, and context as a projection of that trajectory. We've been running one Headlong agent internally at Laude for several weeks. The whole team talks to it over Slack and Telegram, and every conversation lands in its single stream of thought. It works in its own fork of Headlong and we've pulled over 50 of its commits into main. One night, with nobody talking to it, it went back to check whether a recall process it had built was actually wired into its mind, found that it wasn't, diagnosed and fixed the bug, and verified the fix end to end. 48 minutes, no human asked for the fix or was in the loop at any point. Every step is a timestamped line in its log. Things broke too, and we wrote those up. Background thinking costs us $1 to $2 an hour, our agent stopped its own service three times by accident, and self-delegation died on day one. Details in the post. One line installs everything and starts an agent. Use a dedicated sandbox and spend-capped API key; it runs real shell commands and thinks around the clock. Headlong is research software, be careful! curl -fsSL | bash Launch post: Repo: Headlong is a Laude Institute / MIT collaboration.

Andy Konwinski

360,844 Aufrufe • vor 1 Monat