Video wird geladen...
Video konnte nicht geladen werden
🚨 [New Paper] The Adam optimizer is a zombie algorithm... It senses and adapts the learning rate, sure. But the update rule itself? Fixed, frozen. Decided before even the training starts. It works in some regions of the loss landscape and fails in others. What if the optimizer itself... show more
18,589 Aufrufe • vor 4 Monaten •via X (Twitter)
16 Kommentare

PILOT receives a signal at every step: Does the current gradient agree with my past ones? It's basically the cosine similarity between consecutive gradients, smoothed into a running agreement score we call 'rho'. If rho is high, then gradients agree, and the landscape is smooth, PILOT should go aggressive. If rho is low or negative, then the landscape is noisy and chaotic; PILOT should be cautious.

Since rho is our signal, it gets fed into three learned polynomials that learn three control knobs that define PILOT's policy: 1- p_m: How much should we trust the momentum against the new raw gradient? 2- p_v: How much should we normalize the second moment, mimic Adam? skip normalization? or something in between? 3- p_s: Should we use the full magnitude? or compress to pure sign updates Each of these values is learned with weights, biases, and sigmoid activation. If degree 'd' is 2 (quadratic polynomial), then the entire policy is learned by 9 parameters (3 values * 2 coefficients + 3 biases)

The policy reshapes the update rule at every step... For example, since p_m controls trust in momentum vs raw gradient, then the blended direction 'n' at step 't' is: n_t = p_m * momentum_t + (1 - p_m) * gradient_t The rest of the update rule is shown in the equation attached. Notice how if we set p_m = 1, p_v = 0.5, and p_s = 0, then we recover Adam's update rule exactly! But if we set p_v = 0, and p_s = 1, then it's a pure sign update like Lion (which outperforms Adam in many cases) PILOT interpolates through many known and novel update rules freely, and online during the model training, it learns and updates its policy with the same loss as your model.

We trained ResNet-18 and a small CNN on benchmark datasets, CIFAR-10 and FashionMNIST. We compared the training of various optimizers that are closest to PILOT (Adam and AdamW, Lion, Sophia, and AdaBelief) for each combination. PILOT outperformed all the other optimizers in terms of accuracy, Loss, and on CIFAR-10, lowest Loss Variance (training stability). Please refer to the paper for the full benchmarking report and discussion.

Here, PILOT follows a distinct trajectory through the loss landscape and converges to a lower-loss region compared to Adam, AdamW, Lion, and Sophia.

The wildest finding, though, is that PILOT's learned policy can generalize over unseen datasets! We trained PILOT on CIFAR-10, froze the policy, and dropped it on FashionMNIST with a fresh model. The frozen PILOT beat AdamW on Accuracy and F1-Score. Also, it converged faster, reaching 90% in only 4 epochs vs 7 epochs of AdamW! The frozen PILOT (pretrained on CIFAR-10) also showed competitive performance with that of a new, fresh PILOT variant trained online on FashionMNIST. This shows that the policy encodes broadly applicable optimization dynamics rather than dataset-specific patterns.

Huge thanks to my co-author @Lama_s1, and to supervisors Dr. Muhammad Mubashar & @ProfNaeemKhan 🫡 Special thanks to @KAUST_Academy & @Dr_S_Albarakati for their support. For the full paper, installation, & code you can find it all here: 🔗:

Wallahi elegant approach using the directional agreement (cosine similarity of consecutive gradients) as a proxy for landscape stability is a clever way to bypass expensive second-order or Hessian approximations I do have a question regarding the polynomial policy… did you notice sensitivity or stability issues when scaling the degree d past 3 - 4?

Thanks 🙏 And for your question, from our runs and experiments the polynomial degree did not cause any stability or sensitivity issues. The deviation is consistent across degrees 1-15 (std 0.05-0.07) We found setting degree to 1-2 would yield slightly better results.

Okayyy staying that stable all the way up to degree 15 is good . It really proves how robust the formulation is against overfitting. Thanks for the solid explanation Sattam great work on PILOT

قوووووة 🤍🤍

Super impressive

Thanks Jude!

Im so amazed by the approach you did its just outstanding 👏👏

Thank you Abdullah, im glad

Goat 🐐
