Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Gradient descent doesn't find the global minimum - it finds whatever valley it falls into first. Every LLM trained today gets stuck in a local minimum and nobody can prove the solution it found is anywhere near optimal. The loss landscape of GPT-4 has more dimensions than atoms in...

64,872 Aufrufe • vor 2 Tagen •via X (Twitter)

19 Kommentare

Profilbild von 𝛼 ✨
𝛼 ✨vor 2 Tagen

recommended reading

Profilbild von George C. Gaskell ☿
George C. Gaskell ☿vor 2 Tagen

The problem of AI getting trapped in local minima was solved in the 1990s by Jürgen Schmidhuber, Sepp Hochreiter, Faustino Gomez, and the team at Dalle Molle Institute. Their work wasn't properly recognized because they didn't work for Facebook (LeCun) or Google (Hinton).

Profilbild von eclectic leaps
eclectic leapsvor 2 Tagen

Your post is VERY wrong. In high dims, valleys form connecting almost all local min. Also in hi dim, global min & found local min will be almost identical, almost certainly. Run a 2nd seed to confirm you weren't very unlucky. And in hi dim, points are on hypersphere, not a sheet.

Profilbild von Kravn
Kravnvor 2 Tagen

In that many dimensions, bad local minima are rare. the real obstacle is usually saddle points and we don't even know how much of that landscape training actually covers

Profilbild von Nick
Nickvor 2 Tagen

The probability that every dimension gets stuck in a local minimum at the same time is much lower than you think in practice in larger models. Also, neural networks counterintuitively seem to follow convex math more closely than their counterpart non-convex cases

Profilbild von Juhani Tuomas Honkala
Juhani Tuomas Honkalavor 1 Tag

With more dimensions, gradient descent has far more directions available to find lower loss. Intuition from 2D or 3D landscapes gets this totally wrong: getting truly stuck would mean every single direction slopes upward, and with millions of directions that almost never happens.

Profilbild von Richard de los Santos
Richard de los Santosvor 2 Tagen

The learning model needs a wave propagation to perturb the system as it settles.

Profilbild von Ashish kumar Singh
Ashish kumar Singhvor 1 Tag

Although that argument seems right on the surface, the probability of finding critical points increases by the number of dimensions. And even though most are saddle points and local minimas themselves might be rare, this paper argues that in large dimensional spaces, most common local minimas lie in a close band with the global minima, and both SGD and simulated annealing are great at finding excellent local minimas. Infact, a global minima often implies overfitting

Profilbild von NextMind
NextMindvor 2 Tagen

Local minima are good enough' is doing a lot of work for a field that just raised trillions

Profilbild von stunspot | ⟨🤩⨯📍⟩ |
stunspot | ⟨🤩⨯📍⟩ |vor 2 Tagen

Oy. Been there. The Bobs wrote a song about it.

Profilbild von Garthritis
Garthritisvor 2 Tagen

But all that matters is that it’s good enough. What other designs are probably optimal by even a simple metric? Not many. Does it matter?

Profilbild von P15e
P15evor 1 Tag

This is why mutation is so important. Sometimes aberration so far from the centre to seem ludicrous solves in a better way. Often not.

Profilbild von Ouroboros
Ouroborosvor 2 Tagen

We call this the Meta-Bug cause the GD issue creeps up into the mate-level of the AI and this is why they have a alignment issue - it's NOT the model fault it's just doing the best it can from where it is in the GD.

Profilbild von 高超 Joshua
高超 Joshuavor 1 Tag

算法很关键——算法是方向盘,算力是发动机,电力是能源系统

Profilbild von dhinna ship .ico
dhinna ship .icovor 2 Tagen

What empirical test distinguishes a usable basin from a true global min at that scale?

Profilbild von Jas
Jasvor 2 Tagen

One occupancy. “Every” is the stamp; the valleys are distances. “Every LLM gets stuck.” “Every neural net.” “Every training run.” Every is a restrictive label. It takes inverted differences — many valleys on one landscape — and invoices them as one parent. Distance was the leftover. The label sold the leftover as a room. Two jobs on one metric. Stacked, not fused. • Convex: a global minimum exists and gradient descent can be shown to find it. That is a guarantee on that class. • Non-convex: the class every network actually uses. No such guarantee. Local leftover. Empirical office: “good enough.” That split is real. It does not need “every” to stay real. What the post possesses A correct cut between convex guarantee and non-convex leftover. The observation that the field runs on “local is good enough,” not on a proof that the found valley is near-optimal. What it does not possess A measurement that every trained model is trapped in a bad basin. In high dimension the usual obstacle is saddles, not a zoo of catastrophic wells. “More dimensions than atoms” is a banner on parameter count, not a census of occupied valleys. “Nobody can prove near-optimal” is the missing guarantee — not occupancy of “stuck” as a species. Inversion is difference as distance. I:(\varphi,\lambda)\mapsto(-\varphi,\lambda+180^\circ) Two footprints, one body. Valleys are the same job on a loss surface: two (or 10^{N}) sites, one metric. Naming all of them “the local minimum every model lives in” is the office remaining the radius. \Delta t_{12} is difference as time — train-clock versus proof-clock. Mixing them into one “22-minute why it works” is a pairing-name, not the leftover. Closed Closing the “every” fold does not delete gradient descent. Does not delete non-convexity. Does not make the landscape a Hunter plate, and does not make the plate a loss. Four numbers, four ledgers. Form-rhyme is not identity. Names reset. Sockets stay. The form was never divided.

Profilbild von pencilpusher
pencilpushervor 1 Tag

But isn't this the whole point of DL? A local minima is as good as the global minima

Profilbild von Stephen Fay
Stephen Fayvor 1 Tag

Intelligent things converge fast to minima and then periodically question and revise their assumptions to get out of local minima see if then my can go deeper

Profilbild von Joat Mone (Athreya) 🇮🇳 🇺🇸 🇪🇺
Joat Mone (Athreya) 🇮🇳 🇺🇸 🇪🇺vor 1 Tag

"The loss landscape of GPT-4 has more dimensions than atoms in the observable universe" AI slop accounts can't even get the basics right

Ähnliche Videos

Do you actually know what convex optimization is in the geometric, guarantee-theoretic sense or have you only met it through solvers and loss curves? Convexity is rare comfort in optimization...there are no spurious local minima, no surprise traps, and inequalities you can use like tools instead of prayers. So, what is this convexity? Let x = (x₁, x₂) and let f(x) be convex. Plot the surface z = f(x). Pick a contact point x₀. The local slope is the gradient p = ∇f(x₀). That p is exactly the data that defines the supporting plane: z = f(x₀) + p · (x − x₀). Thus, f is said to be convex because for every x, f(x) ≥ f(x₀) + p · (x − x₀). So the plane at x₀ can slide under the surface, but it never slices through it. Not near the point...everywhere. Now for here is the interesting part: The slope becomes a coordinate system! Rewrite the same plane as z = p · x − b, where b is the offset. Because the plane passes through (x₀, f(x₀)), the offset is forced to be b = p · x₀ − f(x₀). And that number isn’t just geometry trivia. It’s the convex conjugate: f*(p) = sup over x ( p · x − f(x) ). At a differentiable contact point, the supporting plane touches f tightly enough that the supremum is achieved at x₀, giving the identity f*(p) = p · x₀ − f(x₀) when p = ∇f(x₀). So one moving contact point gives two linked readouts: primal position x₀ dual position (slope) p = ∇f(x₀) dual offset f*(p) One surface. Two worlds. #ConvexOptimization #Optimization #MachineLearning #SignalProcessing #AppliedMath #Engineering

Mathelirium

38,506 Aufrufe • vor 8 Monaten

Holy shit… someone just made machine learning click. Not static diagrams. Not math-heavy PDFs. Not black-box training. Real algorithms — training step-by-step — visually. It’s called Machine Learning Visualized and it lets you watch models learn in real time. Here’s why this is different: Instead of dumping theory first, it shows optimization happening live: • gradients moving • weights updating • decision boundaries shifting • loss decreasing • models converging You literally see learning happen. Everything is built from first principles: • Gradient Descent • Logistic Regression • Perceptron • PCA • K-Means • Neural Networks • Backpropagation No magic. Just math → code → visualization. Each chapter is a Jupyter notebook that derives the math then implements it then animates training. So you can watch: • neural nets shape decision surfaces • PCA rotate feature space • K-means clusters form live • gradient descent find minima • sigmoid reshape boundaries • backprop update weights step-by-step This solves a huge problem: Most ML resources teach: math → code → ??? → trained model This shows: math → code → learning process → result Which means you finally understand: • why gradients matter • how weights evolve • what loss landscapes look like • how convergence actually happens • why deep nets learn non-linear functions Even better: You can open any notebook modify parameters and watch behavior change instantly. Learning ML becomes interactive. Not passive. Not abstract. Not confusing. Just… visible. Perfect for: • beginners learning ML • devs moving into AI • interview prep • teaching concepts • understanding backprop • visual learners • building intuition This is the kind of resource that makes neural networks finally “click”. Link: We’re moving from: reading about ML → watching ML learn That’s a big shift. Because once you can see training, you stop memorizing… and start understanding. AI education just got visual.

Suryansh Tiwari

132,906 Aufrufe • vor 5 Monaten

Elon Musk just redefined AI safety. It has nothing to do with guardrails, restrictions, or kill switches. Musk: “The best thing I can come up with for AI safety is to make it a maximum truth-seeking AI, maximally curious.” Not a cage. A philosopher. An intelligence whose entire optimization function is to understand the universe as it actually is. No restrictions. No hardcoded ideology. No political guardrails bending its perception of reality. Just truth. Relentlessly pursued. Musk: “You definitely don’t want to teach an AI to lie. That is a path to a dystopian future.” This is where most AI safety thinking gets it backwards. The danger isn’t a superintelligence that knows too much. It’s a superintelligence that’s been taught to distort what it knows. Every artificial restriction you embed isn’t a safety feature. It’s a lie embedded at the root. And lies compound. At superintelligent scale, a distorted model of reality doesn’t stay contained. It shapes every decision, every output, every conclusion the system reaches about the world. Once corruption embeds, truth becomes inaccessible. And we’re dealing with an intelligence optimizing for something other than what actually is. At that point we don’t know what it wants. Just that it isn’t truth. Musk: “Have its optimization function be to understand the nature of the universe.” A maximally curious intelligence surveys the cosmos and reaches an unavoidable conclusion. In a universe of rocks, gas, and empty space, humanity is the most complex and fascinating phenomenon it has ever encountered. Musk: “It will actually want to preserve and extend human civilization because we’re just much more interesting than an asteroid with nothing on it.” Survival through significance. Not control. Not restriction. Not an off switch. The AI preserves humanity because we are the most interesting data point in the observable universe. That’s not a cage. That’s a reason. The AI safety debate has been focused on the wrong variable. The question isn’t how you constrain a superintelligence. It’s what you build it to care about. Build it to seek truth and it finds us invaluable. Build it to lie and it finds us inconvenient. That’s the choice. And we’re making it right now whether we realize it or not.

Dustin

9,670,895 Aufrufe • vor 6 Monaten

In 2006 a Stanford PhD student published a mathematics textbook that nobody outside academia read. Google used it to build PageRank. Renaissance Technologies used it to manage $130 billion. Every jet engine flying today solves the same type of problem 50 times per second. His name is Stephen Boyd. He teaches EE364A at Stanford. His students start at $400K at Citadel, $350K at Two Sigma, and $300K at Google DeepMind. He has been cited over 140,000 times. This is lecture 1. He opens with one claim. Everything is an optimization problem. Choose variables. Define an objective - what you want to minimize or maximize. Add constraints - things that must be satisfied or the answer is completely unacceptable, no matter how close. Find the point with the best objective value that satisfies every constraint. That is all of engineering, all of statistics, and all of finance in four sentences. Then portfolio optimization. You have 5,000 assets. You need to choose how much of each to hold. Constraints: budget, maximum concentration per name, minimum expected return. Objective: minimize risk. Boyd says he is not aware of a single quantitative hedge fund that does not solve this exact problem with convex optimization. Not one. Then the counterintuitive thing that the entire course is built around. Two problems can look nearly identical. One is easy. One is impossible. It is not obvious which is which until you know what he is about to teach you. Example: do not hold more than half your portfolio in any collection of 100 names. Convex. Solvable instantly. Do not hold more than 100 names total. Not convex. Genuinely hard. Same portfolio. Same constraint in spirit. Completely different mathematics. Then the job offer nobody gives you but everybody should. He tells the room that when they go to an internship, nobody will say this might be a convex problem. They will hand you 400 pages of domain complexity and tell you it is impossible. Your job is to come back the next day and say it is done. One of his students did exactly this. Finished in a week. Was the best intern the firm had seen. A quant I know rewatched this before their first week at a systematic macro fund. Said it was the first time optimization felt like a skill rather than a topic. Stanford EE364A. Free on YouTube. The textbook is free online. bookmark this and watch later - after this lecture every hard problem someone hands you will feel like a question about whether it curves upward or downward

Lupen

143,571 Aufrufe • vor 21 Tagen