Загрузка видео...

Не удалось загрузить видео

На главную

📁 Yann LeCun says that scaling models will not get us to human intelligence. He explains that the industry remains obsessed with making LLMs bigger, but that this path is fundamentally broken. It does not matter how many parameters we add or how many clusters we build, because the...

158,876 просмотров • 9 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Yann LeCun (Yann LeCun ) beautifully explains how the architecture and principles used to train LLMs can not be extended to teach AI the real-world intelligence. In 1 line: LLMs excel where intelligence equals sequence prediction over symbols. Real-world intelligence requires learned world models, abstraction, causality, and action planning under uncertainty, which current next-token training does not provide. He says current LLMs learn by predicting the next token. That objective works very well when the task itself can be reduced to manipulating discrete symbols and sequences. Math, physics problem solving on paper, and coding fit this pattern because success largely comes from searching and composing the right sequences of symbols, equations, or program tokens. With enough data and scale, these models get very good at that kind of structured sequence prediction. Real-world intelligence is different. The physical world is continuous, noisy, uncertain, and high dimensional. To act in it, a system needs internal models that capture objects, dynamics, causality, constraints from the body, and the outcomes of actions over time. Humans and animals build abstract representations from rich sensory streams, then make predictions in that abstract space, not at the raw pixel level. That is why a child can learn intuitive physics, plan multi-step actions, and adapt quickly in new situations with little data. His claim about saturation follows from this gap. Scaling token prediction keeps improving symbol manipulation tasks like math and code, but it hits limits on embodied reasoning and common sense because text alone does not provide the right learning signals for world models. Predicting the next word cannot efficiently teach contact forces, affordances, occlusion, friction, or how actions change the state of the environment. For that, he argues we need architectures that learn abstractions from sensory data and predict futures in abstract latent spaces, then use those predictions to plan actions toward goals with built-in guardrails. --- From 'Pioneer Works' YT Channel (link in comment)

Rohan Paul

104,460 просмотров • 8 месяцев назад

Yann LeCun just told the most well-funded industry in human history it is solving the wrong problem. LeCun: “Babies learn this around the age of eight or nine months, that objects don’t float, they fall.” No dataset. No labels. No reward signal. A nine month old drops a spoon and builds a physics engine no machine can match. LeCun: “Most of us can learn to drive in about 20 or 30 hours of training without ever crashing, causing any accident.” Twenty hours. Tesla has built the most capable driving system on the road. It took billions of miles of data to get there. A sixteen year old gets there over a long weekend. Not because the teenager is the better driver. Because the teenager is not learning to drive. They are deploying a model of reality they have been building since birth. LeCun: “If we drive next to a cliff, we know that if we turn the wheel to the right, the car is going to run off the cliff and nothing good is going to come out of this.” You simulate the crash. You see the wreckage. You feel the fall. You turn the wheel. None of it was real. All of it was intelligence. Every AI has to crash a thousand times to learn what you imagined once and never did. That is not a performance gap. That is an architecture gap. LeCun: “The main problem we need to solve is how do we learn models of the world.” Not bigger models. Not more compute. Not another trillion tokens. World models. A machine that can run reality forward before it acts. The industry is scaling language. LeCun says language is a compression of thought. Not thought itself. You understood gravity before you could say the word. You grasped cause and effect before your first sentence. The deepest intelligence you will ever possess was built in total silence. And every lab on Earth is trying to reconstruct the mind from words alone. Physics does not care about your context window. A baby who learns that cups fall in a kitchen already knows that rocks fall off cliffs. No retraining. No fine-tuning. One model. Every environment. That is what intelligence actually is. Not prediction. Not pattern matching. Not scale. A simulation of reality so precise you rehearse the future before it exists. Every infant on Earth builds one. No machine ever has.

Dustin

122,149 просмотров • 2 месяцев назад

Without World Models, There Is No AGI. Google Just Proved It. If AGI ever happens, it will not come from bigger chatbots alone. From the very start of this interview, one thing is crystal clear: without world models, we will never reach AGI. And right now, Google is leading with its world simulator Genie 3. Here is the core of what Demis Hassabis explains in this conversation: • World models are the missing core of AGI Hassabis says his deepest long term focus has always been world models and simulations. Not just language. Not just prediction. Actual internal simulations of reality. • LLMs are impressive, but incomplete Language models understand more about the world than expected because human language encodes a lot of reality. Still, language is only a shadow of the real thing. • What text can never fully teach Reality includes things text struggles to express: •3D space and spatial dynamics •Physical causality and mechanics •Sensorimotor experience like movement, force, smell, or balance • Experience beats description To close the gap, AI must learn from interaction and experience, not just static text. That is how you build an internal world simulator. • Why Genie 3 matters With Google DeepMind pushing systems like Genie 3, AI starts to model reality itself, not just talk about it. • Robots and real world assistants depend on this True robotics, smart glasses, and universal assistants require AI that understands the physical world you live in, not just your screen. Bottom line: AGI will not emerge from better text prediction. It will emerge from systems that can simulate, predict, and understand reality itself. Right now, Google is clearly ahead on that path. Curious what you think. Are world models the real AGI unlock, or just another stepping stone?

VraserX e/acc

23,784 просмотров • 8 месяцев назад