What Yann LeCun is prophesying
Yann LeCun gave a talk recently that I keep turning over in my head, and I am no expert on world models so please take this with a bag of salt, but the core argument is simple enough to fit in one line which is that machine learning sucks at learning. He means it the way an insider means it, because he has spent decades inside the field, and his complaint is that a ten year old can do what no robot can, which is walk into a kitchen she has never seen and make herself useful on the first try without any training at all.
This is the old Moravec paradox wearing new clothes, which says the things that feel easy to us are brutally hard for machines while the things that feel hard to us are easy for them. A computer beats grandmasters at chess and proves math theorems and integrates symbolically, and the same computer cannot learn to drive a car in the few hours it takes any teenager, despite self-driving companies feeding it literally millions of hours of training data. I have been there with this one because I trained openpilot in the Himalayas years ago on roads that had no lane markings at all, and Yann is right that the industry is still stuck at level two or three for consumer cars, with the robo-taxis only working because they are engineered within an inch of their lives.
His explanation for why babies beat billion-dollar models starts with how fast infants build a model of the world just by watching it. A two month old figures out the world is three dimensional because distance is the best explanation for how the view shifts when the head moves, and object permanence and rigidity come quickly after, and the heavier stuff like gravity and inertia takes about nine months, which is why an eight month old on a high chair keeps throwing every toy on the floor and watching what happens. The kid is running the experiment, and Yann’s point is that the learning happens mostly by passive observation, not by being rewarded or punished.
Then he does the arithmetic that the whole talk is really about, and the numbers are worth saying out loud. A modern LLM trains on something like twenty trillion words which is roughly ten to the fourteenth bytes, and that is about four hundred thousand years of reading for any human. A four year old, awake sixteen hours a day with two million optic nerve fibers carrying about a byte per second each, has seen roughly the same ten to the fourteenth bytes through vision alone. So the entire text record of humanity equals what a toddler absorbs just by looking around, and Yann’s conclusion is that we will never get anything like human intelligence from text alone because the data volume and the data type are both wrong. Video looks wasteful next to text but the redundancy is a feature, since self-supervised learning needs redundancy the way a student needs repetition.
What he wants instead of bigger LLMs is a different way of thinking entirely, and this is the part I had to read twice. An LLM produces each token by pushing it through a fixed number of layers, the same computation every time, and the only way to make it reason is to trick it into generating more tokens, which is not how anyone reasons. We reason inside our heads, without words, by imagining actions and checking whether they work. His architecture does the same thing with what he calls a world model, which predicts what happens if you take an action, plus an objective that measures whether the task got done, and the system searches for the action sequence that minimizes that objective at inference time. In plain English, instead of blurting the first answer the network computes, it imagines several futures and picks the one that works, the way you plan a trip to Thailand by first deciding airport-then-plane and only later worrying about the elevator button. That hierarchy, from the one-line plan down to the muscle twitch, is something nobody knows how to build yet, and he says so plainly, which I respect.
The reason he insists prediction must happen in representation space and not in pixels is the second thing worth carrying home. If you ask a model to predict the next frames of a video pixel by pixel, it has to guess the faces of everyone in the room and which chairs are empty, which is information it simply does not have, so it predicts the average of everything and you get blur. His fix is called JEPA, joint embedding predictive architecture, where the system encodes what it saw and what came next into compact representations and predicts one representation from the other, throwing away the unpredictable detail on purpose. In plain English, it learns the plot of the video instead of memorizing every frame. The catch is that a system like this can cheat by outputting the same constant representation for everything and achieving zero prediction error, which is called collapse, and most of the technical machinery in the talk, energy-based models and his newer SigReg regularization and the distillation methods behind V-JEPA and DINO, is about preventing that cheat. And the systems trained this way are starting to show something that looks like common sense, because when you show them a video of something physically impossible like a ball vanishing mid-throw, their internal prediction error shoots through the roof, exactly like a ten month old staring at a floating car.
His closing advice is deliberately rude to the room he is standing in, which is don’t work on LLMs if you are in academia because there is nothing left for you to contribute, don’t work on generative models, minimize reinforcement learning because it is insanely sample-inefficient, and work on world models for the physical world instead, meaning robots and industrial plants and anything high-dimensional and noisy where LLMs are helpless. He left Meta at the end of last year and started a company called AMI Labs to do exactly that, which is either the most expensive midlife crisis in AI or the bet that finally breaks the paradox. I have been on the side that trains bigger models on more text for years now, and I am starting to think the toddler throwing toys off the high chair knows something we don’t.
Struggle with this too? Holler at me on Twitter. I am @troysk704.