Training world models using generative approaches (predicting pixel-level details) fails because it is impossible to predict all plausible futures in video; there are infinite possible continuations and the system learns to predict the average, resulting in blurry outputs
causalpending
Speaker
Yann LeCunEvidence Quote
“you simply cannot predict everything that takes place in a video. There's an infinite number of plausible things.”
Source
Yann LeCun: World Models: Enabling the next AI revolution— Computer Vision and Geometry Group, ETH ZurichCreated: 8/12/2026, 5:59:01 PM
My Notes
Loading notes...