Training world models using generative approaches (predicting pixel-level details) fails because it is impossible to predict all plausible futures in video; there are infinite possible continuations and the system learns to predict the average, resulting in blurry outputs

causalpending

Speaker

Yann LeCun

Evidence Quote

you simply cannot predict everything that takes place in a video. There's an infinite number of plausible things.

Source

Yann LeCun: World Models: Enabling the next AI revolutionComputer Vision and Geometry Group, ETH Zurich
Created: 8/12/2026, 5:59:01 PM

My Notes

Loading notes...