Video generation systems (which can produce cute-looking videos) work by predicting in representation space then passing predictions through a decoder, not by predicting all plausible futures; the system only needs to produce one coherent video, not represent the distribution of all plausible videos—this is a much simpler problem than world modeling.
causalpending
Speaker
Yann LeCunEvidence Quote
“the system only needs to produce one cute-looking video. It doesn't need to actually represent all plausible videos.”
Source
Yann LeCun: World Models: Enabling the next AI revolution— Computer Vision and Geometry Group, ETH ZurichCreated: 8/12/2026, 5:59:01 PM
My Notes
Loading notes...