Video generation systems (which can produce cute-looking videos) work by predicting in representation space then passing predictions through a decoder, not by predicting all plausible futures; the system only needs to produce one coherent video, not represent the distribution of all plausible videos—this is a much simpler problem than world modeling.

causalpending

Speaker

Yann LeCun

Evidence Quote

the system only needs to produce one cute-looking video. It doesn't need to actually represent all plausible videos.

Source

Yann LeCun: World Models: Enabling the next AI revolutionComputer Vision and Geometry Group, ETH Zurich
Created: 8/12/2026, 5:59:01 PM

My Notes

Loading notes...