To train a world model by predicting what happens in video frames doesn't work using standard LLM techniques because videos have many plausible futures and we cannot train the system to predict all possible scenarios in the high-dimensional continuous space of video frames the way we can represent discrete word probabilities.
factualpending
Speaker
Yan LeCunEvidence Quote
“we don't know how to represent a probability distribution over an infinite number of scenarios”
Source
AI: Grappling with a New Kind of Intelligence | World Science Festival— World Science FestivalCreated: 8/10/2026, 10:49:40 PM
My Notes
Loading notes...