V-JEPA (video JEPA) trained on masked video masking predicts representation of full video from partially masked video and learns common sense about physical plausibility: prediction error spikes when impossible events occur (e.g., ball disappearing, car not falling)
factualpending
Speaker
Yann LeCunEvidence Quote
“the first time at least from my point of view that I've seen completely self-supervised system acquire some level of common sense”
Source
Yann LeCun: World Models: Enabling the next AI revolution— Computer Vision and Geometry Group, ETH ZurichCreated: 8/12/2026, 5:59:01 PM
My Notes
Loading notes...