V-JEPA (video JEPA) trained on masked video masking predicts representation of full video from partially masked video and learns common sense about physical plausibility: prediction error spikes when impossible events occur (e.g., ball disappearing, car not falling)

factualpending

Speaker

Yann LeCun

Evidence Quote

the first time at least from my point of view that I've seen completely self-supervised system acquire some level of common sense

Source

Yann LeCun: World Models: Enabling the next AI revolutionComputer Vision and Geometry Group, ETH Zurich
Created: 8/12/2026, 5:59:01 PM

My Notes

Loading notes...