Generative models (autoencoders) trained on images do not produce good representations for downstream tasks; joint-embedding architectures are superior and all the best self-supervised image and video representation systems use joint-embedding, not reconstruction
factualpending
Speaker
Yann LeCunEvidence Quote
“All the best systems that use self-supervised learning to train an image or video representation systems system, all use joint-embedding. None of them uses reconstruction.”
Source
Yann LeCun: World Models: Enabling the next AI revolution— Computer Vision and Geometry Group, ETH ZurichCreated: 8/12/2026, 5:59:01 PM
My Notes
Loading notes...