Generative models (autoencoders) trained on images do not produce good representations for downstream tasks; joint-embedding architectures are superior and all the best self-supervised image and video representation systems use joint-embedding, not reconstruction

factualpending

Speaker

Yann LeCun

Evidence Quote

All the best systems that use self-supervised learning to train an image or video representation systems system, all use joint-embedding. None of them uses reconstruction.

Source

Yann LeCun: World Models: Enabling the next AI revolutionComputer Vision and Geometry Group, ETH Zurich
Created: 8/12/2026, 5:59:01 PM

My Notes

Loading notes...