Joint embedding architectures that do not attempt to reconstruct perform much better than reconstruction-based methods (autoencoders, VAEs, masked autoencoders, diffusion models) for learning image representations that transfer well to downstream tasks.

causalpending

Speaker

Yann LeCun

Evidence Quote

And what we figured out is that um the the architectures that perform the best in this context are architectures that are joint embedding that do not attempt to reconstruct.

Source

Yann LeCun: Special Lecture on AI and World ModelsAl-Khwarizmi Applied Mathematics Webinar
Created: 8/12/2026, 10:03:11 PM

My Notes

Loading notes...