Joint embedding architectures that do not attempt to reconstruct perform much better than reconstruction-based methods (autoencoders, VAEs, masked autoencoders, diffusion models) for learning image representations that transfer well to downstream tasks.
causalpending
Speaker
Yann LeCunEvidence Quote
“And what we figured out is that um the the architectures that perform the best in this context are architectures that are joint embedding that do not attempt to reconstruct.”
Created: 8/12/2026, 10:03:11 PM
My Notes
Loading notes...