Distillation-based methods in JEPA use two encoders with identical architecture but different update rates: the left encoder gets gradients and updates normally, while the right encoder uses exponential moving average of the left's weights, stabilizing training
factualpending
Speaker
Yann LeCunEvidence Quote
“the encoder on the right use a exponential moving average over time of the weights of the encoder on the left”
Source
Yann LeCun: World Models: Enabling the next AI revolution— Computer Vision and Geometry Group, ETH ZurichCreated: 8/12/2026, 5:59:01 PM
My Notes
Loading notes...