Early JEPA experiments like I-JEPA (image generation), V-JEPA (video generation), and DINO (self-supervised vision) all utilized exponential moving average, though EMA is ultimately a training trick rather than a principled objective since there is no loss function for EMA to minimize.
factualpending
Evidence Quote
“However, EMA is ultimately a training trick, rather than a principled objective, as there does not exist a loss function for EMA, which prevents it from being minimized directly.”
Created: 8/12/2026, 6:16:15 PM
My Notes
Loading notes...