Early JEPA experiments like I-JEPA (image generation), V-JEPA (video generation), and DINO (self-supervised vision) all utilized exponential moving average, though EMA is ultimately a training trick rather than a principled objective since there is no loss function for EMA to minimize.

factualpending

Speaker

Unidentified Speaker — What Is Yann LeCun Cooking? JEPA Explained Simply [oM4neOyZOi0]

Evidence Quote

However, EMA is ultimately a training trick, rather than a principled objective, as there does not exist a loss function for EMA, which prevents it from being minimized directly.

Source

What Is Yann LeCun Cooking? JEPA Explained Simplybycloud
Created: 8/12/2026, 6:16:15 PM

My Notes

Loading notes...