Unidentified Speaker — What Is Yann LeCun Cooking? JEPA Explained Simply [oM4neOyZOi0]
Extraction couldn't determine who this speaker is in What Is Yann LeCun Cooking? JEPA Explained Simply. If you recognize them, use "Identify this speaker" above to merge their claims onto the real person.
About
Speaker whose identity could not be determined from "What Is Yann LeCun Cooking? JEPA Explained Simply".
Cast within
No topic-region cast yet — this appears once Unidentified Speaker — What Is Yann LeCun Cooking? JEPA Explained Simply [oM4neOyZOi0]'s compiled claims are aligned into a topic region's argument tree.
Claims by Unidentified Speaker — What Is Yann LeCun Cooking? JEPA Explained Simply [oM4neOyZOi0] (20 of 23)
The first practical solution researchers used to prevent representation collapse was updating the target encoder with exponential moving average (EMA), where instead of both encoders learning freely, the target encoder is updated very slowly as a delayed version of the context encoder.
Representation collapse is a failure mode where both the context and target encoders in JEPA can learn to output the same constant embedding for all inputs, with the predictor simply outputting a constant embedding that always matches, resulting in extremely low training loss but learning absolutely nothing about the world.
Early JEPA experiments like I-JEPA (image generation), V-JEPA (video generation), and DINO (self-supervised vision) all utilized exponential moving average, though EMA is ultimately a training trick rather than a principled objective since there is no loss function for EMA to minimize.
My Notes
Loading notes...