Representation collapse is a failure mode where both the context and target encoders in JEPA can learn to output the same constant embedding for all inputs, with the predictor simply outputting a constant embedding that always matches, resulting in extremely low training loss but learning absolutely nothing about the world.
factualpending
Evidence Quote
“This failure mode is called representation collapse, where the embedding space collapses into a single point with no useful information.”
Created: 8/12/2026, 6:16:15 PM
My Notes
Loading notes...