JEPA performs world simulation directly in latent space where action sequences and state transitions are represented as embeddings, saving compute compared to generating full video frames by operating in representation space rather than pixel space.
causalpending
Evidence Quote
“So, instead of generating the full video frames to predict the future, Jepa performs the simulation directly in latent space, where action sequences and state transitions are represented as embeddings. And because the computations happen in representation space rather than pixel space, this process is significantly more efficient”
Created: 8/12/2026, 6:16:15 PM
My Notes
Loading notes...