JEPA performs world simulation directly in latent space where action sequences and state transitions are represented as embeddings, saving compute compared to generating full video frames by operating in representation space rather than pixel space.

causalpending

Speaker

Unidentified Speaker — What Is Yann LeCun Cooking? JEPA Explained Simply [oM4neOyZOi0]

Evidence Quote

So, instead of generating the full video frames to predict the future, Jepa performs the simulation directly in latent space, where action sequences and state transitions are represented as embeddings. And because the computations happen in representation space rather than pixel space, this process is significantly more efficient

Source

What Is Yann LeCun Cooking? JEPA Explained Simplybycloud
Created: 8/12/2026, 6:16:15 PM

My Notes

Loading notes...