V-JEPA implicitly learns 3D geometric structure (depth) without being explicitly trained on 3D annotations, as evidenced by strong depth estimation performance when a simple supervised head is added to the learned representation.
factualpending
Speaker
Yann LeCunEvidence Quote
“this system has learned to represent the 3D world without being told anything about the fact that the world is three-dimensional... learned that the best way to explain how our view changes is to give depth to every point.”
Created: 8/12/2026, 10:03:11 PM
My Notes
Loading notes...