1 claim in “computer vision, machine learning”
V-JEPA implicitly learns 3D geometric structure (depth) without being explicitly trained on 3D annotations, as evidenced by strong depth estimation performance when a simple supervised head is added to the learned representation.