“V-JEPA is the first system to learn video representations that allow a supervised classifier to identify actions in the video with high accuracy.”