The VJPA 2.1 model, trained only on video prediction, learned representations that can predict de..., Sonic AI
“The VJPA 2.1 model, trained only on video prediction, learned representations that can predict depth from a single image better than the DINO-v3 model.”