“The DINO architecture currently produces the best generic representations of images for vision tasks.”