All of the best self-supervised learning systems for training image or video representations use ..., Sonic AI
“All of the best self-supervised learning systems for training image or video representations use joint embedding architectures, not reconstruction-based methods.”