“Self-supervised learning techniques that work well for text do not work when applied to predict videos at the pixel level.”