“Unsupervised pretraining with models like BERT has been successful in Natural Language Processing (NLP).”