“Applying dropout and training for many more epochs on existing text data could yield more capable models without overfitting.”