“The amount of training data required for large models scales with the square root of the amount of computation used.”