“Chinchilla scaling laws provide a 3x or greater efficiency gain for training large language models.”