“The current training objective for language models does not extract the maximum possible value from each token of training data.”