“The pre-training objective of maximizing log probability by predicting the next token results in models that are very calibrated.”