“Language models know their own uncertainty because the pre-training objective of minimizing log loss forces them to output calibrated probabilities.”