“Larger language models can sometimes be more token-efficient on difficult problems than smaller models.”