“The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought.”