“OpenAI's O-1 was the first major release to successfully combine reinforcement learning with large language models.”