“The o1 system is designed by training models on long reasoning chains using reinforcement learning.”