“The performance of OpenAI's o1 model improves with both more reinforcement learning at train-time and more compute at test-time.”