“By scaling up inference compute, OpenAI's O-1 model improved its accuracy on the AIME math test from 20% to over 80%.”