“The high-compute configuration of OpenAI's o3 model was unable to solve approximately 9% of the tasks in the ARC-AGI Public Evaluation set.”