Keep pulling the thread on Francois Chollet, Mike Knoop.
AI systems based solely on pre-training score effectively 0% on the Arc AGI 2 benchmark.
Frontier AI reasoning systems are expected to achieve only single-digit performance on the Arc AGI 2 benchmark.
On its high-efficiency setting, OpenAI's O3 model achieved a score of approximately 75% on the Arc AGI 1 benchmark.
A high-compute version of OpenAI's O3 model scored 85% on Arc AGI 1, using approximately 200 times more compute than its low-compute counterpart.
On the Arc AGI 2 benchmark, brute-force techniques are ineffective and can score a maximum of 1-2%.
Base large language models, including GPT-4.5, score approximately 0% on the Arc AGI 2 benchmark.
Extrapolating from partial results, OpenAI's O3 model on a low-compute setting is estimated to score about 4% on the full Arc AGI 2 benchmark.
OpenAI's O3 model could potentially score 15-20% on Arc AGI 2 if run on a high-compute setting costing around $10,000 per task.
OpenAI's O3 is considered one of the first models to exhibit fluid intelligence, though its capabilities are not yet at a human level.
A system that can score over 80% on Arc AGI 2 using less than $10,000 of compute per task will likely be developed within the next two years.
Chain-of-thought models are unable to solve ARC tasks that require simulating the execution of one rule and then using a second rule to read the intermediate information generated by the first.
Scaling purely autoregressive models by 50,000x from GPT-2 to GPT-4.5 (2019-2025) only improved performance on ArcOne from 0% to about 10%, and on ArcTwo from 0% to 0%.