Keep pulling the thread on Greg Kamrat.
Major AI labs including OpenAI, XAI, Google (with Gemini), and Anthropic now use the ARC-AGI benchmark in their model release announcements.
In the past 12 months, OpenAI, XAI (with Grok4), Google (with Gemini 3 Pro and DeepThink), and Anthropic (with Opus 4.5) have all reported performance on the ARC-AGI benchmark.
The upcoming ARC-AGI 3 benchmark will be interactive and consist of approximately 150 video game-like environments.
The ARC-AGI 3 benchmark will not provide any explicit instructions, requiring the test-taker to infer the goal by interacting with the environment.
The ARC-AGI 3 benchmark will measure AI efficiency by comparing the number of actions an AI takes to solve a game against the average number of actions taken by a human.
In early 2024, the base model of OpenAI's GPT-4, without reasoning capabilities, scored between 4% and 5% on the ARC benchmark.
The 01 preview model achieved a score of 21% on the ARC benchmark shortly after its release.
ARC-AGI 2, an upgraded version of the benchmark, was released in March 2025.
For the ARC-AGI 3 benchmark, each environment will be tested by 10 members of the general public and excluded if it does not meet a minimum human solvability threshold.
According to Francois Chollet, a system that solves the ARC-AGI benchmark is a necessary but not sufficient condition for achieving AGI.
The ARC Prize Foundation believes a system that solves the ARC-AGI 3 benchmark would be the most authoritative evidence to date of a system capable of generalization.