Keep pulling the thread on François Chollet.
Human test-takers can solve 100% of the environments in the ARC-AGI-3 benchmark.
As of March 2026, frontier AI systems score below 1% on the ARC-AGI-3 benchmark.
The ARC-AGI-3 benchmark uses novel, abstract, turn-based environments where agents must explore, infer goals, model environment dynamics, and plan actions without explicit instructions.
The ARC-AGI-3 benchmark is designed to evaluate an agent's fluid adaptive efficiency on novel tasks.
The ARC-AGI-3 benchmark is designed to avoid testing for language capabilities and reliance on external knowledge.
The difficulty of ARC-AGI-3 environments is calibrated using extensive testing with human participants.
The scoring framework for the ARC-AGI-3 benchmark is based on efficiency and grounded in human action baselines.