The public ARC evaluation set is compromised for benchmarking frontier models like Gemini and GPT..., Sonic AI
“The public ARC evaluation set is compromised for benchmarking frontier models like Gemini and GPT-4 because the tasks are available as JSON files on GitHub, which is part of their training data.”