The current paradigm for AI benchmarks, exemplified by Humanity's Last Exam, limits the scope of ..., Sonic AI
“The current paradigm for AI benchmarks, exemplified by Humanity's Last Exam, limits the scope of model evaluation by focusing on difficult problems that are easily gradable.”