“The field of AI is in an 'eval crisis' because canonical gold-standard benchmarks, such as the SAT, are now fully saturated by models.”