“A study by Scale AI found that some models, such as those from Mistral, were extremely overfit to public benchmarks.”