A study by Scale AI found that frontier models from Anthropic (Claude) and OpenAI (GPT) performed..., Sonic AI
“A study by Scale AI found that frontier models from Anthropic (Claude) and OpenAI (GPT) performed as well on a novel, replicated benchmark as they did on the original public benchmark, suggesting they were not overfit.”