“Models like GPT-4 and Gemini achieve approximately 90% accuracy on the MMLU benchmark, effectively solving it three years after its creation.”