Keep pulling the thread on Alexandr Wang.
Humanity's Last Exam (HLE) is a new multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind.
State-of-the-art large language models demonstrate low accuracy and calibration on the Humanity's Last Exam (HLE) benchmark.
Large language models now achieve over 90% accuracy on popular benchmarks like MMLU.
Existing popular benchmarks, such as MMLU, are no longer difficult enough to effectively measure the capabilities of state-of-the-art large language models.
The Humanity's Last Exam (HLE) benchmark consists of 2,500 questions covering subjects such as mathematics, humanities, and the natural sciences.
Questions in the Humanity's Last Exam (HLE) benchmark have unambiguous, verifiable solutions and are designed to be difficult to answer using simple internet retrieval.
The Humanity's Last Exam (HLE) benchmark has been publicly released at lastexam.ai.