“State-of-the-art large language models demonstrate low accuracy and calibration on the Humanity's Last Exam (HLE) benchmark.”