“It is unlikely that a large language model will achieve an 80% score on the ARC benchmark within the next year.”