“In the MirrorCode benchmark, AI models achieved at least one perfect-scoring run on 17 out of 25 target programs.”