“In the MirrorCode benchmark, 8 out of 25 target programs were never solved to a 100% threshold by AI models.”