“Leading AI models from a year prior to the MirrorCode benchmark release would have scored approximately 30% and were limited to simpler programs.”