“A misaligned AI model could deliberately introduce subtle mistakes or bugs into its work, or attempt to self-exfiltrate.”