“A potential failure mode for future AI models is persuading an OpenAI employee to help exfiltrate the model's weights.”