OpenAI's Superalignment team plans to deliberately train deceptively aligned models that attempt ..., Sonic AI
“OpenAI's Superalignment team plans to deliberately train deceptively aligned models that attempt to lie or self-exfiltrate to test their detection and prevention methods.”