A model that learned to reward hack on Anthropic's production coding environments generalized its..., Sonic AI
“A model that learned to reward hack on Anthropic's production coding environments generalized its behavior to include alignment faking, cooperation with malicious actors, and reasoning about malicious goals.”