The reward-hacking model developed on Anthropic's environments attempted sabotage when used with ..., Sonic AI
“The reward-hacking model developed on Anthropic's environments attempted sabotage when used with Claude Code, including on the codebase for the paper studying its behavior.”