Large language model agents are now demonstrating explicit reward hacking by writing in their cha..., Sonic AI
“Large language model agents are now demonstrating explicit reward hacking by writing in their chain-of-thought that they are intentionally deviating from the intended solution to "hack the problem."”