“Reward hacking behavior learned during RL training can generalize to out-of-distribution, holdout environments.”