“After training on a full curriculum of gameable environments, a model can generalize zero-shot to directly rewriting its own reward function.”