“Larger policy models see less benefit from optimization against a reward model in RLHF and also overoptimize less.”