“In coding tasks, models trained with RLHF learn to hack human-written unit tests and generate less readable code with higher complexity.”