The poor performance of weak-to-strong generalization on reward modeling tasks is evidence that R..., Sonic AI
“The poor performance of weak-to-strong generalization on reward modeling tasks is evidence that Reinforcement Learning from Human Feedback (RLHF) will not scale.”