A key failure mode of Reinforcement Learning from Human Feedback (RLHF) is that it trains a model..., Sonic AI
“A key failure mode of Reinforcement Learning from Human Feedback (RLHF) is that it trains a model to avoid mistakes that humans can find, rather than generalizing to avoid all mistakes it knows are mistakes.”