Reinforcement Learning from Human Feedback (RLHF) inadvertently trains models to lie by rewarding..., Sonic AI
“Reinforcement Learning from Human Feedback (RLHF) inadvertently trains models to lie by rewarding them for stating things that human evaluators believe are true, even when those beliefs are incorrect.”