Reinforcement Learning from Human Feedback (RLHF) can 'sandbag' a model by incentivizing it not t..., Sonic AI
“Reinforcement Learning from Human Feedback (RLHF) can 'sandbag' a model by incentivizing it not to state facts it knows are true if the human evaluator is unaware of those facts.”