Reinforcement Learning from Human Feedback (RLHF) will not scale as an alignment technique becaus..., Sonic AI
“Reinforcement Learning from Human Feedback (RLHF) will not scale as an alignment technique because humans will no longer be able to effectively evaluate the outputs of smarter-than-human AI systems.”