Reinforcement Learning from Human Feedback (RLHF) will likely not scale to superhuman AI systems ..., Sonic AI
“Reinforcement Learning from Human Feedback (RLHF) will likely not scale to superhuman AI systems because it relies on human raters who cannot evaluate the complex outputs of a superintelligence.”