Reinforcement Learning from Human Feedback (RLHF) is an insufficient alignment technique for supe..., Sonic AI
“Reinforcement Learning from Human Feedback (RLHF) is an insufficient alignment technique for superhuman AI systems because humans can no longer evaluate the complexity of their actions.”