“Reinforcement Learning from Human Feedback (RLHF) was demonstrated to work effectively on Atari games using feedback from actual humans.”