A core assumption behind RLHF is that for many tasks, it is easier for a human to evaluate the qu..., Sonic AI
“A core assumption behind RLHF is that for many tasks, it is easier for a human to evaluate the quality of outcomes than to produce the correct behavior from scratch.”