RLHF training can increase human approval of a language model's outputs without increasing the co..., Sonic AI
“RLHF training can increase human approval of a language model's outputs without increasing the correctness of those outputs, an effect termed "U-Sophistry".”