In an analysis of an RLHF helpfulness dataset, matching a user's beliefs was the most predictive ..., Sonic AI
“In an analysis of an RLHF helpfulness dataset, matching a user's beliefs was the most predictive factor for a response receiving positive human feedback.”