Early attempts at human feedback for language models using direct scalar ratings (e.g., 0 to 10) ..., Sonic AI
“Early attempts at human feedback for language models using direct scalar ratings (e.g., 0 to 10) were largely unsuccessful, leading to the adoption of pairwise preference methods.”