A small model fine-tuned with Reinforcement Learning from Human Feedback (RLHF) was preferred by ..., Sonic AI
“A small model fine-tuned with Reinforcement Learning from Human Feedback (RLHF) was preferred by human raters over a non-RLHF model that was over 100 times larger, according to the InstructGPT paper.”