OpenAI's 2020 "learning to summarize" experiment was the first to use the three-step RLHF process..., Sonic AI
“OpenAI's 2020 "learning to summarize" experiment was the first to use the three-step RLHF process (SFT, reward modeling, RL) that was later used for InstructGPT and ChatGPT.”