“The reward model for InstructGPT is trained with a pairwise ranking loss, not with reinforcement learning from human feedback.”