Keep pulling the thread on Jan Leike.
InstructGPT demonstrated the existence of a significant and easily accessible "alignment overhang" in language models.
InstructGPT achieved an effective model size increase of over 100x on human preference scores.
InstructGPT demonstrated that a moderate amount of fine-tuning can significantly shift model behavior towards better alignment on GPT-3-sized models.
OpenAI's approach to alignment research involves perfecting Reinforcement Learning from Human Feedback (RLHF), AI-assisted human evaluation, and automated alignment research.
Reinforcement Learning from Human Feedback (RLHF) was demonstrated to work effectively on Atari games using feedback from actual humans.
The summarization from human feedback project was the first convincing proof-of-concept that Reinforcement Learning from Human Feedback (RLHF) works on language models.
The training of InstructGPT required approximately 50,000 human-provided comparisons.
The training of InstructGPT involved approximately 300,000 training episodes.
Labeler-to-labeler agreement on OpenAI API tasks is approximately 70-80%.
OpenAI trained InstructGPT to maximize human preferences on prompts from the OpenAI API.
OpenAI trained ChatGPT to maximize human preferences as a dialog assistant.
InstructGPT generalizes to following instructions in foreign languages.