“InstructGPT demonstrated the existence of a significant and easily accessible "alignment overhang" in language models.”