“RLHF has not been shown to improve model capabilities on standard academic and professional exams, according to OpenAI's GPT-4 technical report.”