“Nathan Lambert is releasing a technical textbook titled "Reinforcement Learning from Human Feedback" in a few weeks.”