“Proximal Policy Optimization (PPO) is the most widely used algorithm in its space and is used as part of ChatGPT's training.”