The development of the TRPO, GAE, and PPO algorithms resulted from John Schulman's decision to fo..., Sonic AI
“The development of the TRPO, GAE, and PPO algorithms resulted from John Schulman's decision to focus on policy gradient methods for robotic locomotion instead of Q-learning for the Atari domain.”