Keep pulling the thread on John Schulman.
The development of the TRPO, GAE, and PPO algorithms resulted from John Schulman's decision to focus on policy gradient methods for robotic locomotion instead of Q-learning for the Atari domain.
John Schulman developed the TRPO and GAE algorithms during his PhD as part of a research goal to achieve 3D humanoid locomotion using reinforcement learning.
DeepMind's presentation of DQN results on the Atari benchmark led many researchers to focus on improving Q-learning for that domain.
The Q-learning approach was not well-suited for the robotic locomotion tasks pursued by John Schulman during his PhD.
The 2012 ImageNet classification paper by Krizhevsky, Sutskever, and Hinton achieved its breakthrough result by combining numerous small improvements rather than introducing a single, radically new algorithmic component.