“OpenAI used the Proximal Policy Optimization (PPO) algorithm to train the policy for its Rubik's Cube-solving robot hand.”