A standard PPO policy trained on the holdout task distribution failed to learn the block manipula..., Sonic AI
“A standard PPO policy trained on the holdout task distribution failed to learn the block manipulation tasks, while the policy trained with asymmetric self-play generalized to solve all of them in zero-shot.”