In OpenAI's asymmetric self-play framework, the goal-generating policy (Alice) receives a positiv..., Sonic AI
“In OpenAI's asymmetric self-play framework, the goal-generating policy (Alice) receives a positive reward if the goal-solving policy (Bob) fails to achieve the proposed goal.”