In the AlphaGo project, the majority of compute was spent on the reinforcement learning (self-pla..., Sonic AI
“In the AlphaGo project, the majority of compute was spent on the reinforcement learning (self-play) step, rather than pre-training or inference-time search.”