“The initial success of AlphaGo was not purely from reinforcement learning; it relied heavily on a large pre-training process using human data.”