“For convolutional networks, a well-tuned SGD optimizer will almost always slightly outperform the Adam optimizer.”