“The Adam optimizer with a learning rate of 3e-4 is a good default choice for early-stage model baselining.”