“The Attention Residuals (AttnRes) architecture causes gradient norms to be distributed more uniformly across layers during training.”