“The 'Block AttnRes' architecture matches the loss of a baseline Transformer model trained with 1.25x more compute.”