“The Kimi Linear architecture with Attention Residuals (AttnRes) was pre-trained on 1.4 trillion tokens.”