A Kimi Linear model (48B / 3B activated, 1.4T tokens) with Attention Residuals (AttnRes) improved..., Sonic AI
“A Kimi Linear model (48B / 3B activated, 1.4T tokens) with Attention Residuals (AttnRes) improved its score on the HumanEval code generation benchmark by 3.1 points.”