“The pretrained Kimi Linear model is based on a layerwise hybrid of Kimi Delta Attention (KDA) and Multi-Head Latent Attention (MLA).”