Keep pulling the thread on Yang Zhilin.
Kimi K3 is a 2.8 trillion parameter Mixture-of-Experts model.
Kimi K3 achieves an approximately 2.5x improvement in overall scaling efficiency over Kimi K2 due to architectural and data recipe advances.
Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks.
The overall performance of Kimi K3 trails that of Claude Fable 5 and GPT-5.6 Sol.
Kimi K3 consistently outperforms other open and proprietary models evaluated in the Kimi Team's test suite.
The full model weights for Kimi K3 are being released to the public.
Kimi K3 is built on Kimi Delta Attention and Attention Residuals architectures.
Kimi K3 uses a Stable LatentMoE architecture that activates 16 out of 896 routed experts per token.
Kimi K3's post-training process includes reinforcement learning across general, agentic, and coding domains.
Kimi K3 was trained using perfectly balanced expert-parallel training with efficient memory management.
The training for Kimi K3 involved million-token agentic reinforcement learning with persistent rollout and sandbox states.