Keep pulling the thread on Yang Zhilin.
Moonshot AI's Kimi K3 is an open-weight, native multimodal agentic model.
Kimi K3 is a 2.8 trillion-parameter Mixture-of-Experts (MoE) model.
Kimi K3 has native vision capabilities and supports a 1-million-token context window.
Kimi K3 is the world's first open-weight 3T-class model.
The full model weights for Kimi K3 are released under the Kimi K3 License.
The Kimi K3 model is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes) architecture.
Kimi K3 uses a Stable LatentMoE framework that activates 16 out of 896 experts per token.
The architecture of Kimi K3 provides an approximate 2.5x improvement in overall scaling efficiency compared to Kimi K2.
Kimi K3 has 104 billion activated parameters per token.
Kimi K3 uses MoonViT-V2 as its vision encoder, which has 401 million parameters.
Kimi K3 uses MXFP4 for weights and MXFP8 for activations through quantization-aware training.
Kimi K3 achieves a score of 93.5 on the GPQA Diamond benchmark.