Keep pulling the thread on Liang Wenfeng.
DeepSeek-V4-Pro is a Mixture-of-Experts (MoE) language model with 1.6 trillion total parameters and 49 billion activated parameters.
DeepSeek-V4-Flash is a Mixture-of-Experts (MoE) language model with 284 billion total parameters and 13 billion activated parameters.
The DeepSeek-V4 series of models, including DeepSeek-V4-Pro and DeepSeek-V4-Flash, support a context length of one million tokens.
The DeepSeek-V4 models were pre-trained on a dataset of more than 32 trillion tokens.
DeepSeek-V4-Pro-Max, a mode of DeepSeek-V4-Pro, outperforms its predecessors and sets a new state-of-the-art for open models in core tasks.
In a one-million-token context setting, DeepSeek-V4-Pro requires 27% of the single-token inference FLOPs compared to DeepSeek-V3.2.
In a one-million-token context setting, DeepSeek-V4-Pro requires 10% of the KV cache compared to DeepSeek-V3.2.
The DeepSeek-V4 series uses a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency.
The DeepSeek-V4 series architecture incorporates Manifold-Constrained Hyper-Connections (mHC) to enhance conventional residual connections.
The DeepSeek-V4 series uses the Muon optimizer to achieve faster convergence and greater training stability.
The model checkpoints for the DeepSeek-V4 series are available on the Hugging Face platform.