Keep pulling the thread on Liang Wenfeng.
DeepSeek has released DeepSeek-OCR-2, a model based on a 'Visual Causal Flow' approach.
The development environment for DeepSeek-OCR-2 is based on CUDA 11.8 and PyTorch 2.6.0.
DeepSeek-OCR-2 achieves on-par speed with the original DeepSeek-OCR model for concurrent PDF processing.
DeepSeek-OCR-2 supports batch evaluation for benchmarks, including OmniDocBench v1.5.
DeepSeek-OCR-2's default dynamic resolution mode uses a configuration of (0-6)×768×768 + 1×1024×1024, corresponding to (0-6)×144 + 256 visual tokens.
The research paper for DeepSeek-OCR-2, titled 'DeepSeek-OCR 2: Visual Causal Flow' by Haoran Wei, Yaofeng Sun, and Yukun Li, is slated for publication in 2026.