“DeepSeek R1 Zero is the first open model trained with large-scale reinforcement learning without a preliminary supervised fine-tuning step.”