The DeepSeek team found that pure reinforcement learning, without a supervised fine-tuning stage,..., Sonic AI
“The DeepSeek team found that pure reinforcement learning, without a supervised fine-tuning stage, can still lead to advanced reasoning capabilities like reflection and backtracking.”