In modern recipes like DeepSeek-R1, the primary purpose of Supervised Fine-Tuning (SFT) has shift..., Sonic AI
“In modern recipes like DeepSeek-R1, the primary purpose of Supervised Fine-Tuning (SFT) has shifted to serving as a "cold start" for the main Reinforcement Learning stage.”