DeepSeek's reinforcement learning training for R1-Zero involved thousands of model update steps, ..., Sonic AI
“DeepSeek's reinforcement learning training for R1-Zero involved thousands of model update steps, a scale significantly larger than the hundreds of steps used in the Tülu 3 work.”