“The DeepSeek research report outlines a training curriculum that starts with RL on math and code tasks before moving to more general RL tasks.”