“DeepSeek-style RL-first post-training recipes may not be beneficial for smaller models in the 7B to 30B parameter range.”