“Using reinforcement learning (RL) consistently increases the performance of Qwen models in reasoning tasks such as math and coding.”