“Applying reinforcement learning to a 32 billion parameter Qwen model increased its AIME-2024 benchmark score from approximately 65 to 80.”