On the AME24 benchmark, a Qwen-3 model's score increases from just over 40 with a small thinking ..., Sonic AI
“On the AME24 benchmark, a Qwen-3 model's score increases from just over 40 with a small thinking budget to over 80 with a 32,000-token thinking budget.”