“Qwen uses synthetic data, such as thinking paths and agent trajectories, during the mid-training stage to improve model performance.”