“The training data for the Qwen series of models has scaled from 2 trillion tokens to 36 trillion tokens across its development versions.”