“The Qwen team has internally scaled their models to a 256,000 token context window during pre-training.”