“The Qwen team plans to scale the context length to at least 1 million tokens for most of its models this year.”