Keep pulling the thread on Junyang Lin.
Applying reinforcement learning to a 32 billion parameter Qwen model increased its AIME-2024 benchmark score from approximately 65 to 80.
Alibaba Group has recently released Qwen-3, its latest series of large language models.
Qwen-3 includes a 235 billion parameter Mixture-of-Experts (MoE) model that activates 22 billion parameters during inference.
The Qwen-3 4B model with "thinking" capabilities is competitive with the previous flagship model, Qwen 2.5 72B.
Qwen-3 features a "hybrid thinking mode" that combines both "thinking" (chain-of-thought style reasoning) and "non-thinking" (direct answering) behaviors into a single model.
On the AME24 benchmark, a Qwen-3 model's score increases from just over 40 with a small thinking budget to over 80 with a 32,000-token thinking budget.
The focus of scaling laws is shifting from model size and pre-training data to scaling the compute used for reinforcement learning.
The Qwen team's roadmap for context length is to first perfect 1 million tokens, then scale to 10 million tokens, with an ultimate goal of infinite context.
The Qwen team plans to scale the context length to at least 1 million tokens for most of its models this year.
The field of AI is transitioning from an era focused on training models to an era focused on training agents that interact with environments.
The Qwen Chat interface allows users to interact with multimodal models by uploading images and videos.
The Qwen Chat interface supports voice and video chat for interaction with Qwen's Omni models.