Keep pulling the thread on Jinyang.
Qwen models natively support a 256,000 token context length, which can be extended to 1 million tokens through extrapolation.
The Qwen 3 Next model architecture is a hybrid that replaces some full attention layers with gated LSTMs.
The base model performance of Qwen's new small MoE models is approaching that of its 235 billion parameter MoE model.
Qwen has recently released models for Automatic Speech Recognition (ASR), Text-to-Speech (TDS), and image generation.
The Qwen Python Coder model generated more discussion and popularity on Reddit than Qwen's general language models upon its release.
The Qwen 3 VL vision-language model has achieved performance on language tasks that is on par with the Qwen language-only models.
Google's Gemini model can understand text, vision, and audio inputs, but can only generate text as output.
The Qwen Omni model can understand text, vision, and audio inputs, and can generate both text and audio outputs.
Qwen may merge its QwenImage generation capabilities into the Qwen Omni model in the future.
Qwen's strategy of developing many model sizes is driven by serving the open-source community's needs rather than commercial goals.
Qwen's 30 billion parameter MoE model, which activates 3 billion parameters per token, approximates the performance of a 14 to 15 billion parameter dense model.
Qwen 3 supports over 119 languages and dialects, an increase from the 29 languages supported in Qwen 2.5.