“The Qwen Omni model can understand text, vision, and audio inputs, and can generate both text and audio outputs.”