“The Qwen 7B omni-model outperforms the specialized Qwen 2.5 VL 7B model in vision-language understanding tasks.”