“The Qwen team values Mixture-of-Experts (MoE) architecture primarily because it accelerates the training process, not just inference serving speed.”