“Qwen's experiments showed that full attention layers are a performance bottleneck but cannot be completely removed from their hybrid architecture.”