“In the Qwen 3 Next architecture, for every four layers, three are based on gated LSTMs and one uses full attention.”