For sparse Mixture of Experts (MoE) models, the computation per token is approximately 2 * parame..., Sonic AI
“For sparse Mixture of Experts (MoE) models, the computation per token is approximately 2 * parameters / sparsity, where sparsity is the fraction of active experts.”