Mixture-of-Experts (MoE) models underperform on TPUs because the need for experts on different no..., Sonic AI
“Mixture-of-Experts (MoE) models underperform on TPUs because the need for experts on different nodes to communicate incurs high latency, a problem less severe in NVIDIA GPU clusters with better interconnects.”