“Google's early work on mixture-of-experts models, which used 2,048 experts, demonstrated a 10x to 100x improvement in training efficiency.”