“Google's use of sparse model architectures provides an 8x reduction in training compute for the same level of accuracy compared to dense models.”