“At Google, inference for a newly trained model typically consumed 10 to 20 times more compute than the initial training phase.”