“At Google, inference for a new model historically consumed 10 to 20 times more compute than the initial training.”