“In Transformer models, the amount of computation in flops for each generated token is approximately two times the number of parameters.”