“To achieve the lowest possible cost-per-token on a GPU, a very large batch size must be used, which sacrifices inference speed.”