“The pre-training or imitation learning phase consumes the majority of compute resources when training large language models.”