“Flash Attention enables Transformer training that is approximately 1.5x faster than NVIDIA's Megatron-LM at a sequence length of 2,000 tokens.”