“For generative models like GPT, Flash Attention directly speeds up the initial prompt processing phase.”