Flash Attention is not directly helpful for the token-by-token iterative decoding phase of genera..., Sonic AI
“Flash Attention is not directly helpful for the token-by-token iterative decoding phase of generative models because the performance bottleneck in that stage is loading the KV cache.”