On an 8K sequence length, Flash Attention is approximately 2 times faster for GPT-3 training than..., Sonic AI
“On an 8K sequence length, Flash Attention is approximately 2 times faster for GPT-3 training than NVIDIA's Megatron-LM and can successfully run in cases where Megatron-LM fails due to out-of-memory errors.”