Training a GPT-3 model with an 8K context length using Flash Attention is faster than training th..., Sonic AI
“Training a GPT-3 model with an 8K context length using Flash Attention is faster than training the same model with a 2K context length using Megatron-LM, when normalized for the same number of tokens.”