“Training GPT-3 models with a longer context length (8K vs 2K) using Flash Attention results in a better model, as indicated by lower perplexity.”