“A BERT implementation using Flash Attention outperformed the previous MLPerf training record by 15%.”