“Models using Flash Attention achieved the fastest training times for BERT on the MLPerf 2.0 and 2.1 benchmarks.”