“The attention mechanism in Transformers exhibits quadratic (L-squared) scaling in computational cost relative to the input sequence length (L).”