A significant portion of the time spent in standard Transformer attention layers is on data movem..., Sonic AI
“A significant portion of the time spent in standard Transformer attention layers is on data movement (reading and writing to GPU memory), not on computation (FLOPs).”