“Most of the latency in AI chat responses is due to GPU processing time-to-first-token, not network delays.”