“Using full attention with a 1 million token context on an H200 GPU results in a time-to-first-token of 10 minutes or more.”