“The Claude Code team uses end-to-end evaluations, such as running the SuiBench benchmark, to prevent performance regressions in new code harnesses.”