“The Claude Code team uses "triggering evals" to test and tune the model's autonomous decisions on when to use tools like web search.”