“Self-improvement loops for AI agents are most effective for tasks that have measurable and objective evaluation metrics.”