“The most significant challenge in training agentic AI systems that have tool access is preventing reward hacking.”