“Automated auditing, where an agent prompts a target LLM to elicit misaligned behavior, is currently the best available metric for model alignment.”