“Anthropic has a goal to develop interpretability techniques that can reliably detect most model problems by 2027.”