“Anthropic is pursuing interpretability research with the long-term, ambitious goal of reverse-engineering how neural networks function.”