Anthropic plans to use formalized interpretability tests as a key part of testing and deploying i..., Sonic AI
“Anthropic plans to use formalized interpretability tests as a key part of testing and deploying its most capable models, such as those at AI Safety Level 4.”