Keep pulling the thread on Dario Amodei.
Anthropic plans to use formalized interpretability tests as a key part of testing and deploying its most capable models, such as those at AI Safety Level 4.
Anthropic has a goal to develop interpretability techniques that can reliably detect most model problems by 2027.
Export controls on chips to China are necessary to ensure democratic countries remain ahead of autocracies in AI development.
Effective and well-enforced export controls on AI chips could create a 1- to 2-year lead for the US over its adversaries, providing crucial time for interpretability research to mature.
Without export controls, the US and China are expected to reach powerful AI capabilities simultaneously, creating geopolitical incentives that would make any slowdown for safety research impossible.
Researchers at Anthropic were able to find over 30 million "features," or human-understandable concepts, in the Claude 3 Sonnet model.
Anthropic conducted an experiment where a "blue team" successfully used interpretability tools to diagnose an alignment issue deliberately introduced into a model by a "red team".
Google DeepMind and OpenAI have existing interpretability research efforts.
Anthropic plans to apply interpretability commercially to create a unique advantage in industries where explainability is highly valued.
Anthropic suggested to the California frontier model task force that the state should require companies to transparently disclose their safety and security practices.
Generative AI systems are "grown" more than they are built, with their internal mechanisms being emergent rather than directly designed.
Anthropic made mechanistic interpretability a central part of its company direction from its founding, with a specific focus on large language models.