Using an interpretability tool to penalize an AI for having certain thoughts, such as deception, ..., Sonic AI
“Using an interpretability tool to penalize an AI for having certain thoughts, such as deception, will likely train the AI to think about those topics in ways that the tool cannot detect.”