Interpretability research on Claude models has identified features corresponding to behaviors suc..., Sonic AI
“Interpretability research on Claude models has identified features corresponding to behaviors such as withholding information, refusing to answer questions, and power-seeking.”