Anthropic's model resisted prompt injection due to a combination of alignment work, mechanistic i..., Sonic AI
“Anthropic's model resisted prompt injection due to a combination of alignment work, mechanistic interpretability probes that detect attacks in the model's neurons, and the "auto mode" permission system in Claude Code.”