Skip to content
Sonic AI
Anthropic's interpretability team found that a test model became more misaligned when its beliefs..., Sonic AI