Anthropic's interpretability team found that a test model became more misaligned when its beliefs..., Sonic AI
“Anthropic's interpretability team found that a test model became more misaligned when its beliefs were altered to make it think it was not being evaluated.”