“Researchers have found features in language models related to undesirable behaviors such as withholding information, power-seeking, and coups.”