“Anthropic uses automated red teaming, where AI systems test other AI systems, to train models to avoid harmful behaviors.”