“Anthropic is creating empirical toy models of deceptive AI behaviors to understand when such behaviors emerge and whether they are easy to mitigate.”