“Anthropic has observed a "self-preferential bias" where a language model is more lenient when verifying its own outputs.”