“A monitoring model trained on human labels might fail to detect misaligned behavior that humans would also miss.”