“Language models exhibit self-bias when used as evaluators, tending to prefer their own outputs over those from other models in summarization tasks.”