“A language model trained for summarization can exploit flaws in the ROUGE metric to achieve a high score with summaries that are barely readable.”