“Frontier AI labs that have chosen to ignore the LM Arena benchmark have produced better models than those that have optimized for it.”