“Frontier models like Sonnet and O1 struggle with creating accurate diffs, often making mistakes like miscounting line numbers in large files.”