“Anthropic's Sonnet 3.5 model achieved a score of approximately 50% on the SWE-Bench benchmark for software engineering tasks.”