“Anthropic's Claude 3.5 Sonnet model achieved a score of approximately 50% on the SWE-Bench benchmark for real-world software engineering tasks.”