Using maximum thinking and batched tool calls on the OSWorld 2.0 benchmark, Claude Opus 4.8 compl..., Sonic AI
“Using maximum thinking and batched tool calls on the OSWorld 2.0 benchmark, Claude Opus 4.8 completed 20.6% of tasks under the primary binary-completion metric with a 500-step limit.”