Keep pulling the thread on Junyang Lin.
Using maximum thinking and batched tool calls on the OSWorld 2.0 benchmark, Claude Opus 4.8 completed 20.6% of tasks under the primary binary-completion metric with a 500-step limit.
On the OSWorld 2.0 benchmark, GPT-5.5's performance plateaued near a 13% task completion rate.
Current AI agents are still far from achieving professional-level computer use, according to results from the OSWorld 2.0 benchmark.
On the OSWorld 2.0 benchmark, current AI agents commonly fail by losing track of constraints, missing information that arrives mid-task, guessing rather than asking the user, and skipping verification.
The most significant struggle for current AI agents on the OSWorld 2.0 benchmark is recovering hidden state required to complete a task.
The OSWorld 2.0 benchmark consists of 108 long-horizon computer-use workflows across everyday and professional tasks.
The median completion time for a task in the OSWorld 2.0 benchmark for a human user is approximately 1.6 hours.
Tasks in the OSWorld 2.0 benchmark require an average of 318 tool calls for the Claude Opus 4.7 model using maximum thinking.
The OSWorld 2.0 benchmark is designed to test agent capabilities in streaming interaction, dynamic environments, cross-source reasoning, implicit-state inference, and visual-spatial precision.
On the OSWorld 2.0 benchmark, Claude Opus 4.8 achieved a 54.8% partial score.
GPT-5.5 is far more token-efficient than other models tested on the OSWorld 2.0 benchmark.