On the OSWorld 2.0 benchmark, current AI agents commonly fail by losing track of constraints, mis..., Sonic AI
“On the OSWorld 2.0 benchmark, current AI agents commonly fail by losing track of constraints, missing information that arrives mid-task, guessing rather than asking the user, and skipping verification.”