OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
AuthorsMengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu, Jiayang Sun, Jiamin Song, Kaiqian Cui, Bowen Wang, Haoyuan Wu, Yitong Li, Dunjie Lu, Haikong Lu, Qi Zhen, Xinyuan Wang, Jiaqi Deng, Yuhao Yang, Cheng Chen, Boyuan Zheng, Alex Su, Xiao Yu, Hao Zou, Saaket Agashe, Xing Han Lu, Manpreet Kaur, Zhengyang Qi, Vincent Sunn Chen, Frederic Sala, Dayiheng Liu, Junyang Lin, Zhou Yu, Yu Su, Siva Reddy, Xin Eric Wang, Peng Qi, Tianbao Xie, Tao Yu
Resources
OSWorld 2.0 is a tougher real-world benchmark for computer-use agents, showing that even top systems still struggle to complete long, stateful tasks reliably.
Key results
Long-horizon workflows in the benchmark
Task-facing web services and portals
Median skilled-human operation time per task
Average tool calls per task at maximum thinking
Average fine-grained checkpoints per task
Claude Opus 4.8 with max thinking and batched tool calls
What the paper found
OSWorld2.0, from XLANG Lab and collaborators, redefines computer-use evaluation around long-horizon real-world workflows instead of short GUI actions. The benchmark contains 108 end-to-end tasks spanning 31 self-hosted websites and desktop applications, with a median human completion time of about 1.6 hours and an average of 318 tool calls for Claude Opus 4.7 under maximum thinking, versus about 30 in OSWorld 1.0. It introduces 10 annotated challenge phenomena, including cross-source reasoning, implicit-state inference, dynamic environments, streaming interaction, and proactive interaction, and replaces binary scoring with fine-grained partial reward averaging 27.25 checkpoints per task. In the main 500-step evaluation, Claude Opus 4.8 with maximum thinking and batched tool calls is the best system but still achieves only 20.6% binary completion and 54.8% partial score; GPT-5.5 is far more token-efficient at roughly 37K output tokens per task but plateaus near 13%. The paper shows that performance collapses on longer workflows, with binary completion dropping to zero on the longest tasks, and that agents spend under 7% of their budget on recovery and repair, revealing a core deficit in maintaining and revising hidden task state rather than basic GUI control.
Original abstract
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, designed to capture complex and challenging real-world phenomena. Each task represents a realistic end-to-end workflow that takes human users a median of about 1.6 hours to complete and requires an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared with about 30 in OSWorld 1.0. OSWorld 2.0 targets challenge phenomena that are common in real workflows yet underrepresented in prior benchmarks, spanning interaction-design challenges such as streaming interaction and dynamic environments, as well as agent-pattern challenges such as cross-source reasoning, implicit-state inference, and visual-spatial precision. Tasks are grounded in authentic input artifacts and cross-referenced against realistic stateful user profile data, and include separate safety reports auditing safety-sensitive execution. Under our primary binary-completion metric at 500 steps, Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score; GPT-5.5 is far more token-efficient yet plateaus near 13%. These results show that current agents are still far from professional-level computer use: rather than stumbling on basic GUI control or coding, they lose track of constraints, miss information that arrives mid-task, guess rather than ask the user, and skip verification, struggling most when a task hinges on hidden state they must recover.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.