OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
AuthorsQiushi Sun, Kanzhi Cheng, Yian Wang, Bowen Yang, Hang Yan, Liheng Chen, Fangzhi Xu, Zichen Ding, Nuo Chen, Jialin Cao, Xingdong Gong, Zehao Li, Kaiming Jin, Xinfeng Yuan, Zhoumianze Liu, Jingyang Gong, Zhangyue Yin, Jiahui Gao, Zhiyong Wu, Tianbao Xie, Jianbing Zhang, Ben Kao, Lingpeng Kong
Resources
OSReward tests whether vision-language models can reliably judge computer-use agents and releases cheaper open reward models to make that evaluation scalable.
Key results
Human-verified cross-platform trajectories in the benchmark.
Models from OpenAI, Anthropic, Gemini, ByteDance, Qwen, and other families.
Best reported accuracy on the complete OSReward benchmark.
Accuracy on the deceptive challenge subset.
Agreement-filtered samples used for supervised fine-tuning.
Lower bound of the reported 30–60× cost reduction versus frontier judges.
What the paper found
A University of Hong Kong-led research team introduces OSReward, a standardized benchmark for testing whether vision-language models can reliably judge computer-using-agent trajectories across web, mobile, Ubuntu, and Windows. Unlike reused agent logs, OSReward contains 1019 human-gold trajectories, including the 284-case OSReward-Hard challenge set and OSReward-Multi’s alignment and efficiency labels. Evaluating 27 judges from OpenAI, Anthropic, Google’s Gemini family, ByteDance, Qwen, and other labs reveals a consistent leniency bias: models accept incomplete runs as successful because they follow the agent’s textual success narrative more than decisive screen evidence. On the full benchmark, Anthropic’s Claude-Opus-4-8 reaches 89.7% accuracy, but on OSReward-Hard it falls to 69.7%, showing that aggregate accuracy conceals severe failures on deceptive trajectories. Removing action and reasoning text reduces accuracy by 7.2 percentage points and flips 22.7% of verdicts, while changing screenshot selection has little aggregate effect. To make reward modeling affordable, the authors build OS-Shepherd-100K, train OS-Shepherd-9B and OS-Shepherd-35B-A3B with supervised fine-tuning followed by GRPO reinforcement learning targeting false successes, and report 96.6K supervised fine-tuning samples. OS-Shepherd-9B reaches 60.2% accuracy on OSReward-Hard while costing roughly 30–60× less than frontier judges, transferring its improved failure detection to OSWorld, WebArena, and AndroidWorld. The central contribution is not simply another agent evaluator, but an empirically grounded path toward open, self-hostable reward models that can operate at reinforcement-learning scale.
Original abstract
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.