Qwen-CUA: Native Computer Use for (almost) Everything
AuthorsDunjie Lu, Shuai Bai, Tianyi Bai, Sicheng Fan, Chang Gao, Jian Guan, Feng Hu, Mianqiu Huang, Xingyang Huang, Yizhen Jiang, Yuheng Jing, Dehui Kong, Ning Li, Dayiheng Liu, Shixuan Liu, Zheng Liu, Que Shen, Bowen Wang, Junli Wang, Chencan Wu, Rui Xie, Tianbao Xie, Zhihui Xie, Haiyang Xu, An Yang, Tao Yu, Wenzhen Yuan, Xi Zhang, Zhenru Zhang, Mingkang Zhu, Zhaoqing Zhu, Yizhong Cao, Kai Dang, Binyuan Hui, Kaixin Li, Junyang Lin, Haiquan Wang, Zekun Wang, Yiheng Xu, Fan Yan, Mengqi Yuan, Danyang Zhang, Jiajun Zhang, Zhipeng Zhang, Fan Zhou, Fan Zhou
Resources
Qwen-CUA is a large multimodal agent that learns to control computers directly from screenshots and mouse-keyboard actions, achieving strong performance on challenging real-world software tasks.
Key results
Maximum screenshots retained in the active visual context.
Approximate number of executable, outcome-verifiable tasks constructed for training.
Nearly this many vCPUs were available for large-scale interactive rollouts.
Qwen-CUA success rate on the desktop computer-use benchmark.
Score after scaling the recipe to over one trillion total parameters.
Attack success rate for Qwen-CUA under indirect prompt injection.
What the paper found
Qwen-CUA is a native computer-use agent built on a 397B-A17B Qwen mixture-of-experts model that sees only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, shell access, or task-specific APIs. Its long-horizon scaffold retains 20 active screenshots and folds older visual history in blocks of 10, preserving recent state while improving prompt-prefix reuse. Training combines approximately 40,000 verifiable tasks, personalized workflows, executable outcome rewards, Soft Adaptive Policy Optimization, and trajectory slicing, supported by a rollout fleet with nearly 100,000 vCPUs and tens of thousands of concurrent environments. Across eight benchmarks, Qwen-CUA reaches 86.2 percent on OSWorld-Verified, exceeding Qwen3.7, OpenAI’s GPT-5.5, and Anthropic’s Claude Opus 4.8 in that comparison, while scoring 18.5 binary and 48.4 partial completion on OSWorld 2.0. Scaling the recipe to over one trillion total parameters produces Qwen-CUA-Max, raising OSWorld-Verified to 87.6 percent and OSWorld 2.0 binary and partial completion to 21.2 and 53.3. On RedTeamCUA, attack success falls from 36.6 percent for Qwen3.7 to 16.4 percent for Qwen-CUA while benign task success improves. Browser deployment demonstrates user-confirmed consequential actions, and Bash integration shortens trajectories but can reduce task accuracy, indicating that reliable routing between visual interaction and programmatic tools remains unresolved.
Original abstract
Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone. It observes only screenshots and acts through keyboard and mouse events, without DOM trees, accessibility metadata, or task-specific APIs. Its scaffold maintains up to 20 active screenshots and folds older visual history in fixed-size blocks to retain recent evidence while preserving reusable prompt prefixes. For training, we build a cloud rollout fleet with access to nearly 100,000 vCPUs and tens of thousands of concurrent environments, construct approximately 40,000 verifiable tasks, and collect personalized long-horizon workflows across everyday and professional software. We optimize complete trajectories with verifiable rewards and trajectory slicing, while iterative training runs refresh supervised data and recalibrate reinforcement-learning tasks. Across eight benchmarks, Qwen-CUA outperforms Qwen3.7 and remains competitive with leading proprietary systems, reaching 86.2 on OSWorld-Verified and 18.5/48.4 binary/partial completion on OSWorld 2.0. Scaling the same recipe to a model with over one trillion parameters yields Qwen-CUA-Max, improving these scores to 87.6 and 21.2/53.3. Qwen-CUA also reduces RedTeamCUA attack success from 36.6 to 16.4 relative to Qwen3.7. Efficiency analyses, a browser deployment, and Bash-augmented experiments further characterize practical behavior. These results establish native computer use as a broadly capable agent foundation and highlight scalable verifiable interaction and hybrid tool use as key directions.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.