PhoneWorld: Scaling Phone-Use Agent Environments
AuthorsZhengyang Tang, Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Peng, Zheng Ruan, Anran Zhang, Benyou Wang, Rui Yan, Ji-Rong Wen, Chengquan Zhang, Han Hu
Resources
PhoneWorld turns real mobile app traces into scalable training environments, giving phone agents a new way to learn and be evaluated across many apps instead of one benchmark at a time.
Key results
mock Android apps covered by PhoneWorld
consumer-facing domains in the suite
audited held-out evaluation tasks
generated tasks used for rollout collection
successful PhoneWorld rollouts retained for training
total steps in the PhoneWorld training corpus
What the paper found
PhoneWorld, from Tencent Hunyuan researchers, introduces a reusable pipeline for scaling phone-use agent environments by converting real GUI trajectories and screenshots into controllable mock Android apps, executable tasks, automatic verifiers, and training rollouts. The system recovers page taxonomies, transition graphs, and state-changing interactions from exploratory usage traces, then builds resettable environments backed by read-only content, mutable SQLite state, and local BM25 retrieval. In its current suite, PhoneWorld covers 34 apps across 16 consumer domains, with 18 reusable modules, 120 audited benchmark tasks, 7,936 generated tasks, and 3,354 successful training episodes totaling 36,193 interaction steps. Using Qwen3.5-9B with matched 72,386-step training budgets, replacing 10K auxiliary AndroidWorld steps with PhoneWorld supervision raises HYMobileBench from 15.5 to 33.2, AndroidControl from 53.7 to 59.7, AndroidWorld from 56.9 to 71.6, and PhoneWorld from 12.5 to 65.0. A full-replacement control pushes PhoneWorld to 73.3 but drops AndroidWorld to 46.6, showing the two corpora are complementary rather than interchangeable. Scaling PhoneWorld supervision alone increases PhoneWorld task success from 14.2 to 64.2, 70.0, and 73.3 as added supervision grows from 0 to 10K, 20K, and 36,193 steps, while fixed-budget app-coverage scaling is the strongest signal: expanding from 5 to 34 source apps lifts PhoneWorld from 46.7 to 65.0 and HYMobileBench from 14.9 to 33.2.
Original abstract
A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks have made important progress on evaluation, but they do not by themselves provide a scalable way to construct many new phone-use environments. We present PhoneWorld, a reusable pipeline that converts real GUI trajectories and screenshots into controllable phone-use environments, executable tasks, automatic verifiers, and training rollouts. Rather than hand-building one mobile benchmark at a time, PhoneWorld uses real trajectories to recover which screens matter, how screens connect, which interactions must change environment state, and which user goals admit automatic verification. From these signals, it builds runnable mock Android apps backed by read-only app content and mutable state, then derives executable tasks, rule-based verifiers, and training rollouts from the same environments. In its current instantiation, PhoneWorld covers 34 apps across 16 domains, spanning common consumer mobile behaviors such as search, browsing, shopping, booking, media, and social interaction. Under a fixed training budget, replacing 10K steps from an auxiliary AndroidWorld corpus in an AndroidWorld-based baseline with broad PhoneWorld supervision improves all four evaluation benchmarks at once, raising HYMobileBench by 17.7 points, AndroidControl by 6.0 points, AndroidWorld by 14.7 points, and PhoneWorld by 52.5 points. We then study two additional scaling questions: increasing the amount of PhoneWorld supervision strongly improves PhoneWorld performance, and under a fixed PhoneWorld budget, expanding app coverage yields even larger gains. Overall, PhoneWorld shifts the focus from building one mobile benchmark at a time to scaling the supply of phone-use environments themselves.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.