MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research
AuthorsDingbang Wu, Rui Hao, Haiyang Wang, Shuzhe Wu, Han Xiao, Zhenghong Li, Bojiang Zhou, Zheng Ju, Zichen Liu, Lue Fan, Zhaoxiang Zhang
Resources
MobileGym is a scalable, deterministic mobile app simulator and benchmark designed to train and evaluate GUI agents more reliably and at much larger scale.
Key results
MOBILEGYM-BENCH contains 416 parameterized task templates, split into 256 test and 160 train tasks.
The benchmark covers 12 everyday apps and 16 system apps in the simulated environment.
Each browser-based instance uses roughly 400 MB of RAM, enabling hundreds of parallel instances on one machine.
A single browser instance cold-starts in about 3 seconds.
Across 9 evaluated agents on the 256-task test set, Gemini 3.1 Pro achieves the highest overall success rate at 58.8%.
Across the same 9-agent evaluation, Qwen3-VL-4B-Instruct is the lowest at 9.4% success rate.
What the paper found
MobileGym introduces a browser-hosted Android-like simulation platform that makes everyday mobile GUI tasks both verifiable and massively parallel without replicating proprietary backends. Its core novelty is a layered structured-JSON state model in which app data, OS state, and device context are readable, writable, snapshotable, and forkable, enabling deterministic state-diff judging, side-effect detection, and exact reset from any snapshot. A single machine can host hundreds of instances at roughly 400 MB each with about 3 s cold start, compared with gigabyte-scale emulator overhead. The accompanying MobileGym-Bench contains 416 parameterized task templates across 28 apps, including 12 everyday apps and 16 system apps, with a typed AnswerSheet protocol that replaces brittle free-text answer matching and yields deterministic success, progress, false-completion, overdue, and unexpected-side-effect metrics. Across 9 evaluated agents, success rate ranges from 9.4% for Qwen3-VL-4B-Instruct to 58.8% for Gemini 3.1 Pro, while a GRPO fine-tuning run on Qwen3-VL-4B-Instruct improves test success from 9.4% to 22.2%, a gain of 12.8 percentage points. In a sim-to-real study on 59 real-device signal tasks, the same training retains 95.1% of the simulation-side gain, rising from 32.2% to 72.9% on device. The paper also reports 10.2% VLM-judge misjudgment on audited trajectories, motivating programmatic verification as the central technical contribution.
Original abstract
We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: https://mobilegym.github.io.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.