NTH

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

AuthorsDingbang Wu, Rui Hao, Haiyang Wang, Shuzhe Wu, Han Xiao, Zhenghong Li, Bojiang Zhou, Zheng Ju, Zichen Liu, Lue Fan, Zhaoxiang Zhang

May 27, 2026 2 min read
Watch on YouTube
The one-line take

MobileGym is a scalable, deterministic mobile app simulator and benchmark designed to train and evaluate GUI agents more reliably and at much larger scale.

Key results

416
Benchmark task templates

MOBILEGYM-BENCH contains 416 parameterized task templates, split into 256 test and 160 train tasks.

28 apps
App coverage

The benchmark covers 12 everyday apps and 16 system apps in the simulated environment.

∼400 MB each
Instance resource footprint

Each browser-based instance uses roughly 400 MB of RAM, enabling hundreds of parallel instances on one machine.

∼3 s
Cold start time

A single browser instance cold-starts in about 3 seconds.

58.8%
Best benchmark success rate

Across 9 evaluated agents on the 256-task test set, Gemini 3.1 Pro achieves the highest overall success rate at 58.8%.

9.4%
Worst benchmark success rate

Across the same 9-agent evaluation, Qwen3-VL-4B-Instruct is the lowest at 9.4% success rate.

What the paper found

MobileGym introduces a browser-hosted Android-like simulation platform that makes everyday mobile GUI tasks both verifiable and massively parallel without replicating proprietary backends. Its core novelty is a layered structured-JSON state model in which app data, OS state, and device context are readable, writable, snapshotable, and forkable, enabling deterministic state-diff judging, side-effect detection, and exact reset from any snapshot. A single machine can host hundreds of instances at roughly 400 MB each with about 3 s cold start, compared with gigabyte-scale emulator overhead. The accompanying MobileGym-Bench contains 416 parameterized task templates across 28 apps, including 12 everyday apps and 16 system apps, with a typed AnswerSheet protocol that replaces brittle free-text answer matching and yields deterministic success, progress, false-completion, overdue, and unexpected-side-effect metrics. Across 9 evaluated agents, success rate ranges from 9.4% for Qwen3-VL-4B-Instruct to 58.8% for Gemini 3.1 Pro, while a GRPO fine-tuning run on Qwen3-VL-4B-Instruct improves test success from 9.4% to 22.2%, a gain of 12.8 percentage points. In a sim-to-real study on 59 real-device signal tasks, the same training retains 95.1% of the simulation-side gain, rising from 32.2% to 72.9% on device. The paper also reports 10.2% VLM-judge misjudgment on audited trajectories, motivating programmatic verification as the central technical contribution.

Original abstract

We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: https://mobilegym.github.io.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis