OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
AuthorsRui Yang, Qianhui Wu, Yuxi Chen, Hao Bai, Wenlin Yao, Hao Cheng, Baolin Peng, Huan Zhang, Tong Zhang, Jianfeng Gao
Resources
OpenWebRL shows that visual web agents can be trained directly on live websites with online multi-turn RL, delivering a strong open-source system that rivals proprietary agents.
Key results
Warm-start data used before online RL
Open-ended live-web tasks used for MM-GRPO
OpenWebRL-4B official success rate
OpenWebRL-4B official success rate
OpenWebRL-4B official success rate
Approximate cost per training experiment
What the paper found
OpenWebRL, from UIUC and Microsoft, reframes visual web-agent training as end-to-end online multi-turn reinforcement learning on live websites, rather than imitation from large static trajectory corpora. The framework combines a fault-tolerant browser stack, a 13-tool action interface, multimodal context management that can retain only the current screenshot, and trajectory-level judging with MM-GRPO. Using Qwen3-VL-4B-Thinking, the authors warm-start from just 0.4K curated trajectories, then train on 2.2K open-web RL tasks; the resulting OpenWebRL-4B reaches 74.1% on WebVoyager, 67.0% on Online-Mind2Web, and 64.0% on DeepShop, for a 68.4% average and a new open-source state of the art. The paper shows that supervised warm-starting matters, but online RL is the main gain source: 4B performance rises from 39.3% at the base model to 52.0% after SFT and 68.4% after MM-GRPO. It also distills an 8B judge from 12.5K labeled rollouts, matching GPT-4.1-level training quality while avoiding about $545.5 in judge API cost per experiment, and reports that the distilled judge-8B reaches 89.8% accuracy and 92.1% F1 on held-out trajectory evaluation.
Original abstract
Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain largely proprietary, while open agents still depend heavily on supervised post-training over large collections of curated web trajectories. This dependence creates a major scalability bottleneck: high-quality demonstrations are expensive to collect, and static datasets offer limited coverage of the diverse, ever-changing open web. Although online RL has shown promise for text-based agents, its potential for training visual web agents directly on live websites remains largely underexplored. In this paper, we introduce OpenWebRL, an open framework for training visual web agents with online multi-turn RL on real websites. OpenWebRL covers the full training pipeline, including scalable live-browser infrastructure, supervised initialization, multimodal context management, trajectory-level success judging, and efficient multi-turn policy optimization. Using this framework, we train OpenWebRL-4B, which establishes a new open-source state of the art on challenging live-web benchmarks. With only 0.4K initialization trajectories and 2.2K open-ended RL training tasks, OpenWebRL-4B achieves 67.0% success on Online-Mind2Web and 64.0% on DeepShop, outperforming prior open agents of similar or larger scale and remaining competitive with proprietary systems including OpenAI CUA and Gemini CUA. Beyond strong benchmark performance, we systematically study the key design choices that make online RL effective for visual web agents, and analyze how RL improves agentic reasoning. Overall, our work offers a practical path toward building more capable, reproducible, and cost-efficient open web agents. We will release our training data, models, and code to support future research.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.