NTH

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

AuthorsRui Yang, Qianhui Wu, Yuxi Chen, Hao Bai, Wenlin Yao, Hao Cheng, Baolin Peng, Huan Zhang, Tong Zhang, Jianfeng Gao

June 15, 2026 2 min read
Watch on YouTube
The one-line take

OpenWebRL shows that visual web agents can be trained directly on live websites with online multi-turn RL, delivering a strong open-source system that rivals proprietary agents.

Key results

0.4K
SFT trajectories

Warm-start data used before online RL

2.2K
RL training tasks

Open-ended live-web tasks used for MM-GRPO

74.1%
WebVoyager success

OpenWebRL-4B official success rate

67.0%
Online-Mind2Web success

OpenWebRL-4B official success rate

64.0%
DeepShop success

OpenWebRL-4B official success rate

545.5
Judge API cost

Approximate cost per training experiment

What the paper found

OpenWebRL, from UIUC and Microsoft, reframes visual web-agent training as end-to-end online multi-turn reinforcement learning on live websites, rather than imitation from large static trajectory corpora. The framework combines a fault-tolerant browser stack, a 13-tool action interface, multimodal context management that can retain only the current screenshot, and trajectory-level judging with MM-GRPO. Using Qwen3-VL-4B-Thinking, the authors warm-start from just 0.4K curated trajectories, then train on 2.2K open-web RL tasks; the resulting OpenWebRL-4B reaches 74.1% on WebVoyager, 67.0% on Online-Mind2Web, and 64.0% on DeepShop, for a 68.4% average and a new open-source state of the art. The paper shows that supervised warm-starting matters, but online RL is the main gain source: 4B performance rises from 39.3% at the base model to 52.0% after SFT and 68.4% after MM-GRPO. It also distills an 8B judge from 12.5K labeled rollouts, matching GPT-4.1-level training quality while avoiding about $545.5 in judge API cost per experiment, and reports that the distilled judge-8B reaches 89.8% accuracy and 92.1% F1 on held-out trajectory evaluation.

Original abstract

Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain largely proprietary, while open agents still depend heavily on supervised post-training over large collections of curated web trajectories. This dependence creates a major scalability bottleneck: high-quality demonstrations are expensive to collect, and static datasets offer limited coverage of the diverse, ever-changing open web. Although online RL has shown promise for text-based agents, its potential for training visual web agents directly on live websites remains largely underexplored. In this paper, we introduce OpenWebRL, an open framework for training visual web agents with online multi-turn RL on real websites. OpenWebRL covers the full training pipeline, including scalable live-browser infrastructure, supervised initialization, multimodal context management, trajectory-level success judging, and efficient multi-turn policy optimization. Using this framework, we train OpenWebRL-4B, which establishes a new open-source state of the art on challenging live-web benchmarks. With only 0.4K initialization trajectories and 2.2K open-ended RL training tasks, OpenWebRL-4B achieves 67.0% success on Online-Mind2Web and 64.0% on DeepShop, outperforming prior open agents of similar or larger scale and remaining competitive with proprietary systems including OpenAI CUA and Gemini CUA. Beyond strong benchmark performance, we systematically study the key design choices that make online RL effective for visual web agents, and analyze how RL improves agentic reasoning. Overall, our work offers a practical path toward building more capable, reproducible, and cost-efficient open web agents. We will release our training data, models, and code to support future research.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis