NTH

OpenForgeRL: Train Harness-native Agents in Any Environment

AuthorsXiao Yu, Baolin Peng, Ruize Xu, Hao Zou, Qianhui Wu, Hao Cheng, Wenlin Yao, Nikhil Singh, Zhou Yu, Jianfeng Gao

July 27, 2026 2 min read
Watch on YouTube
The one-line take

OpenForgeRL makes it possible to train real-world tool-using and computer-use agents directly inside the complex harnesses where they operate.

Key results

31.7
ClawEval pass3

OpenForge-Claw score on ClawEval.

55.9
ClawEval pass@3

OpenForge-Claw score using three evaluation attempts.

33.7
QwenClawBench pass@1

OpenForge-Claw single-attempt success score.

37.7
OSWorld-Verified

OpenForge-GUI score for computer-use tasks.

63.0
Online-Mind2Web

OpenForge-GUI browser-use score.

72.3
WebVoyager

OpenForge-GUI browser-use score.

What the paper found

OpenForgeRL, from Columbia University, Dartmouth College, and Microsoft Research, is an open framework for end-to-end reinforcement learning of agents inside real inference harnesses such as Anthropic’s Claude Code, OpenAI’s Codex, and OpenClaw. Its lightweight proxy intercepts harness model calls, reconstructs prompt-response trajectories for standard RL systems such as veRL, and uses Kubernetes to launch each rollout in an isolated cloud container, eliminating the train–deploy mismatch for stateful, multi-process tool use and multimodal GUI control. With GRPO and automatically synthesized environments, a Qwen3-30B-A3B-Thinking agent trained on 343 Claw RL tasks achieved 31.7 pass3 on ClawEval, 55.9 pass@3 on ClawEval, and 33.7 pass@1 on QwenClawBench. A Qwen3-VL-8B GUI agent trained on 252 computer-use and 900 browser-use RL tasks reached 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager, despite using far less data than systems such as MolmoWeb. Analyses show that training across ZeroClaw, OpenClaw, and Codex transfers to unseen harnesses, while RL improves self-verification, specialized-tool coverage, and multi-step reliability; error recovery remains the weakest capability.

Original abstract

Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. OpenForgeRL achieves this with a lightweight proxy that serves the harness's model calls while recording them as training data for a standard RL codebase (e.g., veRL), and a Kubernetes orchestrator that runs each rollout in its own remote container, together enabling training on any harness in any environment at scale. By decoupling training and inference, OpenForgeRL allows researchers to easily train, study, and improve agents directly in the real harnesses and environments they are deployed with. We validate our framework across diverse, complex harnesses and environments, spanning tool/claw-based agents and multimodal GUI browser- and computer-use agents. Using only hundreds to a few thousand tasks, OpenForgeClaw reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval and 33.7 on QwenClawBench. OpenForgeGUI reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager. Both outperform open baselines of similar size on nearly all benchmarks, and in the GUI setting match or surpass models several times larger. Beyond benchmarks, we analyze how harness choice (e.g., ZeroClaw, OpenClaw, Codex) and RL shape agent behavior. We find that some harnesses are substantially harder to learn than others, and that RL improves agentic reliability, such as self-verification, tool coverage, and completing multi-step plans, though critical abilities such as error recovery remain weak.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis