NTH

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

AuthorsYiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai

August 28, 2026 2 min read
Watch on YouTube
The one-line take

LEGO-RL turns real coding-agent harnesses into trainable reinforcement-learning environments and improves SWE-bench performance across three popular platforms.

Key results

70.4%
OpenHands SDK SWE-bench Verified

Final solve rate after training, up from 64.0%.

68.2%
Claude Code SWE-bench Verified

Final solve rate after training, up from 62.4%.

66.6%
OpenCode SWE-bench Verified

Final solve rate after training, up from 57.2%.

0.99
Rollout-training correlation

Probability correlation remained above this level across harnesses.

2,699
OpenSWE-derived training index

Number of training tasks used in the production runs.

What the paper found

LEGO-RL is a harness-native reinforcement-learning framework for coding agents that preserves the original control flow of systems such as Anthropic’s Claude Code and OpenHands SDK instead of rewriting them around an RL-specific interface. Its in-process proxy captures exact token IDs, response masks, log-probabilities, and mixture-of-experts routing decisions at the model API boundary, then uses R3 routing replay and trainer-side recomputation to maintain policy fidelity even when harnesses compact or reserialize history. The framework combines GSPO sequence-level optimization, asynchronous rollout scheduling, isolated Docker or Kubernetes sandboxes, cached images, executable verification, reward-hacking defenses, and a Live UI for trajectory-level diagnosis. Using Qwen3.5-35B-A3B with an OpenAI-compatible or Anthropic API, LEGO-RL trains on a 2,699-task OpenSWE-derived index and improves SWE-bench Verified solve rates from 64.0% to 70.4% with OpenHands SDK, from 62.4% to 68.2% with Claude Code, and from 57.2% to 66.6% with OpenCode. Across all three harnesses, rollout–training probability correlation stays above 0.99, showing that captured behavior closely matches the probabilities used for policy updates. The results argue that reliable coding-agent RL depends not only on the optimizer, but also on execution isolation, reward integrity, exact trajectory capture, and observability that distinguishes infrastructure failures from policy failures.

Original abstract

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis