LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
AuthorsYiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
Resources
LEGO-RL turns real coding-agent harnesses into trainable reinforcement-learning environments and improves SWE-bench performance across three popular platforms.
Key results
Final solve rate after training, up from 64.0%.
Final solve rate after training, up from 62.4%.
Final solve rate after training, up from 57.2%.
Probability correlation remained above this level across harnesses.
Number of training tasks used in the production runs.
What the paper found
LEGO-RL is a harness-native reinforcement-learning framework for coding agents that preserves the original control flow of systems such as Anthropic’s Claude Code and OpenHands SDK instead of rewriting them around an RL-specific interface. Its in-process proxy captures exact token IDs, response masks, log-probabilities, and mixture-of-experts routing decisions at the model API boundary, then uses R3 routing replay and trainer-side recomputation to maintain policy fidelity even when harnesses compact or reserialize history. The framework combines GSPO sequence-level optimization, asynchronous rollout scheduling, isolated Docker or Kubernetes sandboxes, cached images, executable verification, reward-hacking defenses, and a Live UI for trajectory-level diagnosis. Using Qwen3.5-35B-A3B with an OpenAI-compatible or Anthropic API, LEGO-RL trains on a 2,699-task OpenSWE-derived index and improves SWE-bench Verified solve rates from 64.0% to 70.4% with OpenHands SDK, from 62.4% to 68.2% with Claude Code, and from 57.2% to 66.6% with OpenCode. Across all three harnesses, rollout–training probability correlation stays above 0.99, showing that captured behavior closely matches the probabilities used for policy updates. The results argue that reliable coding-agent RL depends not only on the optimizer, but also on execution isolation, reward integrity, exact trajectory capture, and observability that distinguishes infrastructure failures from policy failures.
Original abstract
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.