EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
AuthorsGuhong Chen, Yingcheng Shi, Yongbin Li, Binhua Li, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye
Resources
EvoTrainer is an autonomous RL system that not only improves LLM policies but also evolves the training process itself to better diagnose and fix failures in agentic tasks.
Key results
EvoTrainer final score on SWE-9B
Human reference on SWE-9B
Absolute BC% gain on SWE-9B
EvoTrainer aggregate Math score
Absolute Avg@8 gain on Math
EvoTrainer final Coding score
What the paper found
EvoTrainer, developed by researchers at Shenzhen Institutes of Advanced Technology, Tongyi Lab, and Alibaba Group, reframes autonomous LLM reinforcement learning as co-evolution of policies and the training harness that diagnoses them. Instead of only searching over recipes, it uses version-controlled exploration, harness reflection, persistent case memory, and a reusable skill library to revise rewards, analyzers, filtering, and backtesting logic as training failures shift across domains. In experiments on BigMath-Hard, TACO-verified, and swe-rebench-v6, the framework uses a GRPO-style core with asymmetric Clip-Higher bounds and a behavior-sensitive SWE reward combining correctness, instruction following, search-before-edit, and edit-then-test signals. On SWE-9B, EvoTrainer reaches 38.16 BC%, beating the human-engineered RL reference at 33.77 BC% by +4.39 and the no-RL base at 30.19 by +7.97; on Math it reaches 79.49 Avg@8, up +2.88 over the human baseline, and on Coding it reaches 51.29 Avg@8. A key finding is that richer diagnostics matter: score-only iteration saturates at 33.33 BC% on SWE-9B, while harness evolution detects a Git-leak artifact that would have falsely promoted a 48.80 BC% branch, and a dead-group-aware filter plus an IF judge recovers useful variance. The paper argues that autonomous training should evolve the interpreter, not just the policy.
Original abstract
Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and scalar rewards mask diverse failure modes. We introduce EvoTrainer, an autonomous training framework that co-evolves LLM policies and training-side harnesses through empirical feedback: it diagnoses rollout-level evidence, revises diagnostics, backtests interventions, and accumulates reusable skills. Evaluated on mathematical reasoning, competitive-programming code generation, and repository-level software engineering, EvoTrainer matches or exceeds the human-engineered RL references under the same data, codebase, and evaluation protocol, with the largest gain on long-horizon agentic SWE. Trajectory analyses show that retained strategies diverge across domains, evolving diagnostics prevent invalid high-scoring branches from being promoted, and reusable skills shape later search. Autonomous LLM RL should move beyond recipe search toward joint evolution of policies and the training harnesses that interpret them.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.