NTH

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

AuthorsGuhong Chen, Yingcheng Shi, Yongbin Li, Binhua Li, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye

June 14, 2026 2 min read
Watch on YouTube
The one-line take

EvoTrainer is an autonomous RL system that not only improves LLM policies but also evolves the training process itself to better diagnose and fix failures in agentic tasks.

Key results

38.16
SWE-9B BC%

EvoTrainer final score on SWE-9B

33.77
Human-engineered RL SWE-9B BC%

Human reference on SWE-9B

4.39
SWE-9B improvement over human

Absolute BC% gain on SWE-9B

79.49
Math Avg@8

EvoTrainer aggregate Math score

2.88
Math improvement over human

Absolute Avg@8 gain on Math

51.29
Coding Avg@8

EvoTrainer final Coding score

What the paper found

EvoTrainer, developed by researchers at Shenzhen Institutes of Advanced Technology, Tongyi Lab, and Alibaba Group, reframes autonomous LLM reinforcement learning as co-evolution of policies and the training harness that diagnoses them. Instead of only searching over recipes, it uses version-controlled exploration, harness reflection, persistent case memory, and a reusable skill library to revise rewards, analyzers, filtering, and backtesting logic as training failures shift across domains. In experiments on BigMath-Hard, TACO-verified, and swe-rebench-v6, the framework uses a GRPO-style core with asymmetric Clip-Higher bounds and a behavior-sensitive SWE reward combining correctness, instruction following, search-before-edit, and edit-then-test signals. On SWE-9B, EvoTrainer reaches 38.16 BC%, beating the human-engineered RL reference at 33.77 BC% by +4.39 and the no-RL base at 30.19 by +7.97; on Math it reaches 79.49 Avg@8, up +2.88 over the human baseline, and on Coding it reaches 51.29 Avg@8. A key finding is that richer diagnostics matter: score-only iteration saturates at 33.33 BC% on SWE-9B, while harness evolution detects a Git-leak artifact that would have falsely promoted a 48.80 BC% branch, and a dead-group-aware filter plus an IF judge recovers useful variance. The paper argues that autonomous training should evolve the interpreter, not just the policy.

Original abstract

Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and scalar rewards mask diverse failure modes. We introduce EvoTrainer, an autonomous training framework that co-evolves LLM policies and training-side harnesses through empirical feedback: it diagnoses rollout-level evidence, revises diagnostics, backtests interventions, and accumulates reusable skills. Evaluated on mathematical reasoning, competitive-programming code generation, and repository-level software engineering, EvoTrainer matches or exceeds the human-engineered RL references under the same data, codebase, and evaluation protocol, with the largest gain on long-horizon agentic SWE. Trajectory analyses show that retained strategies diverge across domains, evolving diagnostics prevent invalid high-scoring branches from being promoted, and reusable skills shape later search. Autonomous LLM RL should move beyond recipe search toward joint evolution of policies and the training harnesses that interpret them.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →