NTH

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

AuthorsJinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

AffiliationsLLM Core, Xiaomi · Renmin University of China · Peking University · University of Hong Kong

October 4, 2026 2 min read
Watch on YouTube
The one-line take

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Key results

62.2%
DeepSWE step-28 pass rate

MiMo-V2.6-Flash with Gagar at RL step 28

63.4%
DeepSWE peak pass rate

MiMo-V2.6-Flash with Gagar at RL step 44

15.6%
Interaction-turn reduction

Reduction on DeepSWE v1.1 at the shared step-28 checkpoint

9.9%
Token-length reduction

Reduction on DeepSWE v1.1 at the shared step-28 checkpoint

69.8%
Passing-solution average win rate

Blinded groupwise quality evaluation on DeepSWE v1.1

71.9%
MiMo-V2.6-Pro DeepSWE score

DeepSWE v1.1 avg@3 after mixed-task RL with Gagar

What the paper found

Gagar, or Groupwise Agentic Grading for Advantage Redistribution, addresses a weakness in GRPO-based code-agent reinforcement learning: executable tests provide binary rewards, so every passing trajectory in a rollout group receives identical credit even when implementations differ in precision, minimality, side effects, and codebase consistency. Developed for Xiaomi’s MiMo-V2.6 models, Gagar uses dynamic sampling to retain mixed pass-fail groups, then gives an SFT-trained agentic grader a shared workspace containing the task, repository, patches, full trajectories, and test results. The grader can inspect files and run targeted checks, rank passing solutions into quality tiers, downweight weaker candidates, and proportionally rescale all positive advantages so their original sum is preserved while failed-trajectory advantages remain unchanged. On DeepSWE v1.1, MiMo-V2.6-Flash reached 62.2% pass rate at training step 28, a 12.1 percentage-point improvement over binary-reward training, and peaked at 63.4% at step 44; Gagar also reduced interaction turns by 15.6% and token length by 9.9% at the shared step-28 checkpoint. A blinded evaluation found Gagar-generated passing solutions had a 69.8% average win rate. In mixed-task RL, the 1.02T-parameter MiMo-V2.6-Pro reached 71.9% on DeepSWE v1.1 and 62.7% on SWE-bench Pro, competitive with external systems including Claude Opus 5 and GPT-5.6 Sol. The ablation shows that redistribution, rather than simple quality-based downweighting, is central to stable training.

Original abstract

Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

Read the original paper

More in Code Generation

Browse all 43 papers →
03Code Generation

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo

CodeMidas turns existing codebases into scalable, automatically verified RL environments that train coding agents to perform better across diverse software tasks.

Read analysis