NTH

ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

AuthorsFanrui Zhang, Ruixue Ding, Qiang Zhang, Xi Chen, Boli Chen, Shihang Wang, Qiuchen Wang, Hongmin Zhan, Jinxin Bian, Li xingchao, Peijin Zheng, Hao cheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha

September 12, 2026 2 min read
Watch on YouTube
The one-line take

ARISE-RL trains agents to improve themselves by generating increasingly challenging tasks and learning from fine-grained rubric-based feedback.

Key results

100
ECR-Bench deep-research queries

Number of expert-calibrated ECR-DeepResearch queries.

500
ECR-Bench travel queries

Total ECR-Travel queries across five balanced task types.

0.569
ARISE-RL average score

Average across four reported benchmarks using Qwen3.5-9B.

0.781
ECR-DeepResearch score

Rubric score rate for Qwen3.5-9B with ARISE-RL.

72.1%
RG-SED gate activation

ECR-Travel training groups where memory coaching improved reward.

0.518
RG-SED ablation ECR-Travel

ECR-Travel score after removing RG-SED, versus 0.599 for the full system.

What the paper found

ARISE-RL is a closed-loop reinforcement-learning framework for open-ended agents, designed for tasks such as deep research and travel planning where no single gold answer exists. Its Generator creates tool-grounded queries and verifiable rubrics only after observing real tool outputs, while its Solver earns partial-credit and full-completion rewards for satisfying rubric items through multi-step reasoning and tool use. A difficulty-shaped reward uses multiple Solver attempts to target tasks near the capability boundary, and Reward-Gated Self-Evolution Distillation, or RG-SED, distills memory-based coaching from the same policy only when it improves empirical reward, avoiding blind imitation and teacher–student distribution mismatch. The accompanying ECR-Bench contains 100 ECR-DeepResearch queries and 500 ECR-Travel queries spanning five balanced planning tasks. On Qwen3.5-9B, three evolution cycles raise the average score across ResearchRubrics, ECR-DeepResearch, VitaBench, and ECR-Travel to 0.569, including 0.781 on ECR-DeepResearch, outperforming larger Qwen models and commercial systems such as Google DeepMind’s Gemini3-Pro, OpenAI’s GPT-5.2, and Anthropic’s Claude-4.6-Sonnet on the reported aggregate. The RG-SED gate activates for 72.1% of ECR-Travel training groups, while rejecting harmful or ambiguous coaching elsewhere, and removing RG-SED lowers the ECR-Travel score from 0.599 to 0.518. The results position rubric-mediated co-evolution as a scalable alternative to relying on static human-authored trajectories, although validation remains limited to 8B- and 9B-scale open-source backbones.

Original abstract

Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →