ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
AuthorsFanrui Zhang, Ruixue Ding, Qiang Zhang, Xi Chen, Boli Chen, Shihang Wang, Qiuchen Wang, Hongmin Zhan, Jinxin Bian, Li xingchao, Peijin Zheng, Hao cheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
Resources
ARISE-RL trains agents to improve themselves by generating increasingly challenging tasks and learning from fine-grained rubric-based feedback.
Key results
Number of expert-calibrated ECR-DeepResearch queries.
Total ECR-Travel queries across five balanced task types.
Average across four reported benchmarks using Qwen3.5-9B.
Rubric score rate for Qwen3.5-9B with ARISE-RL.
ECR-Travel training groups where memory coaching improved reward.
ECR-Travel score after removing RG-SED, versus 0.599 for the full system.
What the paper found
ARISE-RL is a closed-loop reinforcement-learning framework for open-ended agents, designed for tasks such as deep research and travel planning where no single gold answer exists. Its Generator creates tool-grounded queries and verifiable rubrics only after observing real tool outputs, while its Solver earns partial-credit and full-completion rewards for satisfying rubric items through multi-step reasoning and tool use. A difficulty-shaped reward uses multiple Solver attempts to target tasks near the capability boundary, and Reward-Gated Self-Evolution Distillation, or RG-SED, distills memory-based coaching from the same policy only when it improves empirical reward, avoiding blind imitation and teacher–student distribution mismatch. The accompanying ECR-Bench contains 100 ECR-DeepResearch queries and 500 ECR-Travel queries spanning five balanced planning tasks. On Qwen3.5-9B, three evolution cycles raise the average score across ResearchRubrics, ECR-DeepResearch, VitaBench, and ECR-Travel to 0.569, including 0.781 on ECR-DeepResearch, outperforming larger Qwen models and commercial systems such as Google DeepMind’s Gemini3-Pro, OpenAI’s GPT-5.2, and Anthropic’s Claude-4.6-Sonnet on the reported aggregate. The RG-SED gate activates for 72.1% of ECR-Travel training groups, while rejecting harmful or ambiguous coaching elsewhere, and removing RG-SED lowers the ECR-Travel score from 0.599 to 0.518. The results position rubric-mediated co-evolution as a scalable alternative to relying on static human-authored trajectories, although validation remains limited to 8B- and 9B-scale open-source backbones.
Original abstract
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.