NTH
AI research

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

AuthorsGaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan, Yanfei Chen, Chun-Liang Li, Long T. Le, Rujun Han, George Lee, Hanghang Tong, Chen-Yu Lee, Tomas Pfister

May 18, 2026 2 min read
Watch on YouTube
The one-line take

RubricEM trains research agents to think in rubric-defined stages, turning feedback from long, messy tasks into reusable experience for better long-form reasoning.

Key results

1,400 RL steps
RL training steps

RubricEM-8B is trained with 1,400 RL steps and reaches its reported long-form benchmark gains after this budget.

55.5 average score
Average long-form score

Across HealthBench, ResearchQA, DeepResearchBench, and ResearchRubrics, RubricEM-8B (RL) achieves a 55.5 average score.

49.2
SFT average score

The structured SFT checkpoint starts at 49.2 average before RL improves it.

53.6
Prior RL baseline average

RubricEM-8B (RL) outperforms DR Tulu-8B-RL, which is reported at 53.6 average.

73.5 average
Out-of-distribution short-form average

On the short-form transfer benchmarks SimpleQA, 2WikiMultihopQA, WebWalker, and DeepSearchQA, RubricEM-8B (RL) reaches a 73.5 average.

What the paper found

RubricEM introduces a rubric-guided meta-RL recipe for deep research agents that operate beyond verifiable rewards, where outputs are long-form reports rather than exact answers. The core novelty is to treat rubrics as a shared interface across planning, judging, and memory: the Qwen3-8B agent first generates self-rubrics in a four-stage scaffold, Plan → Research → Review → Answer, then trains with Stage-Structured GRPO, which assigns causal stagewise credit from rubric scores instead of broadcasting only a terminal reward. A shared-backbone reflection meta-policy is trained in parallel to distill judged trajectories into reusable natural-language guidance stored in an agent rubric bank, enabling both within-episode refinement and cross-episode transfer. The system uses Gemini-3.1-Pro for SFT data generation, Gemini Flash as the rubric judge, Google Search grounding plus Semantic Scholar snippet search, and an asynchronous one-step-delayed reflection pipeline to avoid extra wall-clock bottlenecks. On four long-form benchmarks—HealthBench, ResearchQA, DeepResearchBench, and ResearchRubrics—RubricEM-8B reaches a 55.5 average score after 1,400 RL steps, improving from 49.2 at SFT and outperforming DR Tulu-8B-RL at 53.6, while approaching proprietary systems such as OpenAI Deep Research and Gemini Deep Research. The paper also reports strong out-of-distribution short-form transfer, with RubricEM-8B-RL reaching 73.5 average on SimpleQA, 2WikiMultihopQA, WebWalker, and DeepSearchQA.

Original abstract

Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-truth answers, their trajectories span many tool-augmented decisions, and standard post-training offers little mechanism for turning past attempts into reusable experience. In this work, we argue that rubrics should serve not merely as final-answer evaluators, but as the shared interface that structures policy execution, judge feedback, and agent memory. Based on this view, we introduce RubricEM, a rubric-guided reinforcement learning framework that combines stagewise policy decomposition with reflection-based meta-policy evolution. RubricEM first makes research trajectories stage-aware by conditioning planning, evidence gathering, review, and synthesis on self-generated rubrics. It then assigns credit with Stage-Structured GRPO, which uses stagewise rubric judgments to provide denser semantic feedback for long-horizon optimization. In parallel, RubricEM trains a shared-backbone reflection meta-policy that distills judged trajectories into reusable rubric-grounded guidance for future attempts. The resulting RubricEM-8B achieves strong performance across four long-form research benchmarks, outperforming comparable open models and approaching proprietary deep-research systems. Beyond final performance, we perform thorough analyses to understand the key ingredients of RubricEM.

Read the original paper