RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
AuthorsGaotang Li, Bhavana Dalvi Mishra, Zifeng Wang, Jun Yan, Yanfei Chen, Chun-Liang Li, Long T. Le, Rujun Han, George Lee, Hanghang Tong, Chen-Yu Lee, Tomas Pfister
Resources
RubricEM trains research agents to think in rubric-defined stages, turning feedback from long, messy tasks into reusable experience for better long-form reasoning.
Key results
RubricEM-8B is trained with 1,400 RL steps and reaches its reported long-form benchmark gains after this budget.
Across HealthBench, ResearchQA, DeepResearchBench, and ResearchRubrics, RubricEM-8B (RL) achieves a 55.5 average score.
The structured SFT checkpoint starts at 49.2 average before RL improves it.
RubricEM-8B (RL) outperforms DR Tulu-8B-RL, which is reported at 53.6 average.
On the short-form transfer benchmarks SimpleQA, 2WikiMultihopQA, WebWalker, and DeepSearchQA, RubricEM-8B (RL) reaches a 73.5 average.
What the paper found
RubricEM introduces a rubric-guided meta-RL recipe for deep research agents that operate beyond verifiable rewards, where outputs are long-form reports rather than exact answers. The core novelty is to treat rubrics as a shared interface across planning, judging, and memory: the Qwen3-8B agent first generates self-rubrics in a four-stage scaffold, Plan → Research → Review → Answer, then trains with Stage-Structured GRPO, which assigns causal stagewise credit from rubric scores instead of broadcasting only a terminal reward. A shared-backbone reflection meta-policy is trained in parallel to distill judged trajectories into reusable natural-language guidance stored in an agent rubric bank, enabling both within-episode refinement and cross-episode transfer. The system uses Gemini-3.1-Pro for SFT data generation, Gemini Flash as the rubric judge, Google Search grounding plus Semantic Scholar snippet search, and an asynchronous one-step-delayed reflection pipeline to avoid extra wall-clock bottlenecks. On four long-form benchmarks—HealthBench, ResearchQA, DeepResearchBench, and ResearchRubrics—RubricEM-8B reaches a 55.5 average score after 1,400 RL steps, improving from 49.2 at SFT and outperforming DR Tulu-8B-RL at 53.6, while approaching proprietary systems such as OpenAI Deep Research and Gemini Deep Research. The paper also reports strong out-of-distribution short-form transfer, with RubricEM-8B-RL reaching 73.5 average on SimpleQA, 2WikiMultihopQA, WebWalker, and DeepSearchQA.
Original abstract
Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-truth answers, their trajectories span many tool-augmented decisions, and standard post-training offers little mechanism for turning past attempts into reusable experience. In this work, we argue that rubrics should serve not merely as final-answer evaluators, but as the shared interface that structures policy execution, judge feedback, and agent memory. Based on this view, we introduce RubricEM, a rubric-guided reinforcement learning framework that combines stagewise policy decomposition with reflection-based meta-policy evolution. RubricEM first makes research trajectories stage-aware by conditioning planning, evidence gathering, review, and synthesis on self-generated rubrics. It then assigns credit with Stage-Structured GRPO, which uses stagewise rubric judgments to provide denser semantic feedback for long-horizon optimization. In parallel, RubricEM trains a shared-backbone reflection meta-policy that distills judged trajectories into reusable rubric-grounded guidance for future attempts. The resulting RubricEM-8B achieves strong performance across four long-form research benchmarks, outperforming comparable open models and approaching proprietary deep-research systems. Beyond final performance, we perform thorough analyses to understand the key ingredients of RubricEM.
Read the original paper