Small Language Models as Judges for Rubric-Based Reinforcement Learning
AuthorsFengyu Xie, Yilun Zhao, Bingsen Chen, Arman Cohan, Chen Zhao
Resources
This work shows that tiny language models can act as fast, effective rubric judges, potentially making reward-based training far cheaper than relying on large LLM evaluators.
Key results
Size of the constructed pointwise rubric benchmark
Response–criterion satisfaction labels
RaR-Science-Static criterion-level macro-F1
Final RaR-Science rubric score using the Qwen3-1.7B Probe reward
Final RaR-Science rubric score using the Qwen3-8B Generative reward
Probe requires 10.7 times less cumulative reward-judge time
What the paper found
This paper tests whether small language models can replace expensive generative judges in rubric-based reinforcement learning, where responses receive rewards for satisfying weighted, instance-specific criteria rather than exact-answer checks. It introduces PointRubric, containing 1042 questions and 25008 response–criterion labels, and RaR-Science-Static, a science-domain benchmark derived from RaR-Science. Using Qwen3 backbones from 0.6B to 8B parameters, the study compares three readouts: generated Yes/No verdicts, Yes/No Logprob margins, and Probe judges that train a lightweight linear classifier on frozen hidden states, with OpenAI's GPT-4o providing operational reference labels. On RaR-Science-Static, the Qwen3-1.7B Probe reaches 0.835 criterion-level macro-F1, versus 0.443 for Generative and 0.449 for Logprob on the same backbone, indicating that evaluative information is encoded more reliably in representations than in generated outputs. As a GRPO reward model, the Qwen3-1.7B Probe raises the RaR-Science rubric score from 0.232 to 0.643, outperforming the 0.594 achieved by a Qwen3-8B Generative judge. The Probe requires 10.7 times less cumulative reward-judge time, while its trained policy also improves GPQA-Diamond accuracy from 0.335 to 0.388. Cross-domain experiments show transfer from science to medicine, but the authors caution that GPT-4o supervision is an operational target rather than definitive human ground truth. The central result is that frozen small-model representations, paired with a simple probe, can provide efficient, reusable rewards for open-ended RL.
Original abstract
Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific criteria. However, this makes reward computation expensive: training requires repeated rubric judging, often with proprietary APIs or local generative LLM judges with 7B parameters or more. We study whether smaller language models can serve as efficient and reliable rubric-based judges. To make this question measurable, we construct PointRubric and RaR-Science-Static, two pointwise rubric-based evaluation datasets with instance-specific criteria and itemwise satisfaction labels. We compare three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges. Across both datasets, the Qwen3-1.7B Probe judge achieves the strongest criterion-level agreement among these methods, outperforming Generative and Logprob judges. Used as a GRPO reward model, it trains a policy from 0.232 to 0.643 on RaR-Science rubric score, compared with 0.594 for an 8B Generative judge baseline, while the baseline requires 10.7$\times$ more reward-judge time. Task and domain transfer experiments further suggest that Probe judges preserve criterion-level reward structure across settings.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.