NTH

Small Language Models as Judges for Rubric-Based Reinforcement Learning

AuthorsFengyu Xie, Yilun Zhao, Bingsen Chen, Arman Cohan, Chen Zhao

September 4, 2026 2 min read
Watch on YouTube
The one-line take

This work shows that tiny language models can act as fast, effective rubric judges, potentially making reward-based training far cheaper than relying on large LLM evaluators.

Key results

1042
PointRubric questions

Size of the constructed pointwise rubric benchmark

25008
PointRubric labels

Response–criterion satisfaction labels

0.835
Qwen3-1.7B Probe macro-F1

RaR-Science-Static criterion-level macro-F1

0.643
Probe RL score

Final RaR-Science rubric score using the Qwen3-1.7B Probe reward

0.594
Generative RL baseline

Final RaR-Science rubric score using the Qwen3-8B Generative reward

10.7
Judge-time reduction

Probe requires 10.7 times less cumulative reward-judge time

What the paper found

This paper tests whether small language models can replace expensive generative judges in rubric-based reinforcement learning, where responses receive rewards for satisfying weighted, instance-specific criteria rather than exact-answer checks. It introduces PointRubric, containing 1042 questions and 25008 response–criterion labels, and RaR-Science-Static, a science-domain benchmark derived from RaR-Science. Using Qwen3 backbones from 0.6B to 8B parameters, the study compares three readouts: generated Yes/No verdicts, Yes/No Logprob margins, and Probe judges that train a lightweight linear classifier on frozen hidden states, with OpenAI's GPT-4o providing operational reference labels. On RaR-Science-Static, the Qwen3-1.7B Probe reaches 0.835 criterion-level macro-F1, versus 0.443 for Generative and 0.449 for Logprob on the same backbone, indicating that evaluative information is encoded more reliably in representations than in generated outputs. As a GRPO reward model, the Qwen3-1.7B Probe raises the RaR-Science rubric score from 0.232 to 0.643, outperforming the 0.594 achieved by a Qwen3-8B Generative judge. The Probe requires 10.7 times less cumulative reward-judge time, while its trained policy also improves GPQA-Diamond accuracy from 0.335 to 0.388. Cross-domain experiments show transfer from science to medicine, but the authors caution that GPT-4o supervision is an operational target rather than definitive human ground truth. The central result is that frozen small-model representations, paired with a simple probe, can provide efficient, reusable rewards for open-ended RL.

Original abstract

Rubric-based reinforcement learning extends RL beyond tasks with exact answers or rule-based verifiers by scoring responses against instance-specific criteria. However, this makes reward computation expensive: training requires repeated rubric judging, often with proprietary APIs or local generative LLM judges with 7B parameters or more. We study whether smaller language models can serve as efficient and reliable rubric-based judges. To make this question measurable, we construct PointRubric and RaR-Science-Static, two pointwise rubric-based evaluation datasets with instance-specific criteria and itemwise satisfaction labels. We compare three ways of extracting criterion-level judgments from small models: Generative verdicts, Yes/No Logprob margins, and Probe judges. Across both datasets, the Qwen3-1.7B Probe judge achieves the strongest criterion-level agreement among these methods, outperforming Generative and Logprob judges. Used as a GRPO reward model, it trains a policy from 0.232 to 0.643 on RaR-Science rubric score, compared with 0.594 for an 8B Generative judge baseline, while the baseline requires 10.7$\times$ more reward-judge time. Task and domain transfer experiments further suggest that Probe judges preserve criterion-level reward structure across settings.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →