NTH

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

AuthorsTao Chen, Gangwei Jiang, Pengyu Cheng, Siyuan Huang, Yihao Liu, Jingwei Ni, Jiaqi Guo, Mengyu Zhou, Kai Tang, Junling Liu, Qinliang Su, Xiaoxi Jiang, Guanjun Jiang

June 10, 2026 2 min read
Watch on YouTube
The one-line take

Skill-RM turns reward evaluation into an agent-like skill that dynamically combines rules, references, and rubrics to judge LLM outputs more consistently.

Key results

83.9
Qwen3.5-27B avg baseline

LLM-as-a-Judge baseline average on RewardBench2, RM-Bench, and JudgeBench

86.2
Skill-RM avg

Matched Qwen3.5-27B Skill-RM average on RewardBench2, RM-Bench, and JudgeBench

89.1
Skill-RM + sample-spec avg

Skill-RM with sample-specific resources mounted under Qwen3.5-27B

81.0
Appended-resources avg

Resource-use ablation where resources are directly appended to the prompt

97.8
GSM8K best-of-10

Skill-RM selection accuracy on the JETTS GSM8K pool

What the paper found

Skill-RM, developed by the Qwen Large Model Application Team at Alibaba, reframes reward modeling as an executable Agent Skill rather than a static scalar judge. The framework packages a Reward-Evaluation Skill with a frozen resource bank of rubrics, references, checklists, verifiers, and aggregation rules, then lets a Qwen3.5 judge dynamically retrieve and execute only the evidence relevant to each input. On RewardBench2, RM-Bench, and JudgeBench, the matched Qwen3.5-27B baseline averages 83.9, while Skill-RM raises that to 86.2, and Skill-RM with sample-specific resources reaches 89.1; the stronger Qwen3.5-122B-A10B backbone reaches 86.0 with sample-specific resources. The mechanism study is important: directly appending resources drops the average to 81.0, showing the gain comes from skill-mediated orchestration rather than extra context. In best-of-10 selection on JETTS pools, Skill-RM is nearly saturated on GSM8K at 97.8, while showing clearer gains on IFEval and HumanEval+. For instruction-following reward use, Skill-RM achieves the highest IF-RewardBench Kendall correlation at 0.524 and improves downstream GRPO training to 45.9 average on IFEval, IFBench, and AdvancedIF, outperforming VerIF at 44.7. Overall, the paper’s novelty is the explicit, evidence-bearing reward trace that unifies heterogeneous evaluation criteria into a reproducible reward computation pipeline.

Original abstract

Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) pipelines. However, current reward evaluation relies on heterogeneous criteria such as rule-based verifiers, ground-truth references, procedural checklists, and complex rubrics, where a unified mechanism to integrate all types of evidence remains unexplored. To this end, we propose Skill Reward Model (Skill-RM), a unified framework that reformulates reward modeling as the execution of a reusable Reward-Evaluation Skill. By treating reward computation as a structured agentic task, Skill-RM provides a consistent interface to orchestrate heterogeneous resources, dynamically selecting and aggregating evidence tailored to the specific requirements of each input. This approach enables the reward model to move beyond static evaluation, ensuring consistency and transparency across diverse tasks. Extensive experiments on reward benchmarks and downstream applications, including best-of-N selection and reinforcement learning, demonstrate that Skill-RM consistently outperforms traditional judge baselines. Our findings suggest that Skill-RM not only provides a unified solution for reward modeling but also achieves superior performance through the strategic and dynamic orchestration of evidence. The code is at https://github.com/Qwen-Applications/Skill-RM.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis