NTH

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

AuthorsXiuYu Zhang, Yi Shan, Junfeng Fang, Zhenkai Liang

June 16, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that base LLMs already have hidden self-evaluation skills and that a small amount of targeted training can surface them to better predict how external judges will score their answers.

Key results

160
Unique training examples

SEE uses 160 unique HelpSteer2-derived prompts

31x
Data reduction vs baseline

SEE uses about 31× fewer unique examples than Adapted RLCR

0.7312
HelpSteer2 calibration

SEE calibration score on HelpSteer2 validation

0.6752
Baseline HelpSteer2 calibration

Adapted RLCR calibration score on HelpSteer2 validation

77.07%
Top-5 accuracy

Base model places the judge’s score in its top five score tokens on HelpSteer2 validation

What the paper found

Self Evaluation Is Already There argues that judge-aligned self-assessment in base LLMs is largely latent rather than learned from scratch: Qwen3-4B-Base, prompted in a HelpSteer2-style five-attribute format, already predicts an external judge’s scores well above chance before any task-specific training, with calibration around 0.50–0.70 across benchmarks and top-5 score-token localization of 77.07% on HelpSteer2 validation. The paper’s Self-Evaluation Elicitation (SEE) method surfaces this ability with a short two-phase cycle: Calibration-Coupled RL using GRPO optimizes both response quality and score agreement under a nonlinear calibration reward, then Masked Judge Distillation fine-tunes only the [SELF_EVAL] tokens against judge scores while leaving the answer untouched. Using just 160 unique training examples over 15 cycles, SEE reaches a HelpSteer2 calibration score of 0.7312 versus 0.6752 for an Adapted RLCR baseline that uses roughly 31× more unique data, and it improves open-ended evaluation on LC AlpacaEval 2.0, Arena-Hard-Auto v2.0, and WildBench v2, with calibration rising to 0.746, 0.609, and 0.609 respectively. The key technical result is that the gains transfer to held-out judges, including Claude Sonnet 4.6 and Gemini 3.1 Flash-Lite, showing that SEE elicits a stable notion of quality rather than memorizing GPT-5.4’s preferences.

Original abstract

Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an external judge's multi-attribute quality scores on open-ended responses well above chance across three benchmarks. We introduce Self-Evaluation Elicitation (SEE), a method that surfaces this latent ability through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that sharpens the prediction while leaving the answer untouched. From 160 unique examples, roughly 31x fewer than a reinforcement learning baseline, SEE improves held-out calibration across three benchmarks while preserving answer quality. The elicited self-evaluation is sharply localized within the model's own token distribution and stable across judges it was never trained against, indicating a transferable notion of quality rather than a single judge's preference. These results reframe judge-aligned self-evaluation as a problem of elicitation rather than acquisition.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis