Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
AuthorsXiuYu Zhang, Yi Shan, Junfeng Fang, Zhenkai Liang
Resources
This paper shows that base LLMs already have hidden self-evaluation skills and that a small amount of targeted training can surface them to better predict how external judges will score their answers.
Key results
SEE uses 160 unique HelpSteer2-derived prompts
SEE uses about 31× fewer unique examples than Adapted RLCR
SEE calibration score on HelpSteer2 validation
Adapted RLCR calibration score on HelpSteer2 validation
Base model places the judge’s score in its top five score tokens on HelpSteer2 validation
What the paper found
Self Evaluation Is Already There argues that judge-aligned self-assessment in base LLMs is largely latent rather than learned from scratch: Qwen3-4B-Base, prompted in a HelpSteer2-style five-attribute format, already predicts an external judge’s scores well above chance before any task-specific training, with calibration around 0.50–0.70 across benchmarks and top-5 score-token localization of 77.07% on HelpSteer2 validation. The paper’s Self-Evaluation Elicitation (SEE) method surfaces this ability with a short two-phase cycle: Calibration-Coupled RL using GRPO optimizes both response quality and score agreement under a nonlinear calibration reward, then Masked Judge Distillation fine-tunes only the [SELF_EVAL] tokens against judge scores while leaving the answer untouched. Using just 160 unique training examples over 15 cycles, SEE reaches a HelpSteer2 calibration score of 0.7312 versus 0.6752 for an Adapted RLCR baseline that uses roughly 31× more unique data, and it improves open-ended evaluation on LC AlpacaEval 2.0, Arena-Hard-Auto v2.0, and WildBench v2, with calibration rising to 0.746, 0.609, and 0.609 respectively. The key technical result is that the gains transfer to held-out judges, including Claude Sonnet 4.6 and Gemini 3.1 Flash-Lite, showing that SEE elicits a stable notion of quality rather than memorizing GPT-5.4’s preferences.
Original abstract
Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted training: prompted few-shot, a base model already predicts an external judge's multi-attribute quality scores on open-ended responses well above chance across three benchmarks. We introduce Self-Evaluation Elicitation (SEE), a method that surfaces this latent ability through a short cycle comprising a calibration-coupled reinforcement learning phase that improves the answer and predicts the judge, followed by a masked distillation phase that sharpens the prediction while leaving the answer untouched. From 160 unique examples, roughly 31x fewer than a reinforcement learning baseline, SEE improves held-out calibration across three benchmarks while preserving answer quality. The elicited self-evaluation is sharply localized within the model's own token distribution and stable across judges it was never trained against, indicating a transferable notion of quality rather than a single judge's preference. These results reframe judge-aligned self-evaluation as a problem of elicitation rather than acquisition.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.