Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers
AuthorsMarkus Frohmann, Mahdiyar Ali Akbar Alavi, Elizabeth Lingg, Navid Rekabsaz
AffiliationsThomson Reuters Labs · University of Toronto, Vector InstituteCorrespondence: markus.frohmann@mail.utoronto.ca
LLM judges may rank candidates equally well yet make different decisions when their order changes, and OC-SFT makes those decisions more consistent.
Key results
Maximum nDCG@10 difference among matched trained scorers.
Lowest overlap across permutations among the compared trained scorers.
Overlap achieved by OC-SFT across reordered candidate sets.
Order instability on Qwen3-4B passage reranking.
Multi-document QA answer-flip rate for OC-SFT.
What the paper found
This paper shows that an LLM scorer can achieve essentially the same ranking quality while producing different downstream decisions when candidate documents are reordered. In passage reranking, five trained scorers separated by only 0.010 nDCG@10 produced retained-set overlaps ranging from 0.656 to 0.835 across 10 random permutations, exposing a gap between ranking metrics and threshold, reader, or preference-training outcomes. The proposed remedy, order-consistency supervised fine-tuning, or OC-SFT, adds a consistency penalty between a candidate’s scores under two shuffled views while preserving the relevance objective. On Qwen3-4B across 18 reranking collections, OC-SFT reduced τ-PSI order instability from 0.209 to 0.083, without sacrificing ranking quality, and achieved a multi-document QA answer-flip rate of 0.125, outperforming competing order-aware objectives. It also remained more stable across 12 base models than order-averaged distillation, which requires 10 teacher permutations; OC-SFT needs one teacher pass and serves one permutation. Inference-time batched self-consistency improves stability but multiplies serving cost, whereas OC-SFT amortizes the intervention during training. The effect extends beyond reranking to HotpotQA, MuSiQue, and response-ranking benchmarks trained on UltraFeedback. GPT-5.4 achieved strong nDCG@10, but OC-SFT produced more reproducible retained sets, illustrating why evaluations of systems such as Qwen3, Gemma, or GPT-5.4 should report decision stability alongside ranking quality.
Original abstract
In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.