NTH

Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers

AuthorsMarkus Frohmann, Mahdiyar Ali Akbar Alavi, Elizabeth Lingg, Navid Rekabsaz

AffiliationsThomson Reuters Labs · University of Toronto, Vector InstituteCorrespondence: markus.frohmann@mail.utoronto.ca

October 7, 2026 2 min read
Watch on YouTube
The one-line take

LLM judges may rank candidates equally well yet make different decisions when their order changes, and OC-SFT makes those decisions more consistent.

Key results

0.010
Ranking-quality spread

Maximum nDCG@10 difference among matched trained scorers.

0.656
Single-order retained-set overlap

Lowest overlap across permutations among the compared trained scorers.

0.835
OC-SFT retained-set overlap

Overlap achieved by OC-SFT across reordered candidate sets.

0.083
OC-SFT tau-PSI

Order instability on Qwen3-4B passage reranking.

0.125
QA answer-flip rate

Multi-document QA answer-flip rate for OC-SFT.

What the paper found

This paper shows that an LLM scorer can achieve essentially the same ranking quality while producing different downstream decisions when candidate documents are reordered. In passage reranking, five trained scorers separated by only 0.010 nDCG@10 produced retained-set overlaps ranging from 0.656 to 0.835 across 10 random permutations, exposing a gap between ranking metrics and threshold, reader, or preference-training outcomes. The proposed remedy, order-consistency supervised fine-tuning, or OC-SFT, adds a consistency penalty between a candidate’s scores under two shuffled views while preserving the relevance objective. On Qwen3-4B across 18 reranking collections, OC-SFT reduced τ-PSI order instability from 0.209 to 0.083, without sacrificing ranking quality, and achieved a multi-document QA answer-flip rate of 0.125, outperforming competing order-aware objectives. It also remained more stable across 12 base models than order-averaged distillation, which requires 10 teacher permutations; OC-SFT needs one teacher pass and serves one permutation. Inference-time batched self-consistency improves stability but multiplies serving cost, whereas OC-SFT amortizes the intervention during training. The effect extends beyond reranking to HotpotQA, MuSiQue, and response-ranking benchmarks trained on UltraFeedback. GPT-5.4 achieved strong nDCG@10, but OC-SFT produced more reproducible retained sets, illustrating why evaluations of systems such as Qwen3, Gemma, or GPT-5.4 should report decision stability alongside ranking quality.

Original abstract

In passage reranking, response ranking and multi-document question answering, LLMs can score several candidate documents or responses together in one prompt, each still receiving its own score. Such scorers are selected on ranking quality, but their scores determine a decision: what a score threshold retains, a reader answers, or which chosen/rejected pair enters preference training. Because the candidates share that prompt, reordering them changes their scores. The same query over the same candidates should still yield the same decision. However, equal ranking quality does not imply equal decisions: on passage reranking, five trained scorers within 0.010 nDCG@10 retain sets that overlap by only 0.66-0.84 when reordered. No prompt-time change we test resolves that dependence: the only one that improves ranking quality does not measurably improve decision stability. We introduce order-consistency SFT (OC-SFT), which attenuates it in the weights by penalizing disagreement between a candidate's scores across orderings. It holds ranking quality and leads every decision-stability measure among trained scorers on all three tasks. It is also more stable on 12 base models than order-averaged distillation, which trains on labels averaged across permutations. One OC-SFT permutation retains sets that overlap more than ten averaged off-the-shelf permutations. A comparison of such scorers should therefore report what a threshold retains and a reader answers, not ranking quality alone. Code is available at https://github.com/thomsonreuters/presentation-dependence.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis