NTH

LLM Evaluators are Biased across Languages

AuthorsEj Zhou, Lucas Resck, Zheng Hui, Anna Korhonen

July 17, 2026 3 min read
Watch on YouTube
The one-line take

Multilingual LLM judges can score identical content very differently across languages, allowing harmful material in lower-resource languages to slip past safety filters despite impressive benchmark accuracy.

Key results

23
Evaluation languages

Semantically parallel instruction–response pairs were evaluated across 23 languages using RewardBench and M-RewardBench.

0.5
Pointwise score shift

Equivalent content showed approximately a 0.5-point shift on the 1–5 Likert scale.

-0.81
Resource-score correlation

Reward-model scores had a Spearman -0.81 correlation with language resource availability.

43%
Acceptance-rate disparity

A global threshold produced acceptance-rate differences of up to 43 percentage points across languages.

60.9%
Offset-correction reduction

Per-language score offsets reduced cross-language acceptance-rate disparities by 60.9 percent on average.

What the paper found

Researchers at the University of Cambridge show that multilingual LLM evaluators do not use a language-neutral scoring scale. Testing semantically identical instruction–response pairs in 23 languages on RewardBench and its professionally translated extension M-RewardBench, they evaluated eight systems: four LLM-as-a-Judge models, including Aya Expanse 32B, Qwen 2.5 72B, LLaMA 3.1 70B Instruct, and M-Prometheus 14B, plus four reward models. Scores shifted by about 0.5 points on the 1–5 Likert scale, with lower-resource languages generally scored more generously; for reward models, this pattern correlated with Common Crawl resource level at Spearman -0.81. The bias remained statistically significant across architectures and persisted in frontier judges, including OpenAI's GPT-4.1-mini and Qwen3-32B thinking. Crucially, pairwise accuracy stayed above 90 percent, yet a single global threshold produced acceptance-rate differences of up to 43 percentage points, potentially allowing harmful content in lower-resource languages to evade safety filters and enabling cross-lingual reward hacking in RLHF. The authors link inflated scores to model uncertainty using summed negative log-likelihood, alternative uncertainty measures, and semantic entropy, but regression tests show that language identity remains predictive after uncertainty is controlled. Per-language offset correction reduced acceptance-rate disparities by 60.9 percent on average, although code-switched prompts can defeat language identification and misapply thresholds. The paper therefore recommends multilingual calibration, language-balanced preference data, and reporting cross-language score consistency alongside pairwise accuracy.

Original abstract

LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across languages under a global decision threshold, meaning, for instance, that harmful content in lower-resource languages is more likely to pass safety filters. Per-language thresholds would require language identification, which can be defeated by code-switched prompts. We then investigate why lower-resource languages receive higher rather than lower scores, and we find that model uncertainty is linked with the effect: models tend to give higher scores when less confident, both under negative log-likelihood and under token-free uncertainty measures; however, language identity remains a significant predictor after controlling for uncertainty, and the bias cannot be explained away by content difficulty alone, but is a structural, language-level misalignment.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis