LLM Evaluators are Biased across Languages
AuthorsEj Zhou, Lucas Resck, Zheng Hui, Anna Korhonen
Resources
Multilingual LLM judges can score identical content very differently across languages, allowing harmful material in lower-resource languages to slip past safety filters despite impressive benchmark accuracy.
Key results
Semantically parallel instruction–response pairs were evaluated across 23 languages using RewardBench and M-RewardBench.
Equivalent content showed approximately a 0.5-point shift on the 1–5 Likert scale.
Reward-model scores had a Spearman -0.81 correlation with language resource availability.
A global threshold produced acceptance-rate differences of up to 43 percentage points across languages.
Per-language score offsets reduced cross-language acceptance-rate disparities by 60.9 percent on average.
What the paper found
Researchers at the University of Cambridge show that multilingual LLM evaluators do not use a language-neutral scoring scale. Testing semantically identical instruction–response pairs in 23 languages on RewardBench and its professionally translated extension M-RewardBench, they evaluated eight systems: four LLM-as-a-Judge models, including Aya Expanse 32B, Qwen 2.5 72B, LLaMA 3.1 70B Instruct, and M-Prometheus 14B, plus four reward models. Scores shifted by about 0.5 points on the 1–5 Likert scale, with lower-resource languages generally scored more generously; for reward models, this pattern correlated with Common Crawl resource level at Spearman -0.81. The bias remained statistically significant across architectures and persisted in frontier judges, including OpenAI's GPT-4.1-mini and Qwen3-32B thinking. Crucially, pairwise accuracy stayed above 90 percent, yet a single global threshold produced acceptance-rate differences of up to 43 percentage points, potentially allowing harmful content in lower-resource languages to evade safety filters and enabling cross-lingual reward hacking in RLHF. The authors link inflated scores to model uncertainty using summed negative log-likelihood, alternative uncertainty measures, and semantic entropy, but regression tests show that language identity remains predictive after uncertainty is controlled. Per-language offset correction reduced acceptance-rate disparities by 60.9 percent on average, although code-switched prompts can defeat language identification and misapply thresholds. The paper therefore recommends multilingual calibration, language-balanced preference data, and reporting cross-language score consistency alongside pairwise accuracy.
Original abstract
LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across languages under a global decision threshold, meaning, for instance, that harmful content in lower-resource languages is more likely to pass safety filters. Per-language thresholds would require language identification, which can be defeated by code-switched prompts. We then investigate why lower-resource languages receive higher rather than lower scores, and we find that model uncertainty is linked with the effect: models tend to give higher scores when less confident, both under negative log-likelihood and under token-free uncertainty measures; however, language identity remains a significant predictor after controlling for uncertainty, and the bias cannot be explained away by content difficulty alone, but is a structural, language-level misalignment.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.