NTH

HealMed: Multilingual Evaluation of Large Language Models in Medicine

AuthorsYingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, Aosong Feng, Kevin W. Jin, Xing Wu, Jinghui Lu, Abdul Samad, Akbar Faruqi, Cesar Caraballo, Cibele Brandão, Dhruva, Gupta, Eunji Jeon, Gabriel Madera-Santiago, Geon Lee, Hugo Toshio Itikawa, Insook Cho, Isabelli Martins, Isarar Siddique, Israr Ahmed, Jihyo Kwak, Kanyakorn Veerakanjana, Luis Guilherme Cardoso, Minjin Kim, Piyalitt Ittichaiwong, Renee Dua, Santiago Gudiño-Rosales, Xiujie Chen, Zeo Lapalus, Zixin Xu, Michihiro Yasunaga, Rex Ying, Heuiseok Lim, Jaewoo Kang, Chanjun Park, Hang Jiang, Ethan Goh, Hyunjae Kim, Edison Marrese-Taylor, Yusuke Iwasawa, Yutaka Matsuo, Qingyu Chen, Irene Li

August 30, 2026 2 min read
Watch on YouTube
The one-line take

HealMed tests whether medical language models remain reliable across languages and shows that low-resource performance and translation quality can substantially change the results.

Key results

1,000
Examples per language

Aligned medical benchmark examples provided in each of nine languages.

14
Models evaluated

LLMs assessed on MCQA and NLI, including proprietary, open-source, and medically specialized models.

28.9%
Zulu accuracy decline

Accuracy shift relative to English across MCQA and NLI.

59.7%
Zulu low-scoring QA responses

Open-ended responses scoring below 3 on the five-point evaluation rubric.

5.8%
Zulu expert-review shift

Mean absolute English-adjusted MCQA and NLI shift between machine-translated and expert-reviewed data.

What the paper found

HealMed is an expert-reviewed multilingual medical benchmark containing 1,000 aligned examples in each of nine languages, drawn from nine datasets and spanning multiple-choice question answering, natural language inference, and open-ended QA. Each translated instance underwent two-stage review by bilingual medical experts, enabling direct comparison between machine-translated and clinically revised data. Across 14 LLMs, OpenAI’s GPT-5.4 and Google DeepMind’s Gemini-3-Flash were the most accurate and multilingual-stable systems, while Anthropic’s Claude models also maintained relatively small resource gaps. Open-source families including DeepSeek, Qwen, LLaMA, and Gemma, plus specialized systems such as MedGemma and HuatuoGPT-o1, generally degraded more in lower-resource languages; medical specialization alone did not ensure robustness. The largest multiple-choice and NLI accuracy decline occurred in Zulu, at 28.9 percentage points relative to English, compared with 15.4 percentage points in Swahili, while Thai remained closer to higher-resource languages. In open-ended QA, 59.7% of Zulu responses scored below 3 on the five-point LLM-as-judge rubric. Expert revision materially altered measured performance: adjusted MCQA and NLI shifts reached 5.8 percentage points in Zulu, showing that translation artifacts can distort model rankings and apparent language gaps. HealMed’s central contribution is therefore not only broader multilingual testing, but an auditable evaluation pipeline that separates model capability from translation quality.

Original abstract

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis