NTH

Rethinking the Multilingual Reasoning Gap with Layer Swap

AuthorsMaxence Lasbordes, Amélie Chatelain, Djamé Seddah

May 30, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that multilingual reasoning in LLMs can be improved by swapping in stronger mid-layers from an English model, revealing a shared reasoning core beneath language-specific outer layers.

Key results

500k
Corpus size per language

The translated long-CoT dataset was constructed with approximately 500k examples per language.

32k
Training context length

The multilingual reasoning corpus uses a 32k context window, with samples up to 32,768 tokens.

10B
SFT token budget per language

Matched native and English-pivoted specialists were supervised with roughly 10B tokens per language.

1.9-3.5%
Native reasoning gap

The average native-vs-English-pivoted accuracy gap across the five non-English languages shrank to 1.9–3.5%.

83-89%
Layer Swap gap closed

Layer Swap closed 83%–89% of the native reasoning gap in French and German.

What the paper found

Rethinking the Multilingual Reasoning Gap with Layer Swap argues that the apparent multilingual reasoning deficit in long-chain-of-thought LLMs is much smaller than previously reported when native-language and English-pivoted specialists are trained under matched supervision. Using Qwen/Qwen3-8B-Base, the authors build a 32K-context multilingual reasoning corpus of roughly 500k samples per language across English, French, German, Spanish, Chinese, and Swahili, translated from Dolci-Think-SFT-32B with component-wise chunking to preserve reasoning traces. After supervised fine-tuning at about 10B tokens per language, the native-vs-English-pivoted accuracy gap on MGSM-Rev2, Global-MMLU-Lite, GPQA-Diamond, AIME 24/25, and HumanEvalPlus falls to 1.9%–3.5% on average, with most of the remaining deficit concentrated in AIME 24/25. The key novelty is a weight-space analysis showing that language-specific fine-tuning updates are strongly aligned in the middle transformer blocks, especially layers 13–22, while diverging in early and late layers, suggesting a language-agnostic reasoning core surrounded by language-specific input and output processing. Exploiting this, the paper introduces Layer Swap, a training-free model-merging method that replaces the native specialist’s mid-stack layers with the English specialist’s corresponding layers. This preserves nearly 100% target-language chain-of-thought fidelity while closing 83%–89% of the gap in French and German, 60% in Swahili, 27% in Chinese, and fully matching the English-pivoted ceiling in Spanish. The method improves average accuracy on the target-language evaluation set, for example from 72.36% to 74.74% in French and from 66.98% to 69.09% in Swahili, without switching the reasoning trace back to English.

Original abstract

Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (\emph{native reasoning}) substantially degrades performance relative to allowing the model to reason in English before answering in the input language (\emph{English-pivoted reasoning}). However, most studies of this native reasoning gap rely on inference-time interventions or limited native-language training data. We revisit this comparison at a larger scale and under comparable supervision. We construct long multilingual reasoning datasets across six languages (English, French, German, Spanish, Chinese and Swahili); fine-tune specialists in both native and English-pivoted regimes on top of \texttt{Qwen/Qwen3-8B-Base}, and evaluate across mathematics, science, general knowledge, and code. In this setting, the average native reasoning gap shrinks to 1.9--3.5\% across the five non-English languages, considerably smaller than previously reported. Weight-space analysis of the native specialists reveals aligned fine-tuning updates in the middle layers and divergence in the outer layers. This points to a largely language-agnostic reasoning core surrounded by language-specific layers. Exploiting this structure, we introduce a Layer Swap: transferring the English specialist's stronger reasoning mid-layers into each native specialist, closing most of the native reasoning gap across the five non-English languages while preserving CoT in the target language. We release all models and datasets.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis