Rethinking the Multilingual Reasoning Gap with Layer Swap
AuthorsMaxence Lasbordes, Amélie Chatelain, Djamé Seddah
Resources
This paper shows that multilingual reasoning in LLMs can be improved by swapping in stronger mid-layers from an English model, revealing a shared reasoning core beneath language-specific outer layers.
Key results
The translated long-CoT dataset was constructed with approximately 500k examples per language.
The multilingual reasoning corpus uses a 32k context window, with samples up to 32,768 tokens.
Matched native and English-pivoted specialists were supervised with roughly 10B tokens per language.
The average native-vs-English-pivoted accuracy gap across the five non-English languages shrank to 1.9–3.5%.
Layer Swap closed 83%–89% of the native reasoning gap in French and German.
What the paper found
Rethinking the Multilingual Reasoning Gap with Layer Swap argues that the apparent multilingual reasoning deficit in long-chain-of-thought LLMs is much smaller than previously reported when native-language and English-pivoted specialists are trained under matched supervision. Using Qwen/Qwen3-8B-Base, the authors build a 32K-context multilingual reasoning corpus of roughly 500k samples per language across English, French, German, Spanish, Chinese, and Swahili, translated from Dolci-Think-SFT-32B with component-wise chunking to preserve reasoning traces. After supervised fine-tuning at about 10B tokens per language, the native-vs-English-pivoted accuracy gap on MGSM-Rev2, Global-MMLU-Lite, GPQA-Diamond, AIME 24/25, and HumanEvalPlus falls to 1.9%–3.5% on average, with most of the remaining deficit concentrated in AIME 24/25. The key novelty is a weight-space analysis showing that language-specific fine-tuning updates are strongly aligned in the middle transformer blocks, especially layers 13–22, while diverging in early and late layers, suggesting a language-agnostic reasoning core surrounded by language-specific input and output processing. Exploiting this, the paper introduces Layer Swap, a training-free model-merging method that replaces the native specialist’s mid-stack layers with the English specialist’s corresponding layers. This preserves nearly 100% target-language chain-of-thought fidelity while closing 83%–89% of the gap in French and German, 60% in Swahili, 27% in Chinese, and fully matching the English-pivoted ceiling in Spanish. The method improves average accuracy on the target-language evaluation set, for example from 72.36% to 74.74% in French and from 66.98% to 69.09% in Swahili, without switching the reasoning trace back to English.
Original abstract
Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (\emph{native reasoning}) substantially degrades performance relative to allowing the model to reason in English before answering in the input language (\emph{English-pivoted reasoning}). However, most studies of this native reasoning gap rely on inference-time interventions or limited native-language training data. We revisit this comparison at a larger scale and under comparable supervision. We construct long multilingual reasoning datasets across six languages (English, French, German, Spanish, Chinese and Swahili); fine-tune specialists in both native and English-pivoted regimes on top of \texttt{Qwen/Qwen3-8B-Base}, and evaluate across mathematics, science, general knowledge, and code. In this setting, the average native reasoning gap shrinks to 1.9--3.5\% across the five non-English languages, considerably smaller than previously reported. Weight-space analysis of the native specialists reveals aligned fine-tuning updates in the middle layers and divergence in the outer layers. This points to a largely language-agnostic reasoning core surrounded by language-specific layers. Exploiting this structure, we introduce a Layer Swap: transferring the English specialist's stronger reasoning mid-layers into each native specialist, closing most of the native reasoning gap across the five non-English languages while preserving CoT in the target language. We release all models and datasets.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.