Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
AuthorsAyoub Kirouane, Christos Petrocheilos
Resources
The study shows that SFT teaches models to reason visibly in Greek, while reinforcement learning fixes formatting and leakage problems that accuracy benchmarks fail to detect.
Key results
Changing only the random seed shifted benchmark results by 7.7 points.
Base models produced Greek reasoning on 0 of 1,000 Greek prompts.
SFT checkpoints reasoned in Greek on about 98% of measured traces.
The Qwen fine-tune used 3-fold fewer tokens than its base.
RLVR reduced answer-format fallback from 24.1% to 2.5% and leakage from 3.53% to 0.00%.
Explicit English-reasoning compliance improved by 9.1 percentage points.
What the paper found
This study fine-tunes four sparse mixture-of-experts models—Alibaba’s Qwen3.6-35B-A3B, OpenAI’s Gpt-OSS-20B, and NVIDIA’s NemotronH variants—with LoRA to reason in Greek, using 118,092 Greek training rows and a 5,156-item evaluation suite spanning mathematics, commonsense, and logic. Accuracy barely changes: the best configuration scores 76.5 versus 77.2 for its base, while changing only the random seed shifts results by 7.7 points, larger than every corpus or recipe effect tested. Behavioural metrics reveal the real transformation: base models produce Greek reasoning on 0 of 1,000 Greek prompts, whereas SFT checkpoints reach about 98% Greek-trace fidelity, eliminate in-question language switching, improve grammaticality, and substantially shorten reasoning; Qwen uses 3-fold fewer tokens, while the effect is family-dependent because Greek tokenization costs 2.3–2.5 times more tokens per word than English. SFT also introduces format failures and channel leakage, which accuracy obscures. A pre-registered GRPO-based reinforcement-learning-with-verifiable-rewards experiment repairs answer-format fallback from 24.1% to 2.5% and answer-channel leakage from 3.53% to 0.00%, while improving compliance with an explicit instruction to reason in English by 9.1 percentage points; Greek fidelity remains 98.2%. The paper therefore argues that low-resource reasoning models should be evaluated on language fidelity, reasoning budget, termination, format compliance, and steerability—not accuracy alone, echoing but extending DeepSeek-R1’s use of language-consistency rewards.
Original abstract
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.