NTH

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

AuthorsAyoub Kirouane, Christos Petrocheilos

August 25, 2026 3 min read
Watch on YouTube
The one-line take

The study shows that SFT teaches models to reason visibly in Greek, while reinforcement learning fixes formatting and leakage problems that accuracy benchmarks fail to detect.

Key results

7.7
Seed accuracy variance

Changing only the random seed shifted benchmark results by 7.7 points.

0
Base Greek reasoning fidelity

Base models produced Greek reasoning on 0 of 1,000 Greek prompts.

98%
SFT Greek reasoning fidelity

SFT checkpoints reasoned in Greek on about 98% of measured traces.

3
Qwen token reduction

The Qwen fine-tune used 3-fold fewer tokens than its base.

2.5%
RLVR format fallback

RLVR reduced answer-format fallback from 24.1% to 2.5% and leakage from 3.53% to 0.00%.

9.1%
RLVR override improvement

Explicit English-reasoning compliance improved by 9.1 percentage points.

What the paper found

This study fine-tunes four sparse mixture-of-experts models—Alibaba’s Qwen3.6-35B-A3B, OpenAI’s Gpt-OSS-20B, and NVIDIA’s NemotronH variants—with LoRA to reason in Greek, using 118,092 Greek training rows and a 5,156-item evaluation suite spanning mathematics, commonsense, and logic. Accuracy barely changes: the best configuration scores 76.5 versus 77.2 for its base, while changing only the random seed shifts results by 7.7 points, larger than every corpus or recipe effect tested. Behavioural metrics reveal the real transformation: base models produce Greek reasoning on 0 of 1,000 Greek prompts, whereas SFT checkpoints reach about 98% Greek-trace fidelity, eliminate in-question language switching, improve grammaticality, and substantially shorten reasoning; Qwen uses 3-fold fewer tokens, while the effect is family-dependent because Greek tokenization costs 2.3–2.5 times more tokens per word than English. SFT also introduces format failures and channel leakage, which accuracy obscures. A pre-registered GRPO-based reinforcement-learning-with-verifiable-rewards experiment repairs answer-format fallback from 24.1% to 2.5% and answer-channel leakage from 3.53% to 0.00%, while improving compliance with an explicit instruction to reason in English by 9.1 percentage points; Greek fidelity remains 98.2%. The paper therefore argues that low-resource reasoning models should be evaluated on language fidelity, reasoning budget, termination, format compliance, and steerability—not accuracy alone, echoing but extending DeepSeek-R1’s use of language-consistency rewards.

Original abstract

Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →