Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?
AuthorsRenhao Pei, Yihong Liu, Sampo Pyysalo, Hinrich Schütze, Shaoxiong Ji
Resources
This paper shows that giving LLMs step-by-step grammatical reasoning can boost translation quality for very low-resource languages, especially when the reasoning is used at inference time rather than learned from scratch.
Key results
Maximum BLEU improvement from adding reasoning traces in ICL
Maximum LLMaJ improvement from adding reasoning traces in ICL
What the paper found
This paper asks whether synthetic linguistic reasoning traces can improve low-resource machine translation, and it tests that idea on Xibe and Chintang using instruction-tuned Qwen3 and Gemma 4 models. The authors build an automatic pipeline that converts Universal Dependencies treebanks, dictionaries, and modular grammar-rule banks into step-by-step traces that explain lexical choice, morphosyntactic analysis, and phrasal composition, then evaluate them in in-context learning, supervised fine-tuning, and reinforcement fine-tuning. The key finding is that reasoning traces are most effective as inference-time guidance: in ICL they substantially improve translation quality across most models and metrics, with especially large gains on Chintang, including up to 5.57 BLEU, 11.89 chrF, 19.74 SBERT, and 23.42 LLMaJ for gemma-4-E4B-it and up to 19.74 SBERT and 23.42 LLMaJ for Qwen3-4B-Thinking. By contrast, using the same traces as training supervision is less reliable: SFT tends to produce the right output format but often still learns incorrect reasoning content, and RFT adds no clear gains over SFT. Overall, the work shows that LLMs can use grammatical information effectively when given reliable sentence-specific analyses, but learning to generate those analyses remains the main bottleneck.
Original abstract
Large language models (LLMs) offer a promising approach to machine translation (MT) for extremely low-resource languages by incorporating linguistic resources through in-context learning. However, LLMs often struggle to apply grammatical information effectively during translation. Inspired by recent progress in chain-of-thought reasoning, we investigate whether low-resource MT can benefit from structured intermediate steps of linguistic analysis and grammatical reasoning. We propose a pipeline for automatically generating step-by-step linguistic reasoning traces from Universal Dependencies treebanks, dictionaries, and grammar-rule banks. We evaluate these traces in three settings: in-context learning (ICL), supervised fine-tuning (SFT), and reinforcement fine-tuning (RFT), on Xibe and Chintang as test cases. Our results show that linguistic reasoning traces are most effective as inference-time guidance: in ICL, reliable sentence-specific traces substantially improve translation performance across most models, languages, and metrics. In contrast, using the linguistic reasoning traces as training data yields smaller and less consistent gains, as models learn the trace format but often generate erroneous content. These findings suggest that LLMs can leverage grammatical information for low-resource MT when given reliable linguistic analyses, while learning to generate such analyses remains a major bottleneck.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.