On Language Drift during RLVR Post-Training
AuthorsMichael Sullivan, Alexander Koller
AffiliationsDepartment of Language Science and Technology, Saarland University, Saarbrücken, Germany
Resources
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Key results
Total RLVR and SFT runs across the three model families.
Seeds used to replicate each optimal training configuration.
Validation accuracy after RLVR.
Validation accuracy after RLVR, primarily enabled by behavior sharpening.
RLVR minus SFT AUC for best-performing checkpoints.
RLVR minus SFT AUC for best-performing checkpoints.
What the paper found
This paper examines language drift: the emergence of unusual, internally coherent, or illegible chain-of-thought language during reinforcement learning with verifiable reward, or RLVR. Its central theoretical result is that outcome-only RLVR can produce unbounded KL divergence from human language, because it optimizes final-answer correctness while placing no direct constraint on the reasoning trace; supervised fine-tuning, or SFT, has a finite drift bound. The paper further proves that imposing a fixed language-drift budget necessarily limits maximum expected reward, creating a capability–monitorability trade-off. Empirically, researchers trained gemma-3-1b-pt, Llama-3.2-1B, and Qwen2.5-1.5B on GSM8K using DAPO, a GRPO variant, and SFT, completing 30 training runs across 5 random seeds with generation limited to 256 tokens. On GSM8K, RLVR reached 0.23 for Llama-3.2-1B, 0.12 for gemma-3-1b-pt, and 0.79 for Qwen2.5-1.5B, while Qwen’s improvement was largely behavior sharpening rather than novel reasoning. Trace-legibility AUC differences between RLVR and SFT were -0.24 for Llama, -0.17 for Gemma, and 0.02 for Qwen, indicating substantially greater drift when novel behavior had to be discovered. Different random seeds also produced mutually less-legible reasoning, suggesting drift is path-dependent and unpredictable. The findings contextualize reports of illegible reasoning in DeepSeek-R1, OpenAI’s GPT-5, and Anthropic’s Claude models: as frontier systems pursue novel tasks, higher capability may come with less human-monitorable reasoning.
Original abstract
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.
Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning
Shuvendu K. Lahiri
NFV uses AI agents to translate ordinary code into machine-checkable formal proofs, making software verification more accessible while revealing the limits of end-to-end soundness.