Lost in Context: Addressing Context Anxiety in Large Language Models
AuthorsIfueko Igbinedion, Jillian Ross, Etienne Ricardez, Sertac Karaman, Eric So
Resources
Some reasoning models give up not because they lack the ability to solve a problem, but because they mistakenly think they will run out of context before finishing.
Key results
Valid model-task observations analyzed after API errors.
Anxious models overestimated the tokens required for solutions.
Accuracy loss associated with context anxiety after model and disk fixed effects.
Correct answers produced with context anxiety were longer.
Absolute accuracy improvement from anxiety-filtered Reasoning SFT on unseen 5×5 grids.
Required moves for the 12-disk benchmark instances.
What the paper found
Researchers at MIT identify and measure “context anxiety,” a failure mode in which large language models abandon solvable long-horizon tasks because they overestimate the tokens required, rather than lacking the underlying reasoning capability. On a Tower of Hanoi benchmark spanning 270 problems from 2 to 12 disks, requiring up to 4,095 moves, they evaluated Claude Sonnet 3.7 and 4.5, DeepSeek R1, Gemini 2.5 Flash, Kimi K2 Thinking, and OpenAI o4 Mini, analyzing 1,585 valid observations. Anxiety-driven traces overestimated required token usage by 24 percent, reduced accuracy by 15.3 percent after controlling for model and difficulty effects, and made successful solutions 54 percent longer. The authors then fine-tuned OpenAI’s GPT-OSS-20B using supervised learning on correct, anxiety-free reasoning traces, masking final-answer tokens so the model learned effort-allocation behavior rather than memorized move sequences. This selective Reasoning SFT reduced anxiety across held-out disk counts and transferred to shortest-path grid search: on unseen 5×5 grids, accuracy rose by 26.7 percentage points, while an otherwise identical SFT model trained on all correct traces often underperformed. The results suggest that improving self-assessment and adaptive reasoning policies can increase reliability and efficiency without scaling model size, context limits, or inference-time compute, although the detector relies on explicit verbalized output concerns and the main benchmark is symbolic.
Original abstract
Conventional wisdom suggests that reasoning models fail when problems exceed their capabilities. However, we find that frontier reasoning models sometimes possess the necessary capabilities to solve problems but fail due to premature self-doubt -- a phenomenon informally known as context anxiety. We provide the first systematic study of context anxiety, demonstrating that it arises, in part, from a model's inability to accurately estimate the tokens required to complete a task. We also show that context anxiety leads to material efficiency losses when models operate under perceived constraints. Building on this analysis, we further show that models can learn alternative strategies for solving long-horizon problems without exhibiting context anxiety, suggesting that performance improvements may be achievable not through scaling model capabilities, but by improving models' ability to accurately assess and adapt to their own limitations.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.