Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs
AuthorsGiovanni Dettori, Matteo Boffa, Danilo Giordano, Idilio Drago, Marco Mellia
Resources
The paper shows that LLMs struggle not just with long inputs, but especially with long inputs packed with lots of new information, sharply reducing their effective context window.
Key results
All benchmarks use approximately 12k-token contexts.
MK-NIAH lexical density.
Scene-Rules lexical density.
WordChecker lexical density.
Average change at deepest position vs 2k baseline.
Average change at deepest position vs 2k baseline.
What the paper found
This paper, by researchers at Politecnico di Torino and the University of Turin, argues that long-context failures in open-weight LLMs are not explained by input length and needle position alone: lexical density, measured with Moving-Average Type-Token Ratio, is a third axis that shrinks effective context capacity. Using three matched “find-the-needle” benchmarks with ≈12k-token contexts, the authors show that MK-NIAH is sparse at MATTR=0.58, Scene-Rules is denser at 0.75, and WordChecker is maximally dense at 1.0; across eight open LLMs from 9B to 685B parameters, models that are near-perfect on MK-NIAH collapse below 60% retrieval on the denser benchmarks. At the deepest positions, average accuracy changes by +1% on MK-NIAH, −27% on Scene-Rules, and −31% on WordChecker relative to the 2k-token baseline, and within-benchmark sparsification recovers performance by up to 24% on Scene-Rules and 30% on WordChecker. The failure mode is model-specific under high density: DeepSeek-V3.2, Llama 4 Maverick, and Mistral Small mostly conflate candidate and query text, GPT-OSS models increasingly abstain, and Qwen3.5 models often invert the search procedure and truncate, revealing that dense contexts do not just reduce accuracy but change how models search at all.
Original abstract
Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study lexical density -- the rate at which a context introduces distinct information -- as a third, largely overlooked factor that systematically reduces the effective context window of LLMs. We quantify the impact of lexical density on open-weight LLMs (9B-685B) using three "find-the-needle" style benchmarks with identical length (~12k tokens) and controlled needle position, but increasing density of information. We observe a sharp performance collapse in higher-density benchmarks: models that are near-perfect in sparse contexts drop below 60% retrieval score on denser ones. To rule out task-type confounds, we vary and control the density within each benchmark while keeping all other properties unchanged. Reducing density generally restores performance, especially in the high-density regimes where degradation appears. These results show that effective context capacity is a function of lexical density, with direct implications for real-world LLM systems operating on compact, information-rich inputs.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.