NTH

Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

AuthorsGiovanni Dettori, Matteo Boffa, Danilo Giordano, Idilio Drago, Marco Mellia

July 9, 2026 2 min read
Watch on YouTube
The one-line take

The paper shows that LLMs struggle not just with long inputs, but especially with long inputs packed with lots of new information, sharply reducing their effective context window.

Key results

12k
context length

All benchmarks use approximately 12k-token contexts.

0.58
MATTR sparse

MK-NIAH lexical density.

0.75
MATTR dense

Scene-Rules lexical density.

1.0
MATTR max density

WordChecker lexical density.

1%
avg drop MK-NIAH

Average change at deepest position vs 2k baseline.

27%
avg drop Scene-Rules

Average change at deepest position vs 2k baseline.

What the paper found

This paper, by researchers at Politecnico di Torino and the University of Turin, argues that long-context failures in open-weight LLMs are not explained by input length and needle position alone: lexical density, measured with Moving-Average Type-Token Ratio, is a third axis that shrinks effective context capacity. Using three matched “find-the-needle” benchmarks with ≈12k-token contexts, the authors show that MK-NIAH is sparse at MATTR=0.58, Scene-Rules is denser at 0.75, and WordChecker is maximally dense at 1.0; across eight open LLMs from 9B to 685B parameters, models that are near-perfect on MK-NIAH collapse below 60% retrieval on the denser benchmarks. At the deepest positions, average accuracy changes by +1% on MK-NIAH, −27% on Scene-Rules, and −31% on WordChecker relative to the 2k-token baseline, and within-benchmark sparsification recovers performance by up to 24% on Scene-Rules and 30% on WordChecker. The failure mode is model-specific under high density: DeepSeek-V3.2, Llama 4 Maverick, and Mistral Small mostly conflate candidate and query text, GPT-OSS models increasingly abstain, and Qwen3.5 models often invert the search procedure and truncate, revealing that dense contexts do not just reduce accuracy but change how models search at all.

Original abstract

Input length and the position of relevant information are widely cited as the primary causes of degraded LLM long-context performance. Here, we study lexical density -- the rate at which a context introduces distinct information -- as a third, largely overlooked factor that systematically reduces the effective context window of LLMs. We quantify the impact of lexical density on open-weight LLMs (9B-685B) using three "find-the-needle" style benchmarks with identical length (~12k tokens) and controlled needle position, but increasing density of information. We observe a sharp performance collapse in higher-density benchmarks: models that are near-perfect in sparse contexts drop below 60% retrieval score on denser ones. To rule out task-type confounds, we vary and control the density within each benchmark while keeping all other properties unchanged. Reducing density generally restores performance, especially in the high-density regimes where degradation appears. These results show that effective context capacity is a function of lexical density, with direct implications for real-world LLM systems operating on compact, information-rich inputs.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis