NTH

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

AuthorsZhengpei Hu, Kai Li, Dapeng Fu, Xuechao Zou, Yuanhao Tang, Yue Li, Tengfei Cao, Jianqiang Huang

August 16, 2026 2 min read
Watch on YouTube
The one-line take

Hard prompt compression can delete the crucial context that makes surviving answers understandable, but a lightweight restoration step can recover much of the lost accuracy.

Key results

34%
Beaver dangling range

Lower bound of dangling answer paths at compression ratio 0.30 across HotpotQA, 2WikiMultiHopQA, and MuSiQue.

54%
Beaver dangling range

Upper bound of dangling answer paths at compression ratio 0.30 across the three multi-hop QA datasets.

80
LongBench-v2 affected documents

Every one of the 80 evaluated Single-Document QA documents contained at least one dangling reference.

34
Fixed-budget reselection gain

Maximum accuracy improvement in points from restoring omitted support while removing equal-cost distractors.

4.7
Automatic restoration gain

HotpotQA accuracy improvement in points with Qwen3-8B when the ratio changes from 0.30 to 0.31.

What the paper found

This paper identifies referential dangling, a structural failure in hard prompt compression where independently selected fragments remain relevant but lose the definitions or bridge facts needed to interpret them. At compression ratio 0.30, Beaver, using Qwen3-0.6B embeddings, produces dangling answer paths in 34% to 54% of bridge examples across HotpotQA, 2WikiMultiHopQA, and MuSiQue, while all 80 LongBench-v2 Single-Document QA documents contain at least one dangling reference. The problem generalizes across six compressors, including PartPrompt, Selective-Context, LLMLingua-2, LongLLMLingua, and DAC. On affected examples, fixed-budget reselection with Qwen3-8B restores the omitted supporting paragraph while removing equally costly distractors, improving accuracy by 29 to 34 points; stronger answer models do not reliably compensate, with OpenAI’s GPT-5.5 still 8.8 points less accurate on MuSiQue when support is incomplete. A BERT-based dependency classifier provides annotation-free restoration, improving HotpotQA accuracy by 4.7 points while changing the compression ratio only from 0.30 to 0.31. The central design implication is that compressors must optimize referential completeness jointly with relevance, rather than treating fragments as independently valuable.

Original abstract

Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring units under a budget. We identify a structural failure in this procedure: independent selection can split dependent evidence pairs, retaining one member while deleting the other. When retained text contains an answer but deleted text defines the entity needed to interpret it, we call the result referential dangling. At a compression ratio of 0.30, Beaver, which ranks coherent chunks using Qwen3-0.6B embeddings, leaves the answer path incomplete in 34-54% of bridge examples across three multi-hop question answering datasets. On a shared HotpotQA bridge set, all six hard compressors we test exhibit dangling at rates up to 60%, and every document in LongBench-v2 Single-Document QA contains at least one dangling reference. On dangling examples evaluated with Qwen3-8B, reinserting the missing supporting paragraph while removing nonsupporting paragraphs to maintain the token budget improves accuracy by 29-34 percentage points (p < 0.0001), recovering at least 88% of the gap to contexts retaining both supporting paragraphs. Stronger answer models do not absorb the loss: on MuSiQue, GPT-5.5 is 8.8 points less accurate on compressed contexts than on contexts retaining both supporting paragraphs. Finally, we train a compact classifier to rank omitted sentences by whether they are needed to interpret retained text and reinsert the top-ranked candidates without support annotations at inference. On HotpotQA with Qwen3-8B, this automatic restoration improves accuracy by 4.7 points while changing the compression ratio only from 0.30 to 0.31. Hard compressors should optimize both relevance and referential completeness.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis