Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It
AuthorsRandhir Kumar
Resources
Per-chunk fact checking breaks multi-hop RAG because no single passage contains the whole answer, but verifying against decomposed sub-questions substantially restores reliability.
Key results
Deployable per-chunk entailment verifier performance.
Control result showing entailment verification works when one passage is sufficient.
Difference from unfiltered retrieval in the Qwen2.5-1.5B end-to-end evaluation.
Decomposition repair, rising from 0.546 with the original question.
Retrieval-anchored, off-the-shelf decomposition reaching 31% of the gold-decomposition ceiling.
What the paper found
This paper shows that a standard RAG safeguard—scoring and dropping passages one chunk at a time—systematically fails on multi-hop questions because it assumes each passage is sufficient to answer the original query. On HotpotQA, 2WikiMultihopQA, and MuSiQue, the deployable entailment verifier reaches only 0.643, 0.523, and 0.560 AUC, while the same nli-deberta-v3-base pipeline reaches 0.951 on single-hop SQuAD. The verifier preferentially keeps the paragraph named by the question, usually the first-hop bridge, and rejects the unmentioned paragraph containing the final answer; this bias worsens with hop count and persists across BAAI/bge-small-en-v1.5, intfloat/e5-base-v2, and thenlper/gte-base. End to end, per-chunk gating is the worst selector across all three benchmarks, reducing Exact Match by 13.4 points on HotpotQA relative to unfiltered retrieval, with the penalty growing as generators become stronger. The repair is to decompose the question and verify each later hop against a sub-question: on MuSiQue, gold decomposition raises later-hop AUC from 0.546 to 0.840, a paired lift of +0.355. An off-the-shelf Qwen2.5-7B-Instruct decomposer conditioned on the question and top retrieved paragraph reaches 0.637 AUC, capturing 31% of the ceiling. The authors conclude that systems such as Self-Ask and IRCoT already generate the structure needed for verification, but discard it before filtering; verification should follow the decomposition rather than the original question.
Original abstract
Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail. We show this cannot work for multi-hop questions, and show what does. Per-chunk scoring assumes one chunk is a sufficient premise for the answer. Multi-hop questions are built so that none is, and the paragraph carrying the answer is the one the question does not name. Entailment scoring reaches 0.643, 0.523 and 0.560 AUC on HotpotQA, 2WikiMultihopQA and MuSiQue, against 0.951 on single-hop SQuAD. Seven controls rule out model capacity, premise length, hypothesis template, decision threshold, retriever, answer-matching criterion and prompt. End to end across three datasets, three generator sizes and two prompts, per-chunk gating is significantly worse than not filtering at all in every cell, and its penalty grows with generator capability. The repair is to condition verification on the decomposed sub-question rather than the original query. Using MuSiQue's gold decomposition, entailment on a later hop rises from 0.546, which is chance, to 0.840, a paired lift of +0.355 with a bootstrap interval of [0.331, 0.382]. An off-the-shelf Qwen2.5-7B decomposer, given the question and the top retrieved paragraph, reaches 0.637 and captures 31% of that ceiling; decomposing without retrieval reaches 0.533, below the original question. Iterative retrieval systems already produce such decompositions and discard them before verifying.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.