NTH

Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors

AuthorsZhenghua Bao

August 26, 2026 2 min read
Watch on YouTube
The one-line take

Multi-hop RAG may retrieve more intelligently, but when speech recognition gets an entity wrong, that extra reasoning structure can make the final answer even worse.

Key results

12,000
Spoken evaluation suite

Total spoken queries across HotpotQA, 2WikiMultiHopQA, and MuSiQue.

36.5%
IRCoT+HippoRAG2 amplification on HotpotQA

Relative increase in the clean-text to Nigerian-ASR F1 gap versus Naive RAG.

42.3%
IRCoT+HippoRAG2 amplification on 2WikiMultiHopQA

Relative increase in the clean-text to Nigerian-ASR F1 gap versus Naive RAG.

67.4%
IRCoT+HippoRAG2 amplification on MuSiQue

Relative increase in the clean-text to Nigerian-ASR F1 gap versus Naive RAG.

87%
Entity corruption rate

Lower bound of degradation cases involving corrupted query entities on 2WikiMultiHopQA.

51.8%
Real-speech entity mistranscription

Share of 500 real Nigerian utterances containing named entities with at least one mistranscribed entity.

What the paper found

This study evaluates whether multi-hop retrieval-augmented generation becomes more robust or more fragile when spoken queries contain automatic speech-recognition errors. Its controlled suite contains 12,000 spoken queries from HotpotQA, 2WikiMultiHopQA, and MuSiQue, synthesized in four English accents with Microsoft Edge TTS, transcribed by Whisper-large-v3, and answered by OpenAI’s gpt-4o-mini. The systems compare Naive RAG with HippoRAG2’s entity graph, IRCoT’s iterative reformulation, and their combination. Higher word error rate consistently causes larger F1 losses: on Nigerian-accented input, the combined IRCoT+HippoRAG2 system loses 0.142 on HotpotQA, 0.195 on 2WikiMultiHopQA, and 0.072 on MuSiQue relative to clean text, making its degradation 36.5%, 42.3%, and 67.4% larger than Naive RAG’s. Although structurally richer systems retain higher absolute F1, they amplify upstream errors because corrupted query entities disrupt graph linking and successive reformulations; entity corruption appears in at least 87% of 2WikiMultiHopQA degradation cases. N-best decoding and Double Metaphone-based phonetic entity correction recover no more than 12% of the clean-to-ASR gap. Validation on 500 real Nigerian utterances found entity mistranscription in 51.8% of utterances containing named entities, while replacing Whisper with Meta’s SeamlessM4T-v2-large preserved the same robustness ranking. The central implication is that better retrieval architecture can produce worse error robustness in voice-driven RAG.

Original abstract

Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configurations generally retain higher absolute F1 under ASR input, both extensions amplify the error: the F1 gap from clean text to the highest-WER accent is 36-67% larger under their combination than under naive dense retrieval, on all three benchmarks. The dominant failure mode is corruption of one or more query entities, accounting for 87-96% of degradation cases on 2WikiMultiHopQA across all four methods. Two lightweight surface-form mitigations leave most of the gap intact, indicating that downstream retrieval structure amplifies remaining entity errors. We release code and data at https://github.com/ZhenghuaBao/spoken-multihop-rag .

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis