Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors
AuthorsZhenghua Bao
Multi-hop RAG may retrieve more intelligently, but when speech recognition gets an entity wrong, that extra reasoning structure can make the final answer even worse.
Key results
Total spoken queries across HotpotQA, 2WikiMultiHopQA, and MuSiQue.
Relative increase in the clean-text to Nigerian-ASR F1 gap versus Naive RAG.
Relative increase in the clean-text to Nigerian-ASR F1 gap versus Naive RAG.
Relative increase in the clean-text to Nigerian-ASR F1 gap versus Naive RAG.
Lower bound of degradation cases involving corrupted query entities on 2WikiMultiHopQA.
Share of 500 real Nigerian utterances containing named entities with at least one mistranscribed entity.
What the paper found
This study evaluates whether multi-hop retrieval-augmented generation becomes more robust or more fragile when spoken queries contain automatic speech-recognition errors. Its controlled suite contains 12,000 spoken queries from HotpotQA, 2WikiMultiHopQA, and MuSiQue, synthesized in four English accents with Microsoft Edge TTS, transcribed by Whisper-large-v3, and answered by OpenAI’s gpt-4o-mini. The systems compare Naive RAG with HippoRAG2’s entity graph, IRCoT’s iterative reformulation, and their combination. Higher word error rate consistently causes larger F1 losses: on Nigerian-accented input, the combined IRCoT+HippoRAG2 system loses 0.142 on HotpotQA, 0.195 on 2WikiMultiHopQA, and 0.072 on MuSiQue relative to clean text, making its degradation 36.5%, 42.3%, and 67.4% larger than Naive RAG’s. Although structurally richer systems retain higher absolute F1, they amplify upstream errors because corrupted query entities disrupt graph linking and successive reformulations; entity corruption appears in at least 87% of 2WikiMultiHopQA degradation cases. N-best decoding and Double Metaphone-based phonetic entity correction recover no more than 12% of the clean-to-ASR gap. Validation on 500 real Nigerian utterances found entity mistranscription in 51.8% of utterances containing named entities, while replacing Whisper with Meta’s SeamlessM4T-v2-large preserved the same robustness ranking. The central implication is that better retrieval architecture can produce worse error robustness in voice-driven RAG.
Original abstract
Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configurations generally retain higher absolute F1 under ASR input, both extensions amplify the error: the F1 gap from clean text to the highest-WER accent is 36-67% larger under their combination than under naive dense retrieval, on all three benchmarks. The dominant failure mode is corruption of one or more query entities, accounting for 87-96% of degradation cases on 2WikiMultiHopQA across all four methods. Two lightweight surface-form mitigations leave most of the gap intact, indicating that downstream retrieval structure amplifies remaining entity errors. We release code and data at https://github.com/ZhenghuaBao/spoken-multihop-rag .
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.