ICICLE: Expanding Retrieval with In-Context Documents
AuthorsYu-Chen Den, Yung-Yu Shih, Zhi Rui Tam, Kuan-Yu Chen, Pu-Jen Cheng, Yun-Nung Chen, Eugene Yang
Resources
ICICLE lets generative retrievers add new documents on the fly by using in-context document IDs instead of retraining the model each time.
Key results
ICICLE achieves Hits@1 on unseen MS MARCO documents with 100 in-context candidates.
ICICLE achieves Hits@10 on unseen MS MARCO documents with 100 in-context candidates.
ICICLE achieves Hits@1 on unseen NQ320K documents with 100 in-context candidates.
ICICLE achieves Hits@10 on unseen NQ320K documents with 100 in-context candidates.
Continual fine-tuning methods still require updating the model on the newly added documents, while ICICLE avoids retraining for corpus expansion.
What the paper found
ICICLE, from National Taiwan University and Johns Hopkins University, reframes incremental generative retrieval as an in-context indexing problem rather than a parameter-update problem. Instead of retraining a retriever to learn new document–docid links, it supplies newly added documents and their titles as inference-time evidence, then teaches a Qwen3-1.7B-base model to route between contextual copy and parametric memory using a special [COPY] token, hard-negative in-context templates, Direct Preference Optimization, and a large-title-context LoRA adaptation stage. On MS MARCO and NQ320K, with 90/10 corpus splits and up to 100 in-context candidates, ICICLE reaches 0.607 Hits@1 and 0.800 Hits@10 on unseen MS MARCO documents and 0.649 Hits@1 and 0.772 Hits@10 on unseen NQ320K documents, outperforming incremental baselines such as DSI++, DOME, and no-retrain DSI while avoiding O(M) or O(N+M) corpus-update costs. The paper’s main technical finding is that scaling failures are driven less by retrieval ranking than by routing failure: the model often knows the answer but fails to activate [COPY] as the candidate set grows, and DPO substantially improves this source-selection calibration. The authors also show that abstractive document compression with Qwen2.5-14B-Instruct preserves retrieval quality better than keyword compression, enabling more candidates within the context window without destroying rank-1 accuracy.
Original abstract
Generative retrieval (GR) maps queries directly to document identifiers (docids) using parametric knowledge, However, this design makes corpus expansion costly: adding new documents requires updating model parameters to encode new document-docid associations incurs repeated training and catastrophic forgetting of previously indexed documents. In this work, we revisit incremental GR as an in-context retrieval problem, where newly added documents are supplied as inference-time document-docid evidence. We propose ICICLE, an in-context indexing framework that performs source-aware docid generation over both parametric memory and context-provided document-docid pairs. ICICLE combines a `[COPY]`-based routing mechanism, preference-based calibration, and large context adaptation to distinguish context-grounded retrieval from parametric retrieval. Experiments on MS MARCO and NQ320K show that ICICLE improves retrieval of newly introduced documents while preserving seen-document retention without corpus-specific retraining. Our analysis further shows that high-shot degradation is mainly caused by routing failure, highlighting source-selection calibration as a key bottleneck for scaling in-context generative retrieval.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.