Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
AuthorsQian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou
Resources
IAR teaches language models to memorize and use a document collection without retrieval while preserving their broader abilities.
Key results
IAR beats Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings
Average improvement in domain QA accuracy, measured in percentage points
Average improvement across IFEval, MMLU, and MSBench, measured in percentage points
Retrieval-free domain QA accuracy, compared with 42.4 for Vanilla SFT
Largest mean-general-performance recovery across Qwen3-8B, Qwen3-14B, and Qwen3-32B
What the paper found
The paper introduces IAR—Inject, Align, Recover—a staged post-training method for making a fixed document collection answerable without retrieval. Inject exposes the model to denser document supervision through continuation, rewrite, and instruction-conditioned reconstruction objectives; Align converts that knowledge into answer-only question answering; and Recover merges the adapted checkpoint with the original instruction model using methods such as TIES, SLERP, or task arithmetic to restore general abilities. Tested on Common Corpus, CCI, and Llama-3.2-3B, Phi-4-mini, Qwen3-4B, and SmolLM3-3B, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA and 12.1 percentage points across IFEval, MMLU, and MSBench. On Common Corpus with Qwen3-4B, domain accuracy rises to 50.5% from 42.4% while general metrics also improve. Scaling Qwen3 to 8B, 14B, and 32B shows recovery restoring up to 24.1 percentage points of mean general performance while sacrificing only about one domain point. The results indicate that structured exposure, QA alignment, and capability recovery should be optimized separately; IAR is a strong operating-point strategy rather than a universally dominant recipe. Domain judgments used panels including OpenAI’s gpt-oss-120b and DeepSeek-V3 variants.
Original abstract
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.