NTH

Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization

AuthorsQian Kou, Xiaofeng Shi, Xiaosong Qiu, Hua Zhou

August 30, 2026 2 min read
Watch on YouTube
The one-line take

IAR teaches language models to memorize and use a document collection without retrieval while preserving their broader abilities.

Key results

7
Settings with four-metric improvement

IAR beats Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings

3.6
Average domain QA gain

Average improvement in domain QA accuracy, measured in percentage points

12.1
Average general-capability gain

Average improvement across IFEval, MMLU, and MSBench, measured in percentage points

50.5
Qwen3-4B Common Corpus IAR accuracy

Retrieval-free domain QA accuracy, compared with 42.4 for Vanilla SFT

24.1
Maximum Qwen3 recovery gain

Largest mean-general-performance recovery across Qwen3-8B, Qwen3-14B, and Qwen3-32B

What the paper found

The paper introduces IAR—Inject, Align, Recover—a staged post-training method for making a fixed document collection answerable without retrieval. Inject exposes the model to denser document supervision through continuation, rewrite, and instruction-conditioned reconstruction objectives; Align converts that knowledge into answer-only question answering; and Recover merges the adapted checkpoint with the original instruction model using methods such as TIES, SLERP, or task arithmetic to restore general abilities. Tested on Common Corpus, CCI, and Llama-3.2-3B, Phi-4-mini, Qwen3-4B, and SmolLM3-3B, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA and 12.1 percentage points across IFEval, MMLU, and MSBench. On Common Corpus with Qwen3-4B, domain accuracy rises to 50.5% from 42.4% while general metrics also improve. Scaling Qwen3 to 8B, 14B, and 32B shows recovery restoring up to 24.1 percentage points of mean general performance while sacrificing only about one domain point. The results indicate that structured exposure, QA alignment, and capability recovery should be optimized separately; IAR is a strong operating-point strategy rather than a universally dominant recipe. Domain judgments used panels including OpenAI’s gpt-oss-120b and DeepSeek-V3 variants.

Original abstract

Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time. We study this setting as document knowledge internalization: converting a fixed corpus into usable parametric knowledge for retrieval-free question answering. We propose IAR (Inject, Align, and Recover), a three-stage post-training framework that separates structured document knowledge injection, QA behavior alignment, and general ability recovery. Unlike conventional continued pretraining, Inject converts source documents into continuation, rewrite, and instruction-conditioned reconstruction objectives. Align then adapts the injected model with answer-only QA supervision, while Recover merges the domain-adapted model with the base instruction model to recover general capabilities. Across Common Corpus (CC) and CCI, and across Llama, Phi, Qwen, and SmolLM model families, IAR improves the domain-primary domain-general frontier for retrieval-free document internalization. In the main comparison, IAR improves over Vanilla SFT on all four reported metrics in 7 of 8 dataset-model settings, with average gains of 3.6 percentage points in domain QA accuracy and 12.1 percentage points in mean general performance across IFEval, MMLU, and MSBench. Extended CC baselines show that LoRA and FAPM can win individual general metrics, but among methods that also reach leading or near-leading domain internalization, IAR retains one of the strongest general profiles.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis