NTH

Where Should a Document Live: Context, Representations, or Parameters?

AuthorsNathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert

AffiliationsGonzalo Iglesias · Amazon AGI · {ncarraz, momchilh, gjii, agispert}@amazon.com

September 26, 2026 2 min read
Watch on YouTube
The one-line take

This study asks whether new knowledge belongs in an LLM’s context, hidden representations, or weights, finding that KV-cache-based Cartridges often deliver the best accuracy but can cause forgetting.

Key results

73.4
Compaction average single-document score

Compaction nearly matches the 73.6 in-context baseline in the oracle setting.

83.2
Cartridges LongHealth at k=10

Cartridges improve from 70.8 at one retrieved document to 83.2 at ten.

26.4
Compaction TechQA at k=10

Compaction declines from 65.4 at one retrieved document to 26.4 at ten.

23.5
LoRA TechQA at k=3

Merged LoRA falls from 57.4 at one retrieved document to 23.5 at three.

6%
Cartridges average control-benchmark degradation

Cartridges cause 6% average forgetting on general control benchmarks.

What the paper found

This paper asks whether newly acquired document knowledge should remain in the context window, be stored as compressed representations, or be encoded in model parameters. Using Qwen3-8B, with validation on Gemma-3-12B, it compares in-context learning, Cartridges, Compaction, LoRA, MLP adapters, and full fine-tuning across LongHealth, QASPER, QuALITY, FinQA, and TechQA, using synthetic Self-Study data and distillation from a document-conditioned teacher. In the single-document oracle setting, Compaction reaches an average score of 73.4, nearly matching the 73.6 in-context baseline, while parametric methods trail substantially. In multi-document retrieval, Cartridges are uniquely composable: their LongHealth score rises from 70.8 at one retrieved document to 83.2 at ten, whereas Compaction falls from 65.4 to 26.4 on TechQA and merged LoRA drops from 57.4 to 23.5. The core failure is interference between independently trained artifacts; alternative LoRA merging methods do not solve it. Compaction preserves general capabilities, but Cartridges incur 6% average degradation on control benchmarks and 13% on coding, mainly at high compression. Representation and parametric adapters are cheaper than repeatedly processing raw text: a 10× token reduction yields approximately 100× fewer prefill FLOPs, and loading ten Cartridges takes 30–50 ms versus 400–800 ms for equivalent raw context. The practical conclusion is conditional: Cartridges offer the best accuracy and retrieval composition, Compaction offers stronger retention of general skills, and LoRA offers fixed-size, low-cost serving, but no method dominates every deployment scenario. The evaluation also uses DeepSeek-Distilled-Qwen-32B as a judge for TechQA.

Original abstract

To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost, and performance trade-offs, with no single winner. We present a controlled comparison of representation-based (KV-cache based) and parametric (fine-tuning-based) adaptation methods on five knowledge-intensive benchmarks. We show that in the oracle setting, Cartridges (KV) are the most accurate injection method at nearly every storage budget, outperforming parametric methods by 10 points. Compaction (KV) matches Cartridges only at low compression rates, lagging behind the parametric methods by 10 points at rates higher than $50\times$. In the more realistic multi-document retrieval scenario, Cartridges are the only method that matches in-context learning (ICL), leading the parametric methods by 29 points and Compaction by 15 points. Nonetheless, Cartridges are also the only method, besides full fine-tuning and large MLP adapters, that suffers from catastrophic forgetting, i.e., a 6% performance degradation on control benchmarks, with 13% in coding.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis