Where Should a Document Live: Context, Representations, or Parameters?
AuthorsNathanaël Carraz Rakotonirina, Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert
AffiliationsGonzalo Iglesias · Amazon AGI · {ncarraz, momchilh, gjii, agispert}@amazon.com
Resources
This study asks whether new knowledge belongs in an LLM’s context, hidden representations, or weights, finding that KV-cache-based Cartridges often deliver the best accuracy but can cause forgetting.
Key results
Compaction nearly matches the 73.6 in-context baseline in the oracle setting.
Cartridges improve from 70.8 at one retrieved document to 83.2 at ten.
Compaction declines from 65.4 at one retrieved document to 26.4 at ten.
Merged LoRA falls from 57.4 at one retrieved document to 23.5 at three.
Cartridges cause 6% average forgetting on general control benchmarks.
What the paper found
This paper asks whether newly acquired document knowledge should remain in the context window, be stored as compressed representations, or be encoded in model parameters. Using Qwen3-8B, with validation on Gemma-3-12B, it compares in-context learning, Cartridges, Compaction, LoRA, MLP adapters, and full fine-tuning across LongHealth, QASPER, QuALITY, FinQA, and TechQA, using synthetic Self-Study data and distillation from a document-conditioned teacher. In the single-document oracle setting, Compaction reaches an average score of 73.4, nearly matching the 73.6 in-context baseline, while parametric methods trail substantially. In multi-document retrieval, Cartridges are uniquely composable: their LongHealth score rises from 70.8 at one retrieved document to 83.2 at ten, whereas Compaction falls from 65.4 to 26.4 on TechQA and merged LoRA drops from 57.4 to 23.5. The core failure is interference between independently trained artifacts; alternative LoRA merging methods do not solve it. Compaction preserves general capabilities, but Cartridges incur 6% average degradation on control benchmarks and 13% on coding, mainly at high compression. Representation and parametric adapters are cheaper than repeatedly processing raw text: a 10× token reduction yields approximately 100× fewer prefill FLOPs, and loading ten Cartridges takes 30–50 ms versus 400–800 ms for equivalent raw context. The practical conclusion is conditional: Cartridges offer the best accuracy and retrieval composition, Compaction offers stronger retention of general skills, and LoRA offers fixed-size, low-cost serving, but no method dominates every deployment scenario. The evaluation also uses DeepSeek-Distilled-Qwen-32B as a judge for TechQA.
Original abstract
To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost, and performance trade-offs, with no single winner. We present a controlled comparison of representation-based (KV-cache based) and parametric (fine-tuning-based) adaptation methods on five knowledge-intensive benchmarks. We show that in the oracle setting, Cartridges (KV) are the most accurate injection method at nearly every storage budget, outperforming parametric methods by 10 points. Compaction (KV) matches Cartridges only at low compression rates, lagging behind the parametric methods by 10 points at rates higher than $50\times$. In the more realistic multi-document retrieval scenario, Cartridges are the only method that matches in-context learning (ICL), leading the parametric methods by 29 points and Compaction by 15 points. Nonetheless, Cartridges are also the only method, besides full fine-tuning and large MLP adapters, that suffers from catastrophic forgetting, i.e., a 6% performance degradation on control benchmarks, with 13% in coding.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.