Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
AuthorsTianci Liu, Zihan Dong, Tianchun Li, Yi-Chung Chen, Qiming Cao, Xingchen Wang, Shiyang Wang, Zichen Miao, Linjun Zhang, Haoyu Wang, Jing Gao
Resources
HPSE helps language models turn newly injected passages into usable, multi-step knowledge rather than merely memorized text.
Key results
Each UnKEBench passage contains 5 atomic facts used to test decomposition.
Average relative gain from HPSE across the four language-model backbones.
Average improvement in MQuAKE-uns multi-hop composition accuracy.
Maximum relative MQuAKE-uns improvement for LoRA during continual editing.
HPSE step-in frequency fell to 1.7% by the third rollout round.
What the paper found
This paper addresses a central weakness of unstructured knowledge editing: language models may memorize a newly injected passage but fail to retrieve its individual facts or combine them in multi-hop reasoning. The proposed method, Hybrid-Policy Self-Editing, or HPSE, uses the same model in two states: a student being edited and a privileged copy reading the new passage in context. During self-distillation, the student generates an on-policy rollout, while a confidence- and disagreement-gated step-in replaces only tokens where the student is likely to miss novel facts; an NLL anchor preserves passage-level learning. Unlike methods requiring external supervision, auxiliary models, or synthetic data, HPSE changes only the training signal and plugs into gradient-based editors such as LoRA and FT-M. Experiments span Qwen2.5-7B-Instruct, Qwen3-8B, Llama-3.1-8B-Instruct, and Gemma-2-9B-it, using UnKEBench, whose passages contain 5 atomic facts, and MQuAKE-uns for multi-hop composition. Across four backbones, FT-M gained a relative 67.9% on MQuAKE-uns, while LoRA improved multi-hop composition by 9.9 points on average. In continual editing, LoRA’s relative MQuAKE-uns improvement reached 149%. The intervention also rapidly became unnecessary: its step-in rate fell from 26.8% to 1.7% within 3 rounds, indicating that the edited model was internalizing the missing knowledge. Overall, HPSE improves decomposition and composition while largely preserving locality measured with MMLU.
Original abstract
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.