Knowledge Pull Requests for Continual Document Authoring
AuthorsAlexander Martin, Benjamin Van Durme
AffiliationsJohns Hopkins University · {amart233, vandurme}@jhu.edu
Resources
Knowledge Pull Requests make continually updating documents more transparent by showing exactly which claims changed, where they belong, and how the final text was revised.
Key results
MegaWika 2.0 articles evaluated for cross-lingual revision
Non-English Wikipedia languages used as knowledge sources
MiRAGE precision for claims added to revised Wikipedia articles
MiRAGE recall of cross-lingual information integrated into Wikipedia
Correct answers on the difficult 100-question multilingual QA subset
Percentage of the original Wikipedia article left unchanged
What the paper found
Knowledge Pull Requests, or KPRs, introduce a reviewable alternative to opaque document rewriting. The system decomposes source and main documents into atomic, decontextualized claims, classifies them as supported, absent, or conflicting, filters for coverage and relevance, routes new claims to sections, and rewrites only affected sections. Its ChangeLog separates a knowledge-level Claim Proposal from the resulting text diff, allowing reviewers to approve additions and adjudicate conflicts rather than silently overwriting facts. Evaluated on 599 English Wikipedia articles from MegaWika 2.0, spanning 49 source languages, KPR achieved MiRAGE information precision of 0.878, retained 0.950 of existing information, and added 0.891 of source information, outperforming raw-text ConText and unfiltered claim-based ConClaim. On multilingual question answering, KPR grounding averaged 69.2 accuracy versus 58.5 for ConClaim and 48.4 for ConText, while preserving English-content performance. On a difficult 100-question subset, GPT-5.6 answered 62 questions with a KPR-revised article, compared with 38 using web search; even smaller models such as Llama-3.1-8B benefited. KPR edits preserved 97.3 percent of the original Wikipedia text and required 11.2 contiguous approvals, compared with 25.3 for ConClaim. In RAGTIME report updating, the same claim-level process improved retention and addition across temporal, conflict, and balanced settings, while explicitly withholding contradictory claims. The pipeline used Qwen3.5-27B for decomposition, classification, routing, and rewriting, demonstrating that continual authoring can be both information-complete and auditable.
Original abstract
We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or regenerate from scratch. A KPR integrates new knowledge into a document by extracting claims, filtering and routing them to sections, and flagging conflicts with existing content, producing a ChangeLog that separates what knowledge changes (claim proposal) from how the text changes (document diff). We evaluate KPRs on revising Wikipedia across languages and updating query-driven reports on RAGTIME. KPRs integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated. A KPR-revised article also grounds question answering better than a frontier model with search, which does not surface knowledge documented only in other languages.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.