NTH

Knowledge Pull Requests for Continual Document Authoring

AuthorsAlexander Martin, Benjamin Van Durme

AffiliationsJohns Hopkins University · {amart233, vandurme}@jhu.edu

September 24, 2026 2 min read
Watch on YouTube
The one-line take

Knowledge Pull Requests make continually updating documents more transparent by showing exactly which claims changed, where they belong, and how the final text was revised.

Key results

599
Wikipedia articles

MegaWika 2.0 articles evaluated for cross-lingual revision

49
Source languages

Non-English Wikipedia languages used as knowledge sources

0.878
KPR information precision

MiRAGE precision for claims added to revised Wikipedia articles

0.891
KPR information recall added

MiRAGE recall of cross-lingual information integrated into Wikipedia

62
GPT-5.6 with KPR

Correct answers on the difficult 100-question multilingual QA subset

97.3%
KPR preserved text

Percentage of the original Wikipedia article left unchanged

What the paper found

Knowledge Pull Requests, or KPRs, introduce a reviewable alternative to opaque document rewriting. The system decomposes source and main documents into atomic, decontextualized claims, classifies them as supported, absent, or conflicting, filters for coverage and relevance, routes new claims to sections, and rewrites only affected sections. Its ChangeLog separates a knowledge-level Claim Proposal from the resulting text diff, allowing reviewers to approve additions and adjudicate conflicts rather than silently overwriting facts. Evaluated on 599 English Wikipedia articles from MegaWika 2.0, spanning 49 source languages, KPR achieved MiRAGE information precision of 0.878, retained 0.950 of existing information, and added 0.891 of source information, outperforming raw-text ConText and unfiltered claim-based ConClaim. On multilingual question answering, KPR grounding averaged 69.2 accuracy versus 58.5 for ConClaim and 48.4 for ConText, while preserving English-content performance. On a difficult 100-question subset, GPT-5.6 answered 62 questions with a KPR-revised article, compared with 38 using web search; even smaller models such as Llama-3.1-8B benefited. KPR edits preserved 97.3 percent of the original Wikipedia text and required 11.2 contiguous approvals, compared with 25.3 for ConClaim. In RAGTIME report updating, the same claim-level process improved retention and addition across temporal, conflict, and balanced settings, while explicitly withholding contradictory claims. The pipeline used Qwen3.5-27B for decomposition, classification, routing, and rewriting, demonstrating that continual authoring can be both information-complete and auditable.

Original abstract

We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or regenerate from scratch. A KPR integrates new knowledge into a document by extracting claims, filtering and routing them to sections, and flagging conflicts with existing content, producing a ChangeLog that separates what knowledge changes (claim proposal) from how the text changes (document diff). We evaluate KPRs on revising Wikipedia across languages and updating query-driven reports on RAGTIME. KPRs integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated. A KPR-revised article also grounds question answering better than a frontier model with search, which does not surface knowledge documented only in other languages.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis