NTH

From Knowledge Access to Source Learning: Developing Source-Specific Competence

AuthorsLucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

AffiliationsGeorgia Institute of Technology · Oct 2026 € Website: https://sourcelearn.github.io/ · Introduction Large language model (LLM) agents increasingly rely on external sources to solve knowledge-intensive tasks. Here, a source denotes an identifiable body of external knowledge, such as a do · University of Illinois at Urbana-Champaign · University of California, Los Angeles

October 6, 2026 2 min read
Watch on YouTube
The one-line take

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Key results

13
Best settings

SourceLearn achieved the best result in 13 of 15 benchmark-backend settings.

22.6
Maximum gain over Hybrid RAG

The largest reported improvement over Hybrid RAG was 22.6 points, on APIBench with GPT-5.6-Luna.

14.3
GPT-5.6-Luna average gain

SourceLearn improved over Hybrid RAG by an average of 14.3 points with GPT-5.6-Luna.

40.0%
Test-claim coverage after learning

Coverage of source claims required by test questions increased to 40.0%, from 23.2% initially.

86.6
Accuracy with complete claim coverage

Answer accuracy reached 86.6% when all required claims were represented in the source model.

What the paper found

SourceLearn reframes repeated use of an external knowledge source as source learning: building a persistent, revisable source model that captures structures, conditions, procedures, relations, and applicability rather than merely storing retrieved facts or task memories. Its Self-Directed Source Learning uses model-conditioned inspection and adaptive DEEPEN or CONNECT study actions to resolve gaps, while Task-Guided Source Learning diagnoses failures and recurring representation needs, then reconstructs updates from the authoritative source instead of copying answers. At task time, the activated source model complements Hybrid RAG evidence. Across five benchmarks—MultiDoc2Dial, NarrativeQA, SWE-QA, APIBench, and AppWorld—and three backends—OpenAI’s GPT-5.6-Luna, gpt-oss-120b, and DeepSeek-V4.1-Flash—SourceLearn was best in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and an average GPT-5.6-Luna improvement of 14.3 points. The model’s coverage of claims needed by test questions increased from 23.2% initially to 40.0% after learning, and accuracy rose from 67.5% when no required claims were represented to 86.6% when all were represented. The findings indicate that source learning changes representation toward explicit rules, conditions, procedures, and cross-entity structure, transferring beyond the exact regions exposed during guidance tasks. The main limitation is scope: the method targets persistent, authoritative, relatively stable sources rather than noisy or rapidly changing knowledge.

Original abstract

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis
03Continual Learning

ACLArena: Agent Continue Learning in Multi-stage Post-training

Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang

ACLArena studies how agents can learn new skills over multiple training stages without forgetting old ones, proposing replay and specialized LoRA experts as a practical solution.

Read analysis