NTH

Chain-of-Experience for Continual LLM Improvement

AuthorsHaoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan

August 21, 2026 2 min read
Watch on YouTube
The one-line take

Chain-of-Experience helps LLMs learn from repeated feedback during inference, improving performance and token efficiency across math, coding, and knowledge tasks.

Key results

5.6%
Overall CoE improvement

Average improvement from feedback-driven CoE over no-feedback inference.

19%
API cost reduction

Lower API cost achieved across tasks and models.

8.6%
Coding feedback gain

Average executor-feedback gain on LiveCodeBench (V6) and LiveBench (Code), from 66.4% to 75.0%.

71.0%
Self-feedback average accuracy

Average performance with self-feedback versus 66.8% for the no-feedback baseline.

0.50
Base ability correlation

Average Pearson correlation between initial ability and improvement capacity across five benchmarks.

What the paper found

This study introduces Chain-of-Experience, or CoE, a training-free test-time learning framework in which an LLM repeatedly solves a problem, stores the full trajectory of actions and feedback, and conditions its next attempt on that accumulated experience. Feedback can be absent, generated by the model itself, supplied by a code executor, or provided as a correctness signal. Across math, coding, and knowledge benchmarks, experiments with eight models—including OpenAI’s GPT-5, Gemini-2.5 Pro, and Anthropic’s Claude-4.5 Sonnet—show a 5.6% overall improvement over no-feedback inference while reducing API cost by 19%. On LiveCodeBench (V6) and LiveBench (Code), executor feedback raises accuracy by 8.6%, from 66.4% to 75.0%, while self-feedback reaches 71.0% on average compared with 66.8% for the no-feedback baseline. Combining model feedback with correctness or executor signals can improve results further, reaching 81.2% on LiveBench (Code), although Dynamic CheatSheet, Agentic Context Engineering, and other compressed-memory methods often discard useful intermediate reasoning. Improvement capacity correlates positively with initial model ability, averaging a Pearson correlation of 0.50 across five benchmarks. Most gains appear early in the interaction loop, and models remain partially robust to spurious feedback. Because CoE does not update model parameters, its improvements reflect contextual reuse of experience rather than persistent weight-level learning.

Original abstract

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis