NTH

Make LLM Learn to Synthesize from Streaming Experiences through Feedback

AuthorsZhenlin Hu, Yan Wang, Zhen Bi, Zihao Xue, Bingyu Zhu, Longtao Huang, Xiongtao Zhang, Zeyu Yang, Zhixuan Chu, Jungang Lou

June 4, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks whether LLMs can get better at making synthetic data over time by learning from a stream of past synthesis tasks instead of treating each one in isolation.

Key results

82.50
LLaMA3.1-8B MNLI accuracy

Under the StreamSynth task stream, SynLearner reaches 82.50 accuracy on MNLI with LLaMA3.1-8B, compared with 80.91 for GORP.

86.56
Qwen2.5-7B MNLI accuracy

Under the StreamSynth task stream, SynLearner reaches 86.56 accuracy on MNLI with Qwen2.5-7B, compared with 85.53 for GORP.

2.81
LLaMA3.1-8B GSM8K gain

In the Step 5 generalization experiment, SynLearner improves GSM8K accuracy by 2.81 points over Ori on LLaMA3.1-8B.

7.80
LLaMA3.1-8B MATH-500 gain

In the Step 5 generalization experiment, SynLearner improves MATH-500 accuracy by 7.80 points over Ori on LLaMA3.1-8B.

What the paper found

This paper from Huzhou Normal University and Alibaba Group introduces StreamSynth, a new setting for synthetic data generation in which tasks arrive sequentially and a model is trained to learn transferable synthesis behavior from past experience rather than generating each dataset in isolation. The authors propose SynLearner, which combines Diversity-Aware Initialization with Hierarchical Reward Optimization: prompts are dynamically instantiated and evolved to expand both depth and breadth of synthesis patterns, then the model is fine-tuned and reinforced using a dual reward that mixes sample-level quality—structural validity, fluency, and task relevance—with set-level distinctiveness computed from batch embedding density. Experiments on the Yelp, Amazon, Yahoo, and MNLI stream with LLaMA3.1-8B and Qwen2.5-7B show that SynLearner consistently outperforms direct prompt-only synthesis and continual-learning baselines such as GORP, FAPM, InsCL, and SEEKR, with especially strong gains on later tasks; for example, on LLaMA3.1-8B it reaches 82.50 accuracy on MNLI versus 80.91 for GORP, and on Qwen2.5-7B it improves MNLI to 86.56 versus 85.53. Ablations confirm that removing dynamic prompting or GRPO-based reward optimization degrades performance, and t-SNE plus cross-task heatmaps show broader coverage and better forward transfer. The method also generalizes to reasoning, improving GSM8K and MATH-500 by 2.81 and 7.80 points on LLaMA3.1-8B, with larger gains on Qwen2.5-7B.

Original abstract

Large language models (LLMs) have been widely adopted for synthetic data generation, significantly reducing annotation costs. However, most existing studies treat synthesis as a set of isolated tasks and overlook a more fundamental question: whether a model can learn to synthesize by accumulating experience from past tasks and transferring it to future ones. In this work, we introduce StreamSynth, a new setting in which synthesis tasks arrive sequentially and experience from historical tasks provides informative signals for future synthesis. To address this setting, we propose SynLearner, a general framework that enables synthesis models to acquire reusable synthesis experience over a task stream. Instead of generating data independently for each task, SynLearner encourages the model to explore diverse synthesis patterns, learn from feedback, and balance sample quality with set-level diversity as tasks evolve. Extensive experiments across multiple benchmarks show that SynLearner effectively leverages experience from earlier tasks to improve synthesis performance on later ones, exhibiting consistent cross-task transferability. These findings provide evidence for the feasibility of StreamSynth and highlight synthetic data generation as an experience-driven process that can benefit from task streams.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis