Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
AuthorsChenhao Dang, Jing Ma, Mingjie Liao
Resources
This paper uses reinforcement learning to automatically decide what training data an LLM should see next, making pre-training faster and improving final model quality.
Key results
HDS matches the prior best final perplexity on The Pile with fewer steps
HDS matches the static baseline perplexity with fewer steps
Accuracy gain over the prior best method
Accuracy gain over the prior best method
Training-time speedup to reach the TPW perplexity target
Final validation perplexity achieved by HDS at the last checkpoint
What the paper found
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning, from Chenhao Dang, Jing Ma at Renmin University of China, and Mingjie Liao at Alibaba Group, reframes online data mixing for LLM pre-training as a continuous-control reinforcement learning problem and optimizes it with Soft Actor-Critic. The key novelty is a holistic reward that combines inter-domain gradient alignment, scheduled lexical diversity, and model stability, so the scheduler can favor data that is jointly useful, curriculum-appropriate, and convergence-stable instead of relying on a single loss signal. On The Pile, HDS trains a Pythia-1B model for 50 billion tokens and reaches the same final validation perplexity as the prior best dynamic method with 44% fewer training iterations, while beating the static The Pile Weights baseline with 57% fewer iterations. At the end of training, it lowers validation perplexity by 13.6% versus TPW and improves MMLU by 7.2% on 0-shot accuracy and 4.0% on 5-shot accuracy over the prior best method. The paper also shows that the scheduler adds only 26.5M parameters and about 0.02 seconds per step, yet delivers a 2.21x wall-clock speedup, and the gains persist at scale on Pythia-12B, where HDS reduces perplexity from 7.32 to 4.89 by the final checkpoint.
Original abstract
The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training. Online Data Mixing (ODM), the technique of adaptively adjusting data mixtures during training, has emerged as a promising direction to improve efficiency. However, existing methods are constrained by their reliance on a singular optimization perspective, which fundamentally overlooks the need for complex LLM pre-training to consider the dynamic data composition from multiple dimensions. To overcome this limitation, we introduce the Holistic Data Scheduler (HDS), a novel online data mixing framework. HDS formulates the data scheduling challenge as a reinforcement learning problem in a continuous control space and leverages the Soft Actor-Critic (SAC) algorithm for its stability and sample efficiency in exploring the high-dimensional policy space. At the core of HDS lies a novel multi-objective, holistic reward function that integrates three critical perspectives: a data-driven reward for quality, a loss-driven reward capturing inter-domain influence, and a model-driven reward based on weight norms. To validate our design and determine its optimal configuration, we conducted systematic experiments on LLMs of various sizes. On The Pile benchmark, HDS reaches the final validation perplexity of the next best method with 44% fewer training iterations. Furthermore, it achieves a 7.2% improvement on the MMLU 0-shot task along with consistent gains on other benchmarks, showcasing its ability to enhance both training efficiency and final model capability.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.