NTH

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

AuthorsChenhao Dang, Jing Ma, Mingjie Liao

July 9, 2026 2 min read
Watch on YouTube
The one-line take

This paper uses reinforcement learning to automatically decide what training data an LLM should see next, making pre-training faster and improving final model quality.

Key results

44%
training iteration reduction vs AC-ODM

HDS matches the prior best final perplexity on The Pile with fewer steps

57%
training iteration reduction vs TPW

HDS matches the static baseline perplexity with fewer steps

7.2%
MMLU 0-shot improvement

Accuracy gain over the prior best method

4.0%
MMLU 5-shot improvement

Accuracy gain over the prior best method

2.21x
wall-clock speedup

Training-time speedup to reach the TPW perplexity target

4.89
Pythia-12B perplexity

Final validation perplexity achieved by HDS at the last checkpoint

What the paper found

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning, from Chenhao Dang, Jing Ma at Renmin University of China, and Mingjie Liao at Alibaba Group, reframes online data mixing for LLM pre-training as a continuous-control reinforcement learning problem and optimizes it with Soft Actor-Critic. The key novelty is a holistic reward that combines inter-domain gradient alignment, scheduled lexical diversity, and model stability, so the scheduler can favor data that is jointly useful, curriculum-appropriate, and convergence-stable instead of relying on a single loss signal. On The Pile, HDS trains a Pythia-1B model for 50 billion tokens and reaches the same final validation perplexity as the prior best dynamic method with 44% fewer training iterations, while beating the static The Pile Weights baseline with 57% fewer iterations. At the end of training, it lowers validation perplexity by 13.6% versus TPW and improves MMLU by 7.2% on 0-shot accuracy and 4.0% on 5-shot accuracy over the prior best method. The paper also shows that the scheduler adds only 26.5M parameters and about 0.02 seconds per step, yet delivers a 2.21x wall-clock speedup, and the gains persist at scale on Pythia-12B, where HDS reduces perplexity from 7.32 to 4.89 by the final checkpoint.

Original abstract

The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training. Online Data Mixing (ODM), the technique of adaptively adjusting data mixtures during training, has emerged as a promising direction to improve efficiency. However, existing methods are constrained by their reliance on a singular optimization perspective, which fundamentally overlooks the need for complex LLM pre-training to consider the dynamic data composition from multiple dimensions. To overcome this limitation, we introduce the Holistic Data Scheduler (HDS), a novel online data mixing framework. HDS formulates the data scheduling challenge as a reinforcement learning problem in a continuous control space and leverages the Soft Actor-Critic (SAC) algorithm for its stability and sample efficiency in exploring the high-dimensional policy space. At the core of HDS lies a novel multi-objective, holistic reward function that integrates three critical perspectives: a data-driven reward for quality, a loss-driven reward capturing inter-domain influence, and a model-driven reward based on weight norms. To validate our design and determine its optimal configuration, we conducted systematic experiments on LLMs of various sizes. On The Pile benchmark, HDS reaches the final validation perplexity of the next best method with 44% fewer training iterations. Furthermore, it achieves a 7.2% improvement on the MMLU 0-shot task along with consistent gains on other benchmarks, showcasing its ability to enhance both training efficiency and final model capability.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis