CausalMix: Data Mixture as Causal Inference for Language Model Training
AuthorsZinan Tang, Yukun Zhang, Shaomian Zheng, Zhuoshi Pan, Qizhi Pei, Dingnan Jin, Jun Zhou, Yujun Wang, Biqing Huang
Resources
CausalMix treats LLM data mixing like a causal problem, using learned treatment effects to choose better training mixtures that scale to larger models and data pools.
Key results
Historical Qwen2.5-0.5B proxy trainings used to fit the causal model
Each proxy run used a 100K-example SFT sub-dataset
Largest tulu-3-sft-mixture setting evaluated for scaling and transfer
Best reported average development score for Qwen2.5-0.5B at 800K
Best reported average development score for Qwen2.5-7B at 800K
CausalMix performance on AM-Thinking-v1-Distilled-Code&Math with Qwen3-4B
What the paper found
CausalMix, developed by researchers from Tsinghua University, Ant Group, and Renmin University of China, reframes supervised fine-tuning data-mixture search as a causal inference problem rather than a static regression over past runs. The method uses 512 proxy trainings of Qwen2.5-0.5B on 100K-example mixtures from tulu-3-sft-mixture, extracts state covariates from OpenDataArena-scored-data-2603 using Normalized_Loss, HES, and Writing_Style, and fits a CausalForestDML model with LightGBM nuisance predictors to estimate state-conditioned marginal returns on log-mixture weights. On the main Tulu 3 benchmark suite, CausalMix-A reaches 33.94 AvgDev at 800K data with Qwen2.5-0.5B, while CausalMix-S reaches 62.28 AvgDev on Qwen2.5-7B at the same 800K scale, both exceeding prior data-mixing baselines such as RegMix, DoReMi, ODM, and DMO in the reported settings. The framework also transfers to an unseen LongCoT-style setup: on AM-Thinking-v1-Distilled-Code&Math with Qwen3-4B, CausalMix achieves 66.66 average score, ahead of the 64.74 best grid baseline. An ablation shows that removing covariates or DML orthogonalization degrades performance, confirming that the gain comes from state-conditioned causal estimation rather than ordinary mixture regression. The paper also reports a closed-form simplex policy: allocate only to domains with positive estimated causal effect, normalized by ReLU-style weighting.
Original abstract
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of Qwen2.5-0.5B to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on Qwen3-4B-Base. By leveraging causal modeling to isolate confounding biases, CausalMix dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by CausalMix consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall, CausalMix offers a causal and interpretable framework for optimizing LLM data mixtures.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.