NTH

CausalMix: Data Mixture as Causal Inference for Language Model Training

AuthorsZinan Tang, Yukun Zhang, Shaomian Zheng, Zhuoshi Pan, Qizhi Pei, Dingnan Jin, Jun Zhou, Yujun Wang, Biqing Huang

July 7, 2026 2 min read
Watch on YouTube
The one-line take

CausalMix treats LLM data mixing like a causal problem, using learned treatment effects to choose better training mixtures that scale to larger models and data pools.

Key results

512
proxy runs

Historical Qwen2.5-0.5B proxy trainings used to fit the causal model

100K
proxy size

Each proxy run used a 100K-example SFT sub-dataset

800K
main data pool

Largest tulu-3-sft-mixture setting evaluated for scaling and transfer

33.94
CausalMix-A AvgDev

Best reported average development score for Qwen2.5-0.5B at 800K

62.28
CausalMix-S AvgDev

Best reported average development score for Qwen2.5-7B at 800K

66.66
LongCoT Avg

CausalMix performance on AM-Thinking-v1-Distilled-Code&Math with Qwen3-4B

What the paper found

CausalMix, developed by researchers from Tsinghua University, Ant Group, and Renmin University of China, reframes supervised fine-tuning data-mixture search as a causal inference problem rather than a static regression over past runs. The method uses 512 proxy trainings of Qwen2.5-0.5B on 100K-example mixtures from tulu-3-sft-mixture, extracts state covariates from OpenDataArena-scored-data-2603 using Normalized_Loss, HES, and Writing_Style, and fits a CausalForestDML model with LightGBM nuisance predictors to estimate state-conditioned marginal returns on log-mixture weights. On the main Tulu 3 benchmark suite, CausalMix-A reaches 33.94 AvgDev at 800K data with Qwen2.5-0.5B, while CausalMix-S reaches 62.28 AvgDev on Qwen2.5-7B at the same 800K scale, both exceeding prior data-mixing baselines such as RegMix, DoReMi, ODM, and DMO in the reported settings. The framework also transfers to an unseen LongCoT-style setup: on AM-Thinking-v1-Distilled-Code&Math with Qwen3-4B, CausalMix achieves 66.66 average score, ahead of the 64.74 best grid baseline. An ablation shows that removing covariates or DML orthogonalization degrades performance, confirming that the gain comes from state-conditioned causal estimation rather than ordinary mixture regression. The paper also reports a closed-form simplex policy: allocate only to domains with positive estimated causal effect, normalized by ReLU-style weighting.

Original abstract

In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of Qwen2.5-0.5B to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on Qwen3-4B-Base. By leveraging causal modeling to isolate confounding biases, CausalMix dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by CausalMix consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall, CausalMix offers a causal and interpretable framework for optimizing LLM data mixtures.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis