NTH

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

AuthorsYaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Xinyue Bi, Zhaoyi Li, Zhiqiang Shen

June 2, 2026 2 min read
Watch on YouTube
The one-line take

LLMSurgeon tries to infer what data an LLM was trained on just by reading its outputs, offering a new way to audit foundation models without access to their training sets.

Key results

6
LLMScan domains

Coarse-grained LLMScan setup spans 6 broad data domains for auditing general-purpose models.

17
LLMScan mid-grained domains

Mid-grained LLMScan uses 17 Pile-based domains for finer mixture recovery.

87
LLMScan programming languages

Fine-grained LLMScan evaluates StarCoder across 87 programming language categories.

95.14%
LLMSurgeon overlap accuracy

LLMSurgeon achieves 95.14% overlap accuracy on LLaMA1-7B on LLMScan.

94.46%
LLMSurgeon overlap accuracy

LLMSurgeon achieves 94.46% overlap accuracy on OLMo-1B on LLMScan.

27.54%
StarCoder best baseline overlap accuracy

On the fine-grained StarCoder setting, the strongest baseline reaches 27.54% while LLMSurgeon reaches 30.37%.

What the paper found

LLMSurgeon, from MBZUAI’s VILA Lab, tackles a new auditing problem the authors call Data Mixture Surgery: inferring the domain-level pretraining mixture of a large language model from generated text alone, without access to weights or training data. The method assumes label shift and uses a three-stage inverse-correction pipeline: train a proxy domain classifier on reference data, estimate its soft confusion matrix C to capture systematic domain confusions, sample the target model with neutral prompts, then solve a constrained linear inverse problem to recover the latent mixture prior π from the observed classifier outputs. To benchmark this, the paper introduces LLMScan, a verifiable suite built from open-source models with known data recipes, including LLaMA-1, OLMo, Amber, Pythia, GPT-Neo, and StarCoder, spanning 6 coarse domains, 17 mid-grained Pile domains, and 87 programming languages. On LLMScan, LLMSurgeon reaches 95.14% overlap accuracy on LLaMA1-7B and 94.46% on OLMo-1B, while the strongest aggregation-style baselines stay near 50%; even on fine-grained StarCoder it still improves over the best baseline, 30.37% versus 27.54%. Ablations show that fine-tuned DistilBERT gives the best proxy classifier and that merging semantically indistinguishable sources like C4 and Common Crawl is essential to avoid ill-conditioned inversion. The authors also show the method can track training dynamics across checkpoints and estimate injected toxic data proportions with 97% to 98% accuracy in a GPT-2 sandbox, positioning LLMSurgeon as a practical post-hoc tool for transparency, bias analysis, and safety triage of foundation models.

Original abstract

The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{Data Mixture Surgery (DMS)}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{LLMSurgeon}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{LLMScan}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis