LLMSurgeon: Diagnosing Data Mixture of Large Language Models
AuthorsYaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang, Jiacheng Liu, Xinyue Bi, Zhaoyi Li, Zhiqiang Shen
Resources
LLMSurgeon tries to infer what data an LLM was trained on just by reading its outputs, offering a new way to audit foundation models without access to their training sets.
Key results
Coarse-grained LLMScan setup spans 6 broad data domains for auditing general-purpose models.
Mid-grained LLMScan uses 17 Pile-based domains for finer mixture recovery.
Fine-grained LLMScan evaluates StarCoder across 87 programming language categories.
LLMSurgeon achieves 95.14% overlap accuracy on LLaMA1-7B on LLMScan.
LLMSurgeon achieves 94.46% overlap accuracy on OLMo-1B on LLMScan.
On the fine-grained StarCoder setting, the strongest baseline reaches 27.54% while LLMSurgeon reaches 30.37%.
What the paper found
LLMSurgeon, from MBZUAI’s VILA Lab, tackles a new auditing problem the authors call Data Mixture Surgery: inferring the domain-level pretraining mixture of a large language model from generated text alone, without access to weights or training data. The method assumes label shift and uses a three-stage inverse-correction pipeline: train a proxy domain classifier on reference data, estimate its soft confusion matrix C to capture systematic domain confusions, sample the target model with neutral prompts, then solve a constrained linear inverse problem to recover the latent mixture prior π from the observed classifier outputs. To benchmark this, the paper introduces LLMScan, a verifiable suite built from open-source models with known data recipes, including LLaMA-1, OLMo, Amber, Pythia, GPT-Neo, and StarCoder, spanning 6 coarse domains, 17 mid-grained Pile domains, and 87 programming languages. On LLMScan, LLMSurgeon reaches 95.14% overlap accuracy on LLaMA1-7B and 94.46% on OLMo-1B, while the strongest aggregation-style baselines stay near 50%; even on fine-grained StarCoder it still improves over the best baseline, 30.37% versus 27.54%. Ablations show that fine-tuned DistilBERT gives the best proxy classifier and that merging semantically indistinguishable sources like C4 and Common Crawl is essential to avoid ill-conditioned inversion. The authors also show the method can track training dynamics across checkpoints and estimate injected toxic data proportions with 97% to 98% accuracy in a GPT-2 sandbox, positioning LLMSurgeon as a practical post-hoc tool for transparency, bias analysis, and safety triage of foundation models.
Original abstract
The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{Data Mixture Surgery (DMS)}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{LLMSurgeon}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{LLMScan}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.