NTH

Autodata: An agentic data scientist to create high quality synthetic data

AuthorsIlia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston

July 7, 2026 2 min read
Watch on YouTube
The one-line take

Autodata turns an AI agent into a data scientist that generates better training data, and even learns how to improve its own data-making process.

Key results

6.59
CS agentic rounds

Mean rounds per accepted example on computer science papers

0.677
CS weak avg

Weak solver score under CoT Self-Instruct

0.458
CS agentic weak avg

Weak solver score under Agentic Self-Instruct

0.774
CS RL mean@3

Qwen3.5-4B trained on Agentic Self-Instruct data, evaluated on the CoT test

0.393
Legal PRBench-Legal

Qwen3.5-4B RL on Agentic Self-Instruct data, Kimi-K2.6 grader

71.86%
Principia avg@8

Overall combined-validation score for Agentic Self-Instruct in scientific reasoning

What the paper found

Autodata, from FAIR at Meta, is a general framework that turns an AI agent into a data scientist that iteratively creates, evaluates, and revises synthetic training data, then meta-optimizes its own prompting strategy. In the paper’s Agentic Self-Instruct implementation, a Kimi-K2.6 orchestrator coordinates a challenger, weak solver, strong solver, and judge to generate examples that separate capability levels, using Qwen3.5-4B as the weak model and Qwen3.5-397B-A17B as the strong model. On computer science papers from S2ORC, the agent needed 6.59 rounds on average to accept an example, and the accepted set shifted weak-solver quality from 0.677 to 0.458 while raising strong-solver quality from 0.696 to 0.772; training Qwen3.5-4B with GRPO on 1.3k Agentic examples lifted held-out mean@3 from 0.630 to 0.774 on the CoT test and from 0.366 to 0.632 on the harder Agentic test. On legal reasoning from Pile of Law, the same loop reshaped weak rollouts from degenerate all-zero outputs into a usable learning signal, and GRPO training on 2.8k Agentic examples improved PRBench-Legal from 0.343 to 0.393 under Kimi-K2.6 grading, surpassing the larger 397B baseline at 0.358. In scientific reasoning over mathematical objects, Agentic data achieved the best overall avg@8 on the combined validation set at 71.86%, and reduced truncation rates at a 65,536-token budget from 23.75% to 4.09%; a separate meta-optimization loop further raised the CS validation pass rate from 62.1% to 79.6% over 233 iterations.

Original abstract

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning with mathematical objects, where we obtain improved results compared to classical synthetic dataset creation methods. Further, meta-optimizing the data scientist agent itself delivers an even larger performance uplift. Agentic data creation provides a way to convert increased inference compute into higher quality model training. Overall, we believe this direction has the potential to change the way we build AI data.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis