NTH

Can Generalist Agents Automate Data Curation?

AuthorsFeiyang Kang, Hanze Li, Adam Nguyen, Mahavir Dabas, Jiaqi W. Ma, Frederic Sala, Dawn Song, Ruoxi Jia

June 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper asks whether coding agents can take over the tedious loop of data curation, and finds they can do useful work—but only when guided by structured method adaptation rather than vague prompting.

Key results

665K
LLaVA pool size

candidate instruction-tuning pool used for the main selection task

10K
Selection budget

curated subset size selected by the agent for fine-tuning

33.7
Open-prompt Claude Code score

average score on 8 benchmarks after LLaVA-1.5-7B fine-tuning

32.5
Best random baseline score

best of 10 random 10k selections on the same task

34.9
Best scaffolded score

best 10k policy reached by the adapt-papers scaffold

What the paper found

This paper asks whether generalist coding agents, including Anthropic’s Claude Code, OpenAI’s Codex, and open-source backbones run through OpenHands, can automate training-data curation as an iterative research loop rather than a one-shot selection problem. The authors introduce Curation-Bench, a terminal-based benchmark that fixes the model, training recipe, evaluation suite, and contamination checks while giving agents control only over executable data policies. On a 10k-example selection task from LLaVA-665K for fine-tuning LLaVA-1.5-7B, open-prompt agents reliably execute the loop and outperform random selection and published baselines such as ICONS and ARDS, with Claude Code reaching 33.7 average score versus 32.5 for the best random run and recovering 59% of the full-data gain using only 1.5% of the pool. The key failure is that open-ended prompting produces local heuristic edits—source ratios, length thresholds, and seed sweeps—rather than new policy families. Light scaffolds broaden the search vocabulary but do not raise the best score, while heavy scaffolds that force evidence-grounded or method-grounded iteration do: the strongest scaffold composes an EL2N-style top-loss selector with a p95 assistant-loss noise filter and reaches 34.9, exceeding the open-prompt maximum and even a 100k ARDS baseline. Longer sessions also keep improving up to 50 iterations, and the benchmark extends to DataComp Small, where the agent beats the top-30% CLIP ViT-L/14 filtering baseline, showing the framework generalizes beyond instruction tuning.

Original abstract

Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback. We ask whether generalist coding agents can automate this data-curation loop. We introduce *Curation-Bench*, an agent-centric benchmark that fixes the model, training recipe, and evaluation suite while giving agents command-line access to inspect data, implement policies, submit them to a fixed training/evaluation pipeline, and revise. In a vision-language instruction-tuning instantiation, out-of-the-box agents reach strong published data-selection baselines within ten iterations. However, trajectory analysis reveals a persistent *execution-research gap*: agents mainly tune local policy variants rather than explore new policy families, even when given strategy guides and paper references. Scaffolds requiring each iteration to cite, instantiate, and adapt a prior method shift agents toward method-guided exploration. The scaffolded agent autonomously composes -- without human design input -- a data-selection policy that outperforms strong published baselines at one-tenth their data budget. Overall, current agents can run the curation loop, but reliable data research requires scaffolded method adaptation, not open-ended prompting alone. Code and benchmark are open-sourced.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis