NTH

OpenThoughts-Agent: Data Recipes for Agentic Models

AuthorsNegin Raoof, Richard Zhuang, Marianna Nezhurina, Etash Guha, Atula Tejaswi, Ryan Marten, Charlie F. Ruan, Tyler Griggs, Alexander Glenn Shaw, Hritik Bansal, E. Kelly Buchanan, Artem Gazizov, Reinhard Heckel, Chinmay Hegde, Sankalp Jajee, Daanish Khazi, Emmanouil Koukoumidis, Xiangyi Li, Hange Liu, Shlok Natarajan, Harsh Raj, Nicholas Roberts, Ethan Shen, Nishad Singhi, Michael Siu, Ashima Suvarna, Hanwen Xing, Patrick Yubeaton, Robert Zhang, Leon Liangyu Chen, Xiaokun Chen, Steven Dillmann, Saadia Gabriel, Xunyi Jiang, Anurag Kashyap, Boxuan Li, Yein Park, Minh Pham, Sujay Sanghavi, Lin Shi, Ke Sun, Yixin Wang, Zhiwei Xu, Erica Zhang, Siyan Zhao, Wanjia Zhao, Jenia Jitsev, Alex Dimakis, Benjamin Feuer, Ludwig Schmidt

June 26, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows how to build better open training data for AI agents, and demonstrates that the resulting 100K-example dataset improves performance across multiple agentic benchmarks.

Key results

100K
training set size

Final OpenThoughts-Agent-v2 SFT dataset size

44.8%
benchmark average

OpenThinkerAgent-32B average across seven agentic benchmarks

54.0%
SWE-Bench Verified

OpenThinkerAgent-32B result

26.2%
Terminal-Bench 2.0

OpenThinkerAgent-32B result

144.76M
token budget

Matched-token comparison for the ≥5-turn filter

What the paper found

OpenThoughts Agent Data Recipes for Agentic Models presents a fully open data curation pipeline for training agentic language models, developed by researchers from UC Berkeley, Stanford, UCLA, Microsoft, Amazon, and related institutions, and benchmarked mainly with Qwen3 models. The paper’s core contribution is a six-stage supervised fine-tuning recipe, validated through more than 100 ablation experiments that show task-source choice is the dominant lever, stronger models are not necessarily better teachers, and keeping longer multi-turn rollouts improves learning even at a matched 144.76M-token budget. Using a 100K-example dataset distilled from sources such as SWE-Smith, StackExchange SuperUser, StackExchange Tezos, and IssueTasks, the authors fine-tune Qwen3-32B to OpenThinkerAgent-32B, which reaches 44.8% average accuracy across seven agentic benchmarks, including 54.0% on SWE-Bench Verified and 26.2% on Terminal-Bench 2.0, outperforming Nemotron-Terminal-32B by 3.9 percentage points on the seven-benchmark average. The data recipe also scales: OpenThoughts-Agent-v2 expands Tezos task diversity from about 902 to over 21K surface forms and keeps improving up to 100K rows, reaching 55.7% on SWE-Bench Verified-100 and 26.2% on Terminal-Bench 2.0 at 32B scale. In parallel, the reinforcement-learning study at 8B scale shows that the RL data source strongly changes behavior; the best source, pymethods2test, yields 35.67% on SWE-bench Verified, 16.02% on OT-TBLite, and 13.48% on Terminal-Bench 2.0, while the full SFT+RL pipeline reaches 27.9% average across seven benchmarks and improves over the Qwen3-8B base by about 18 points.

Original abstract

Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation pipeline for training agentic models. We conduct more than 100 controlled ablation experiments to systematically investigate each stage of the pipeline, yielding insights on the importance of task sources and diversity. We then assemble a training set of 100K examples from our pipeline and fine-tune Qwen3-32B on this dataset, which yields an average accuracy of 44.8% across seven agentic benchmarks and a 3.9 percentage point improvement over the strongest existing open data agentic model (Nemotron-Terminal-32B, 40.9%). Moreover, our training data exhibits strong scaling properties, outperforming alternative open datasets at every training set size in compute-controlled comparisons. We publicly release our training sets, data pipeline, experimental data, and models at openthoughts.ai to support future open research on agentic model training.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis