NTH

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

AuthorsGuijin Son, Seungyeop Yi, Minju Gwak, Hyunwoo Ko, Wongi Jang, Youngjae Yu

June 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper builds the largest dataset so far for research-level math and shows that filtered AI-generated reasoning traces can help train stronger math models.

Key results

14,056
problem corpus size

ResearchMath-14K final number of research-level mathematical problems

1,233
source documents

Academic sources collected for the extraction pipeline

220K
reasoning trajectories

ResearchMath-Reasoning teacher traces released with the dataset family

54.0%
fake-reference traces

Share of 720 sampled traces containing at least one fake reference

9.2
average fine-tuning gain

Average improvement in percentage points after LoRA fine-tuning on filtered reasoning traces

What the paper found

ResearchMath-14K scales research-level mathematics by converting open problems from the literature into agent-generated training data: the authors curate 1,233 academic sources from arXiv, zbMATH, workshops, and repositories, then use a two-stage extractor-refiner pipeline to produce 14,056 self-contained problems, the largest public research-grade math corpus to date. They also release ResearchMath-Reasoning, 220K teacher trajectories from GPT-OSS-120B and Qwen3-30B-A3B, and show that trace quality is fragile: in 720 sampled traces, 54.0% contain at least one fake reference, while newer model generations emit 5.6× more reference-like mentions and 5.0× more fabricated references per trace. A GPT-5.5 self-containment audit raises extracted questions from 67.2% to 94.2%, and duplicate filtering uses Qwen3-Embedding-8B with a 0.9 similarity threshold. Most importantly, LoRA fine-tuning Qwen3-4B, Qwen3-8B, and Qwen3-30B-A3B on a 5,000-trace filtered subset improves downstream performance by 9.2 percentage points on average, outperforming a DASD-Thinking control in 8 of 9 benchmark cells and indicating that wrong-but-reasonable research attempts can still provide useful supervision after filtering hallucinated citations and non-attempts.

Original abstract

The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully engage with such problems without human intervention. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of $14{,}056$ problems curated from academic sources via a multi-agent pipeline, making it the largest collection of research-level mathematical problems to date. We further generate ResearchMath-Reasoning, $220$K teacher trajectories from two open models, where we observe recurring avoidance behaviors such as non-attempts and fabricated references. Interestingly, across eight open-weight models, newer generations produce $5.6\times$ more references and $5.0\times$ more fake references per trace. After agentic filtering of ResearchMath-Reasoning, fine-tuning Qwen3 models from 4B to 30B parameters improves over base models by $9.2$ points on average. This shows that filtered open-problem attempts can provide useful supervision even without fully correct reasoning traces. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis