ResearchMath-14K: Scaling Research-Level Mathematics via Agents
AuthorsGuijin Son, Seungyeop Yi, Minju Gwak, Hyunwoo Ko, Wongi Jang, Youngjae Yu
Resources
This paper builds the largest dataset so far for research-level math and shows that filtered AI-generated reasoning traces can help train stronger math models.
Key results
ResearchMath-14K final number of research-level mathematical problems
Academic sources collected for the extraction pipeline
ResearchMath-Reasoning teacher traces released with the dataset family
Share of 720 sampled traces containing at least one fake reference
Average improvement in percentage points after LoRA fine-tuning on filtered reasoning traces
What the paper found
ResearchMath-14K scales research-level mathematics by converting open problems from the literature into agent-generated training data: the authors curate 1,233 academic sources from arXiv, zbMATH, workshops, and repositories, then use a two-stage extractor-refiner pipeline to produce 14,056 self-contained problems, the largest public research-grade math corpus to date. They also release ResearchMath-Reasoning, 220K teacher trajectories from GPT-OSS-120B and Qwen3-30B-A3B, and show that trace quality is fragile: in 720 sampled traces, 54.0% contain at least one fake reference, while newer model generations emit 5.6× more reference-like mentions and 5.0× more fabricated references per trace. A GPT-5.5 self-containment audit raises extracted questions from 67.2% to 94.2%, and duplicate filtering uses Qwen3-Embedding-8B with a 0.9 similarity threshold. Most importantly, LoRA fine-tuning Qwen3-4B, Qwen3-8B, and Qwen3-30B-A3B on a 5,000-trace filtered subset improves downstream performance by 9.2 percentage points on average, outperforming a DASD-Thinking control in 8 of 9 benchmark cells and indicating that wrong-but-reasonable research attempts can still provide useful supervision after filtering hallucinated citations and non-attempts.
Original abstract
The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully engage with such problems without human intervention. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of $14{,}056$ problems curated from academic sources via a multi-agent pipeline, making it the largest collection of research-level mathematical problems to date. We further generate ResearchMath-Reasoning, $220$K teacher trajectories from two open models, where we observe recurring avoidance behaviors such as non-attempts and fabricated references. Interestingly, across eight open-weight models, newer generations produce $5.6\times$ more references and $5.0\times$ more fake references per trace. After agentic filtering of ResearchMath-Reasoning, fine-tuning Qwen3 models from 4B to 30B parameters improves over base models by $9.2$ points on average. This shows that filtered open-problem attempts can provide useful supervision even without fully correct reasoning traces. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.