AI Research Agents Narrow Scientific Exploration
AuthorsYixuan Tang, Yi Yang
Resources
This paper shows that today’s AI research agents mostly generate narrow, local variations on existing science rather than broadly expanding the frontier of discovery.
Key results
non-empty structured research ideas retained from all generation runs
DBLP papers from ICLR, NeurIPS, and ICML, 2019-2025
active citation-defined areas used in the main analysis
AI-generated ideas compared with human-authored papers from the same area
AI-generated ideas versus their starting seed papers
What the paper found
This paper studies AI research agents from AI Scientist, ResearchAgent, AgentLaboratory, and a zero-shot baseline as scientific search systems, using six LLMs from Qwen, Llama, and Gemma to generate 37,802 valid research ideas from 34,698 papers in ICLR, NeurIPS, and ICML spanning 2019 to 2025. Across 19 citation-defined research areas, the authors find a consistent pattern: AI-generated ideas are more concentrated than human-authored papers, with within-area cosine similarity of 0.82–0.84 versus 0.77 for human papers, and a centroid distance of 0.091 versus 0.121. The ideas also stay closer to the seed literature than later human follow-on work, at 0.92 similarity versus 0.88, suggesting local extrapolation rather than broad exploration. To probe impact, the paper matches AI ideas to 2,359 semantically similar human papers and finds those papers receive 50.4 citations on average, below the 54.9 citation same-area baseline by 4.47 citations. Finally, the content analysis shows AI agents mostly recombine methods rather than invent new questions: 85.1% of ideas reuse the seed literature’s research question, while only 62.6% reuse the technical method, indicating that current agentic novelty is concentrated at the method-combination level rather than at problem discovery.
Original abstract
AI research agents can now generate research ideas, design experiments, run code, and draft papers, raising the possibility of large-scale AI-assisted scientific discovery. Many current agent frameworks explicitly encourage the generation of novel and high-impact ideas. Yet it remains unclear whether AI-assisted ideation broadens scientific exploration or mainly concentrates around existing work. We study AI research agents as scientific search systems. Using four AI research-agent frameworks and six large language models, we generate 37,802 scientific ideas from shared seed literature across citation-defined research areas in AI and machine learning. We then compare the resulting AI ideas against human-authored papers from the same research areas, follow-on human research emerging from the same seed literature, and the seed literature itself. Across experiments, four consistent patterns emerge. First, AI-generated ideas are substantially more concentrated than human-authored papers from the same research areas. Second, AI-generated ideas remain much closer to their starting literature than later human follow-on work does. Third, papers most similar to AI-generated ideas tend to receive lower subsequent citations. Fourth, when AI-generated ideas differ from prior work, the differences arise primarily from recombining existing technical methods rather than introducing fundamentally new research questions. Overall, current AI research agents appear better suited to local elaboration than to broadening scientific exploration.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.