Multi-Turn Agentic Scientific Literature Search via Workflow Induction
AuthorsJisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning, Xiyao Wang, Yifan Shen, Heng Wang, Yuqing Jian, Xiaoxia Wu, Ben Athiwaratkun, Pan Lu, Jiaxuan You, Bingxin Zhao
Resources
PaperPilot turns scientific paper search into a controllable multi-turn workflow, making literature discovery more editable, inspectable, and effective.
Key results
Retrieval score after training, compared with 58.0 for the base Qwen3.5-9B toolset agent.
Mean reciprocal rank, compared with 47.5 for the base Qwen3.5-9B toolset agent.
Rank-sensitive retrieval score, compared with 26.8 for the base Qwen3.5-9B toolset agent.
PAPER PILOT-9B error rate after workflow-induction training, reduced from 9.5%.
Cases used to generate teacher workflow trajectories across five search directions.
Percentage of human-study sessions judged satisfactory.
What the paper found
Researchers from the University of Illinois Urbana-Champaign, Together AI, the University of Pennsylvania, and Stanford introduce PAPER PILOT, a multi-turn scientific literature-search agent that treats retrieval as workflow induction rather than fixed query matching. Given an anchor paper and an evolving request, the agent builds an executable, type-consistent DAG from operators for keyword search, citation expansion, filtering, scoring, reranking, and evidence extraction, then edits that DAG in response to clarification feedback. Training combines supervised workflow imitation with IPO-style DPO preference optimization, using 2,723 anchor-query cases, 5,540 workflow-supervision examples, and 1,733 hard chosen–rejected pairs created by corrupting valid workflows. On a benchmark spanning five search directions, PAPER PILOT-9B improves the multi-turn Qwen3.5-9B toolset baseline from 58.0 to 77.0 Hit@5, from 47.5 to 59.4 MRR, and from 26.8 to 32.5 nDCG@10, while reducing workflow execution errors from 9.5% to 0%. It also reaches 74.7% human-session success and a 4.2 out-of-5 question-satisfaction score. The system is competitive with much larger models: GPT-5.4 with the PAPER PILOT toolset achieves 84.0 Hit@5, while OpenAI DeepResearch is evaluated as a one-shot baseline. The central result is that explicit, editable workflows translate user feedback into concrete retrieval operations, improving controllability, ranking quality, and interaction alignment.
Original abstract
Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, and evolve through interaction. Existing search agents typically rely on fixed pipelines or implicit language-only reasoning, making their search strategies difficult to control, inspect, and refine. We introduce PaperPilot, a multi-turn literature search agent that frames scientific search as workflow induction. Given an anchor paper and a user query, PaperPilot constructs an executable DAG of paper-search operators, including keyword search, citation expansion, filtering, scoring, reranking, and evidence extraction. User feedback is then used to refine both the query and the workflow itself. We train PaperPilot with supervised workflow imitation and preference optimization over controlled workflow corruptions. Experiments show that PaperPilot-9B improves over the base Qwen3.5-9B toolset agent under multi-turn interaction, increasing Hit@5 from 58.0 to 77.0, MRR from 47.5 to 59.4, and nDCG@10 from 26.8 to 32.5, while reducing workflow execution errors from 9.5% to 0%. These results show that explicit, editable search workflows provide an effective and controllable interface for aligning literature search agents with complex scientific intent.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.