NTH

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

AuthorsJisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning, Xiyao Wang, Yifan Shen, Heng Wang, Yuqing Jian, Xiaoxia Wu, Ben Athiwaratkun, Pan Lu, Jiaxuan You, Bingxin Zhao

July 22, 2026 2 min read
Watch on YouTube
The one-line take

PaperPilot turns scientific paper search into a controllable multi-turn workflow, making literature discovery more editable, inspectable, and effective.

Key results

77.0
PAPER PILOT-9B multi-turn Hit@5

Retrieval score after training, compared with 58.0 for the base Qwen3.5-9B toolset agent.

59.4
PAPER PILOT-9B multi-turn MRR

Mean reciprocal rank, compared with 47.5 for the base Qwen3.5-9B toolset agent.

32.5
PAPER PILOT-9B multi-turn nDCG@10

Rank-sensitive retrieval score, compared with 26.8 for the base Qwen3.5-9B toolset agent.

0%
Workflow execution error rate

PAPER PILOT-9B error rate after workflow-induction training, reduced from 9.5%.

2,723
Training anchor-query cases

Cases used to generate teacher workflow trajectories across five search directions.

74.7%
Human-session success rate

Percentage of human-study sessions judged satisfactory.

What the paper found

Researchers from the University of Illinois Urbana-Champaign, Together AI, the University of Pennsylvania, and Stanford introduce PAPER PILOT, a multi-turn scientific literature-search agent that treats retrieval as workflow induction rather than fixed query matching. Given an anchor paper and an evolving request, the agent builds an executable, type-consistent DAG from operators for keyword search, citation expansion, filtering, scoring, reranking, and evidence extraction, then edits that DAG in response to clarification feedback. Training combines supervised workflow imitation with IPO-style DPO preference optimization, using 2,723 anchor-query cases, 5,540 workflow-supervision examples, and 1,733 hard chosen–rejected pairs created by corrupting valid workflows. On a benchmark spanning five search directions, PAPER PILOT-9B improves the multi-turn Qwen3.5-9B toolset baseline from 58.0 to 77.0 Hit@5, from 47.5 to 59.4 MRR, and from 26.8 to 32.5 nDCG@10, while reducing workflow execution errors from 9.5% to 0%. It also reaches 74.7% human-session success and a 4.2 out-of-5 question-satisfaction score. The system is competitive with much larger models: GPT-5.4 with the PAPER PILOT toolset achieves 84.0 Hit@5, while OpenAI DeepResearch is evaluated as a one-shot baseline. The central result is that explicit, editable workflows translate user feedback into concrete retrieval operations, improving controllability, ranking quality, and interaction alignment.

Original abstract

Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, and evolve through interaction. Existing search agents typically rely on fixed pipelines or implicit language-only reasoning, making their search strategies difficult to control, inspect, and refine. We introduce PaperPilot, a multi-turn literature search agent that frames scientific search as workflow induction. Given an anchor paper and a user query, PaperPilot constructs an executable DAG of paper-search operators, including keyword search, citation expansion, filtering, scoring, reranking, and evidence extraction. User feedback is then used to refine both the query and the workflow itself. We train PaperPilot with supervised workflow imitation and preference optimization over controlled workflow corruptions. Experiments show that PaperPilot-9B improves over the base Qwen3.5-9B toolset agent under multi-turn interaction, increasing Hit@5 from 58.0 to 77.0, MRR from 47.5 to 59.4, and nDCG@10 from 26.8 to 32.5, while reducing workflow execution errors from 9.5% to 0%. These results show that explicit, editable search workflows provide an effective and controllable interface for aligning literature search agents with complex scientific intent.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis