NTH

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

AuthorsJian Xie, Tianhe Lin, Zilu Wang, Yuting Ning, Yuekun Yao, Tianci Xue, Zhehao Zhang, Zhongyang Li, Kai Zhang, Yufan Wu, Shijie Chen, Boyu Gou, Mingzhe Han, Yifei Wang, Vint Lee, Xinpeng Wei, Xiangjun Wang, Yu Su, Huan Sun

June 11, 2026 2 min read
Watch on YouTube
The one-line take

QUEST shows that fully synthetic training can produce open deep research agents that rival proprietary systems on long-horizon search, citation grounding, and synthesis tasks.

Key results

35B
Quest model scale

Largest released Quest model

8K
Quest-8K

Synthetic training set used for SFT and RL

64.6
BrowseComp

Quest-35B score on BrowseComp

30.7
Mind2Web 2

Quest-35B score on Mind2Web 2

48.2
DeepResearch Bench

Quest-35B score on DeepResearch Bench

80.8
GAIA

Quest-35B score on GAIA

What the paper found

Quest, from The Ohio State University with Amazon AGI SF Lab contributors, is an open deep-research agent family spanning 2B to 35B parameters that targets three coupled capabilities: fact seeking, citation grounding, and long-form report synthesis. The core novelty is a fully synthetic training pipeline, Quest-8K, built from rubric trees that convert web-synthesized tasks into verifiable rewards without human annotation, plus a structured Context State that compresses search history into trusted, untrusted, and uncertain claims so the agent can continue across long horizons. The training recipe combines mid-training, supervised fine-tuning, and GRPO-style reinforcement learning, while also using inline citations and a fact-checking reward. On eight benchmarks, Quest-35B reaches 64.6 on BrowseComp, 30.7 on Mind2Web 2, 48.2 on DeepResearch Bench, and 80.8 on GAIA, surpassing OpenAI DeepResearch on Mind2Web 2 and DeepResearch Bench and matching or exceeding GPT-5 on GAIA. The paper also shows that even Quest-2B-SFT remains surprisingly strong on fact-seeking tasks, but open-ended synthesis still benefits most from the full MT+SFT+RL recipe.

Original abstract

Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how humans interact with information. However, frontier systems remain proprietary, while existing open agents often generalize poorly across different task types, leaving unclear how to train a broadly capable deep research agent. We release QUEST, a family of open models (ranging from 2B to 35B) that serve as general-purpose deep research agents designed to handle a wide range of long-horizon search tasks, with strong capabilities in fact seeking, citation grounding, and report synthesis. To build QUEST, we propose an effective training recipe combining mid-training, supervised fine-tuning, and reinforcement learning. Central to this recipe is a curated data synthesis pipeline based on unified rubric trees, which applies to different task types and enables synthesizing training data with verifiable rewards without human annotation. In addition, QUEST incorporates a built-in context management mechanism that enables effective long-horizon reasoning and knowledge synthesis. Using only 8K synthesized tasks, QUEST approaches or even surpasses frontier closed-source agents across eight deep research benchmarks spanning diverse task types, and achieves the best overall performance among recent open-weight agents. We released everything: models, data, and training scripts.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis