AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
AuthorsKailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AffiliationsUniversity of Science and Technology of China · Yuanbao Team, Tencent
Resources
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Key results
AdaTutoRank's overall score across the answer-level RAG and deep-research benchmarks.
AdaTutoRank exceeds the strongest answer-level baseline by 1.44 points.
AdaTutoRank's short-form setwise score, compared with 46.80 for RubricRanker.
The λ value balancing GRPO outcome advantage and tutoring distillation advantage.
Average search calls per query induced by AdaTutoRank on HealthBench.
Average search calls per query induced by AdaTutoRank on DeepResearchBench.
What the paper found
AdaTutoRank, developed for Tencent Yuanbao, reframes retrieval reranking as document-set composition rather than independent relevance sorting. Its Adaptive Tutoring Optimization, or ATO, uses a three-level hierarchy of nine rubric dimensions—document relevance, authenticity, and quality; set complementarity, redundancy, and conflict; and global completeness, density, and reachability—to create silver labels, reinforcement rewards, and dense token-level distillation signals. Starting from Qwen3-8B, the method adaptively tutors each rollout: strong selections receive rubrics, medium ones receive corrective reflections, and weak ones receive a better sibling set generated by a frozen DeepSeek-V4-Pro teacher; DeepSeek-V4-Flash supplies the 0–10 rubric rewards. The combined GRPO and tutoring advantage uses λ = 0.7, while inference requires only the query and candidate documents, with no hints or reference answer. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank reaches an answer-level overall score of 45.28, exceeding RubricRanker by 1.44 points, and achieves 48.97 on short-form SetwiseEvalKit versus 46.80 for RubricRanker. Using top-20 retrieval candidates, it also reduces search effort, averaging 3.03 calls on HealthBench and 3.39 on DeepResearchBench. Downstream experiments use Llama-3.1-8B for RAG generation and an open deep-research agent, showing that better evidence composition—not merely a stronger generator—drives the gains.
Original abstract
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.
Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
Ej Zhou, Suchir Salhan, Catherine Arnett, Anna Korhonen
Independently trained language models may spontaneously learn representations that can be rotated and transferred across languages, suggesting multilingual capabilities without joint training.