NTH

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

AuthorsAdam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein

August 25, 2026 2 min read
Watch on YouTube
The one-line take

A smart router decides when it is worth paying for better model-quality estimates before sending each query to the best specialist.

Key results

16,512
MATH corpus size

Mathematical reasoning problems used to evaluate costly estimation with partial reasoning traces.

123
EmbedLLM routing models

Number of candidate models in the large-scale model-selection benchmark.

0.105
Pandora’s Router on MATH

Average regret-plus-inspection-cost score across tested inspection costs.

0.118
Pandora’s Router on RAG

Average regret-plus-inspection-cost score for retrieval-augmented specialists.

0.386
Pandora’s Router on EmbedLLM

Average regret-plus-inspection-cost score across the 123-model routing benchmark.

What the paper found

Pandora’s AI Model Routing Box treats model selection as a Pandora’s Box search problem: every specialist receives a cheap, noisy value estimate, and the router selectively pays for a refined estimate only when its expected value of information exceeds the inspection cost. Under a Gaussian signal model, closed-form reservation prices determine which specialist to inspect next and when to stop, while a non-obligatory inspection variant can commit to an unopened model. The centralized Pandora’s Router was tested on MATH, retrieval-augmented generation, and EmbedLLM, which contains 123 routing models, including systems such as Gemma3-4B and Gemini-3.1-Flash-Lite. The MATH evaluation used a 16,512-problem corpus, with costly estimation based on the first 20 reasoning tokens and Gemini-2.5-Flash-Lite; retrieval experiments compared Wikipedia and PubMed specialists. Averaged across inspection costs, Pandora’s Router achieved regret-plus-cost scores of 0.105 on MATH, 0.118 on RAG, and 0.386 on EmbedLLM, matching or improving on exhaustive estimation while querying the expensive estimator selectively. The paper also introduces Pandora’s Bidder, where each specialist decides whether to buy a refined self-assessment before accepting a posted price; this improves allocative efficiency when competing estimates are accurate, but noisy or weak bids can raise the strategic specialist’s utility while reducing overall welfare.

Original abstract

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of others.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis