Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
AuthorsAdam Fisch, Shubhendu Trivedi, Fantine Huot, William W. Cohen, Michael Kaisers, Mirella Lapata, Kate Larson, Jacob Eisenstein
Resources
A smart router decides when it is worth paying for better model-quality estimates before sending each query to the best specialist.
Key results
Mathematical reasoning problems used to evaluate costly estimation with partial reasoning traces.
Number of candidate models in the large-scale model-selection benchmark.
Average regret-plus-inspection-cost score across tested inspection costs.
Average regret-plus-inspection-cost score for retrieval-augmented specialists.
Average regret-plus-inspection-cost score across the 123-model routing benchmark.
What the paper found
Pandora’s AI Model Routing Box treats model selection as a Pandora’s Box search problem: every specialist receives a cheap, noisy value estimate, and the router selectively pays for a refined estimate only when its expected value of information exceeds the inspection cost. Under a Gaussian signal model, closed-form reservation prices determine which specialist to inspect next and when to stop, while a non-obligatory inspection variant can commit to an unopened model. The centralized Pandora’s Router was tested on MATH, retrieval-augmented generation, and EmbedLLM, which contains 123 routing models, including systems such as Gemma3-4B and Gemini-3.1-Flash-Lite. The MATH evaluation used a 16,512-problem corpus, with costly estimation based on the first 20 reasoning tokens and Gemini-2.5-Flash-Lite; retrieval experiments compared Wikipedia and PubMed specialists. Averaged across inspection costs, Pandora’s Router achieved regret-plus-cost scores of 0.105 on MATH, 0.118 on RAG, and 0.386 on EmbedLLM, matching or improving on exhaustive estimation while querying the expensive estimator selectively. The paper also introduces Pandora’s Bidder, where each specialist decides whether to buy a refined self-assessment before accepting a posted price; this improves allocative efficiency when competing estimates are accurate, but noisy or weak bids can raise the strategic specialist’s utility while reducing overall welfare.
Original abstract
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of others.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.