The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs
AuthorsXu Wan, Speed Zhu, Jianwei Cai, Guang Chen, XiMing Huang, Wiggin Zhou, Mingyang Sun
Resources
This paper treats LLM reasoning like an economy, using shadow prices to decide which queries deserve more compute so limited budget yields the best overall accuracy.
Key results
On the Balanced reasoning stream, CLEAR (Lambert) improves accuracy over Uniform by 11.6 points at the 256-token budget.
On the Mostly-Easy stream, CLEAR (Lambert) improves accuracy over Uniform by 24.0 points at the 256-token budget.
On the U-Shaped stream, CLEAR (Lambert) improves accuracy over Uniform by 14.2 points at the 256-token budget.
For code generation, CLEAR improves over Uniform by up to 6.5 points, with the largest gain reported on MBPP+.
What the paper found
The paper, from researchers at Zhejiang University, Tencent HY Team, and Peking University, reframes LLM inference as a constrained economic allocation problem rather than a per-query decoding heuristic. Using Qwen2.5-Math-7B and Qwen3-30B-A3B-Instruct as frozen backbones, the authors show that reasoning utility follows an S-shaped compute-utility curve with three phases: Strict, Surge, and Ample. They model each query’s latent utility as a shifted-surge function with threshold τ, initial velocity α, and decay β, then solve the batch-level optimization with a global shadow price λ that equalizes marginal utility across active queries. This yields a closed-form token allocation via the principal branch of the Lambert W function and a plug-and-play system called CLEAR, which combines threshold prediction with bisection-based market clearing and rational abandonment for queries whose expected surplus is negative. On mixed reasoning streams built from MATH-500, AMC-23, AIME-24, AIME-25, Minerva, and OlympiadBench, CLEAR consistently improves the cost-accuracy Pareto frontier; at a 256-token budget it raises accuracy by 11.6 points on a balanced stream, 24.0 points on mostly-easy traffic, and 14.2 points on a U-shaped stream, reaching up to a 3× gain in global accuracy over uniform allocation in the most resource-scarce settings. The same allocation principle also transfers to code generation, improving HumanEval+, MBPP+, and BigCodeBench by 3.6 to 6.5 points over uniform budgeting.
Original abstract
Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets. In this work, we formulate inference budget allocation as a global constrained optimization problem governed by economic principles. By modeling per-query reasoning utility with a shifted-surge function, we derive an optimal allocation policy based on a global shadow price that equilibrates marginal utility under resource scarcity. Based on this theory, we propose Constrained Latent-utility Equilibrium Allocation for Reasoning (CLEAR). It performs rational abandonment and reallocates resources from insolvent queries to solvable queries near their emergence thresholds. Extensive experiments on several reasoning tasks with different traffic streams demonstrate that CLEAR significantly improves the Pareto frontier of total token cost versus mean accuracy. In resource-scarce regimes, CLEAR achieves up to a 3x improvement in global accuracy compared to uniform allocation.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.