NTH

Bandits in Prod: Hyperparameter Optimization at Inference Time

AuthorsLouis Abraham, Tuan-Anh Nguyen, Nicolas Devatine

September 10, 2026 2 min read
Watch on YouTube
The one-line take

This paper treats live configuration tuning for models and agents as a bandit problem, enabling systems to learn better inference-time settings directly from noisy production feedback.

Key results

0.5
IMOSS active-set exponent

Experimental setting for active-set growth as t^β.

5000
Discrete benchmark budget

Pulls used to compare IMABO oracles on segment, credit-g, and numerai28.6.

19%
Returned feedback under censoring

Approximate fraction of pulls that produced an observable reward in delayed-feedback experiments.

60
Delay-aware active arms

Approximate active-set size under delayed and censored feedback.

100
Naive active arms

Approximate active-set size reached by the delay-blind policy.

What the paper found

This paper introduces online hyperparameter optimization, or OHPO, for production systems where configurations can be evaluated only on live requests with noisy feedback. It models each mixed or conditional configuration—such as retrieval depth, prompt template, decoding temperature, or model choice—as an arm in an infinitely many-armed bandit, and proposes IMABO, which separates allocation among known configurations from an oracle that proposes new ones. Its IMOSS policy is anytime, restart-free, and grows the active set as t^β; with β = 0.5, it achieves an expected cumulative quantile-regret bound of O(pρ^-1/β + T^(1+β)/2). The framework combines IMOSS with random sampling, Tree-structured Parzen Estimation, incumbent mutation using KL-UCB and Parzen Estimation, or the TabPFN-3 tabular foundation model. Across HPOBench and OpenML tasks, the learned oracles consistently beat uniform proposals; on 5000-pull discrete benchmarks, mutate-KL×PE achieved the lowest regret on segment, credit-g, and numerai28.6. In a 5000-question HotpotQA retrieval-augmented QA stream, configurations varied across top-k retrieval, prompts, temperature, and models including DeepSeek V4 Flash and Meta’s Llama 3.2 1B Instruct; IMOSS-TabPFN achieved the lowest online average regret. For delayed feedback, where only about 19% of pulls returned rewards, the delay-aware rule concentrated exploration on about 60 active arms instead of about 100 for the naive policy, improving ranking reliability.

Original abstract

Many production systems can assess a configuration only by using it on live requests and observing noisy feedback. Modern agentic systems are a prominent example, with inference-time choices such as model selection, retrieval depth, prompting strategy, and decoding temperature, yet often with no representative validation data. We formalize this setting as Online Hyperparameter Optimization (OHPO) and cast it as an infinitely many-armed bandit over mixed and conditional search spaces. We introduce IMABO, a general framework that combines any bandit policy for choosing among already sampled configurations with any oracle for proposing new ones. We instantiate it with IMOSS, a restart-free anytime policy whose active set grows as $t^β$, and prove an expected cumulative quantile-regret bound of $O(p_ρ^{-1/β} + T^{(1+β)/2})$, where $β\in(0,1)$ controls active-set growth and $p_ρ$ lower-bounds the probability that a proposed configuration falls in the top-$ρ$ fraction of the search space. We combine IMOSS with three practical oracles: a Tree-structured Parzen Estimator, an incumbent-mutation oracle driven by a per-coordinate bandit, and a pretrained tabular foundation model, all three improving over the uniform random oracle baseline. IMABO outperforms all baselines in terms of regret across diverse OHPO settings, from tuning classical machine-learning models to configuring LLM-based agents. Our implementation is available at https://github.com/Tiime-Software/IMABO.

Read the original paper

More in Optimization

Browse all 36 papers →