Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
AuthorsZipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
AffiliationsSouth China University of Technology · South China Normal University · Singapore Management Univeristy
Resources
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Key results
Number of HADA code-evolution iterations in the main experiments
Improvement over the existing state-of-the-art MetaBBO baseline
Improvement over the existing state-of-the-art MetaBBO baseline
Improvement over the existing state-of-the-art MetaBBO baseline
HADA's average normalized score across held-out single-objective functions
What the paper found
Hyper Algorithm Design Agent, or HADA, turns Meta-Black-Box Optimization into an executable code-evolution problem. Starting from a naive MetaBBO template, a task agent edits the optimizer and meta-level policy, while a hyper agent edits the task agent’s instructions and its own reasoning logic, creating recursive self-improvement. HADA records code patches, execution logs, domain knowledge, and normalized performance in a tree-based evolution history that balances exploration and exploitation. Using DeepSeek-V4-Pro as the coding-agent backbone, the system ran an evolution horizon of 100 iterations and produced variants for COCO-BBOB single-objective optimization, COCO-constrained optimization, WFG multi-objective optimization, and UAV path planning. Relative to existing state-of-the-art MetaBBO baselines, the reported performance leaps were 121.7% for single-objective, 114.7% for constrained, and 104.3% for multi-objective optimization. On held-out COCO-BBOB functions, HADA reached an average normalized score of 0.9674, compared with 0.8942 for GLEET. An ablation also compared DeepSeek-V4-Pro with GPT-5.5, Grok-4, and Qwen3.7-Max, while showing that allowing policy modifications was more important than merely changing the initial optimizer. The evolution logs indicate that HADA moved beyond parameter tuning toward SHADE-like search, progress-aware state representations, expanded DQN policies, and hybrid local-search mechanisms, suggesting that self-modifying coding agents can discover MetaBBO designs that are difficult to obtain through manually fixed pipelines.
Original abstract
Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm's bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While MetaBBO helps advance the performance lower bound of the resulted optimization system, it is currently handcrafted and customized case by case to adapt different optimization problems, which inevitably introduces inherent subjectivity and hence restricts the performance upper bound and usability in practice. In this paper, we address this issue by regarding MetaBBO's design loop as coding task, where we could introduce openendedness into MetaBBO with recursive self-improvement capability of advanced coding agents. Specifically, we propose a dual-agent framework: i) a task agent continuously refines the codebase of a target MetaBBO approach through code evolution; ii) a hyper agent progressively modifies the task agent and itself to provide open-ended design behavior; iii) the evolved MetaBBO codebase is evaluated and all in-execution information is fed back to the agents for recursive self-referential improvement. As a result, given a naive MetaBBO template, our framework automates a design evolution and finds novel variants superior to up-to-date human-made MetaBBO baselines. Surprisingly, the experimental results also demonstrate that our framework supports fast adaption across different optimization domains. Solid interpretation analysis further reveals interesting design principles emerge in such open-ended process. This work serves as the first exploration on automating design of complex learning-assisted optimization algorithms.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.
Adaptively Incorporating Directional Hints into Zeroth-Order Optimization
Alexander Ryabchenko, Jian Qian, Wenlong Mou
A new zeroth-order optimizer adaptively uses unreliable directional hints to approach first-order performance without needing to know how good those hints are.