NTH

Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero

AuthorsZipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao

AffiliationsSouth China University of Technology · South China Normal University · Singapore Management Univeristy

September 29, 2026 3 min read
Watch on YouTube
The one-line take

A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.

Key results

100
Evolution horizon

Number of HADA code-evolution iterations in the main experiments

121.7%
Single-objective performance leap

Improvement over the existing state-of-the-art MetaBBO baseline

114.7%
Constrained performance leap

Improvement over the existing state-of-the-art MetaBBO baseline

104.3%
Multi-objective performance leap

Improvement over the existing state-of-the-art MetaBBO baseline

0.9674
Held-out COCO-BBOB average score

HADA's average normalized score across held-out single-objective functions

What the paper found

Hyper Algorithm Design Agent, or HADA, turns Meta-Black-Box Optimization into an executable code-evolution problem. Starting from a naive MetaBBO template, a task agent edits the optimizer and meta-level policy, while a hyper agent edits the task agent’s instructions and its own reasoning logic, creating recursive self-improvement. HADA records code patches, execution logs, domain knowledge, and normalized performance in a tree-based evolution history that balances exploration and exploitation. Using DeepSeek-V4-Pro as the coding-agent backbone, the system ran an evolution horizon of 100 iterations and produced variants for COCO-BBOB single-objective optimization, COCO-constrained optimization, WFG multi-objective optimization, and UAV path planning. Relative to existing state-of-the-art MetaBBO baselines, the reported performance leaps were 121.7% for single-objective, 114.7% for constrained, and 104.3% for multi-objective optimization. On held-out COCO-BBOB functions, HADA reached an average normalized score of 0.9674, compared with 0.8942 for GLEET. An ablation also compared DeepSeek-V4-Pro with GPT-5.5, Grok-4, and Qwen3.7-Max, while showing that allowing policy modifications was more important than merely changing the initial optimizer. The evolution logs indicate that HADA moved beyond parameter tuning toward SHADE-like search, progress-aware state representations, expanded DQN policies, and hybrid local-search mechanisms, suggesting that self-modifying coding agents can discover MetaBBO designs that are difficult to obtain through manually fixed pipelines.

Original abstract

Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm's bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While MetaBBO helps advance the performance lower bound of the resulted optimization system, it is currently handcrafted and customized case by case to adapt different optimization problems, which inevitably introduces inherent subjectivity and hence restricts the performance upper bound and usability in practice. In this paper, we address this issue by regarding MetaBBO's design loop as coding task, where we could introduce openendedness into MetaBBO with recursive self-improvement capability of advanced coding agents. Specifically, we propose a dual-agent framework: i) a task agent continuously refines the codebase of a target MetaBBO approach through code evolution; ii) a hyper agent progressively modifies the task agent and itself to provide open-ended design behavior; iii) the evolved MetaBBO codebase is evaluated and all in-execution information is fed back to the agents for recursive self-referential improvement. As a result, given a naive MetaBBO template, our framework automates a design evolution and finds novel variants superior to up-to-date human-made MetaBBO baselines. Surprisingly, the experimental results also demonstrate that our framework supports fast adaption across different optimization domains. Solid interpretation analysis further reveals interesting design principles emerge in such open-ended process. This work serves as the first exploration on automating design of complex learning-assisted optimization algorithms.

Read the original paper

More in Optimization

Browse all 36 papers →