OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
AuthorsSiyuan Li, Jiabao Pan, Yumou Liu, Zhuoli Ouyang, Xin Jin, Xinglong Xu, Jingxuan Wei, Shengye Pang, Jintao Che, Xuanhe Zhou, Conghui He, Cheng Tan
Resources
OmniOpt organizes the chaotic world of modern optimizers into a unified taxonomy and benchmark so researchers can compare methods more systematically.
Key results
C4 short-context screening at 1B parameters
C4 short-context screening at 1B parameters
C4 at 32k tokens in the sequence-length ablation
Muon core-mechanism ablation after removing second-moment scaling
What the paper found
OmniOpt, developed by researchers from Shanghai AI Laboratory, Shanghai University, Westlake University, Shanghai Jiao Tong University, UCAS, Zhejiang University, and SUSTech, reframes optimizer design as a mechanism-aware system rather than a flat leaderboard. The paper unifies 108 optimizers with a five-stage meta-pipeline, an LMO-driven four-axis geometry, and a dual taxonomy that separates method families from effect objectives. In benchmarking 24 representative optimizers, it shows that no single method dominates: on C4 short-context pretraining, APOLLO reaches the best 1B perplexity at 13.53, while Muon and MARS-Shampoo are close behind at 13.72; yet the cheapest runtime comes from Lion, and the lowest optimizer-state memory comes from AdaFactor. Across FineWeb-Edu 32k long-context runs, SOAP is the most stable cross-scenario optimizer, holding the top perplexity position in 7 of 8 scenarios, while MARS-AdamW is the most stable AdamW-style enhancement. APOLLO is the sharpest cautionary case: it is best at 256 tokens with 13.53 perplexity but collapses to 35.40 at 32k, showing that aggressive state compression is rank-bounded and can fail as context length grows. The Muon ablation isolates Newton–Schulz orthogonalization as the core gain, recovering perplexity from 70.74 to 16.86 after removing AdamW-style second-moment scaling, and then improving to 13.58 when symmetric learning-rate scaling and post-orthogonalization Nesterov are combined at 1B. On CIFAR100, optimizer choice is backbone-dependent: AdaBelief leads ResNet50 with 80.53%, Muon leads DeiT-S with 77.38%, and Adan leads CAFormer-S12 with 84.89%, reinforcing the paper’s central claim that optimizer geometry must match model topology and training constraints.
Original abstract
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning budget, and task diversity, yet the landscape of over one hundred methods remains fragmented. We therefore present OmniOpt, a unified survey and benchmark cookbook of optimizers for the research community. OmniOpt rests on four coupled components. First, we treat every optimizer update as a structured transformation through a five-stage meta-pipeline, and show that most methods engage only one or two of these stages. Second, we use norm-constrained linear minimization oracles (LMOs) to unify different optimizers. Third, these two views ground a dual-dimension taxonomy, one dimension assigning each method to a mechanism family and the other recording the measurable training objectives it aims to improve. Fourth, and at the core of this paper, we instantiate the full taxonomy in a unified cross-domain benchmark spanning representative optimizers, model scales, and training regimes from language model pretraining to image classification, systematically analyzing each method family across multiple effect objectives and laying out their trade-offs. OmniOpt thus supplies the research community with an operational coordinate system for selecting optimizers under explicit mechanism and objective assumptions, and charts a direction for the future development of the optimizer community.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.