NTH
Research collection

Optimization research

Research on algorithms that train machine learning models. Compare convergence, stability, memory use, and computational cost.

36 papers · Latest edition September 30, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Optimization papers

Newest editions first.

07Optimizer

Bandits in Prod: Hyperparameter Optimization at Inference Time

Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine

This paper treats live configuration tuning for models and agents as a bandit problem, enabling systems to learn better inference-time settings directly from noisy production feedback.

Read analysis
13Optimizer

Approximate Muon with low-rank adapters

Ben Anson, Conor Houghton, Edward Milsom

sMuon adapts the Muon optimizer to low-rank fine-tuning, offering a simpler and sometimes more effective alternative for LoRA-based training.

Read analysis
18Optimizer

The Loss Does Not See the Basis, but Adam Does

Devender Singh

The paper argues that Adam’s coordinate-wise behavior breaks a hidden symmetry that lets gradient descent favor low-rank solutions, showing that optimizer geometry—not just the loss—determines which answer a model learns.

Read analysis
23Optimizer

Hyperball May Not Be a Free Lunch

Yihao Xiao, Jialong Sun, Zitian Gao, Zeming Wei, Chutian Wang, Ran Tao, Jiaye Teng, Bryan Dai

Hyperball optimizers may not win because of better update directions after all—their apparent advantage largely depends on how their effective learning rate evolves over training.

Read analysis
25Optimizer

ISO: An RLVR-Native Optimization Stack

Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang

ISO speeds up RLVR training by preserving a model’s singular-value spectrum while optimizing the input and output directions that encode new reasoning behaviors.

Read analysis
26Optimizer

When Does Muon Help Agentic Reinforcement Learning?

Kai Ruan, Jinghao Lin, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun

Muon may substantially improve agentic RL training, but stronger evidence across seeds, tasks, and optimization settings is still needed.

Read analysis
28Optimizer

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers

Siyuan Li, Jiabao Pan, Yumou Liu, Zhuoli Ouyang, Xin Jin, Xinglong Xu, Jingxuan Wei, Shengye Pang, Jintao Che, Xuanhe Zhou, Conghui He, Cheng Tan

OmniOpt organizes the chaotic world of modern optimizers into a unified taxonomy and benchmark so researchers can compare methods more systematically.

Read analysis
30Optimizer

Why Muon Outperforms Adam: A Curvature Perspective

Shuche Wang, Fengzhuo Zhang, Jiaxiang Li, Dirk Bergemann, Zhuoran Yang

This paper explains why the Muon optimizer can beat Adam in large language model training by showing it makes smaller curvature-induced mistakes, especially in imbalanced and high-curvature settings.

Read analysis
34Optimizer

Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers

Nikhil Nayak, Julia White, Urchade Zaratiana, Kelton Zhang, Henrijs Princis, Dhruv Atreja, Henry Fawcett, Matthew Thomas, George Hurn-Maloney, Ash Lewis

This paper shows that popular optimizer updates for language models are subtly biased on small batches, and it introduces a practical correction that can make AdamW, Sophia, and Shampoo train a bit better.

Read analysis
35Optimizer

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics

Igor Ignashin, Anna Radovskaya, Andrew Semenov, Egor Lopatin, Stanislav Potapov, Aleksandr Kovalenko, Andrey Veprikov, Aleksandr Shestakov, Andrey Leonidov, Aleksandr Beznosikov

This paper argues that SGD should be viewed less like Brownian motion and more like motion in a randomly changing landscape, revealing when training trajectories diffuse versus stay trapped.

Read analysis