NTH

LLMs Can Design Near-Optimal OR Algorithms

AuthorsJackie Baek

August 28, 2026 2 min read
Watch on YouTube
The one-line take

The study finds that frontier LLMs can independently produce near-competitive algorithms for important operations-research problems.

Key results

23
Lost-sales level-1 wins

Number of 26 lost-sales instances beaten by gpt-5.6-sol at level 1.

21
Lost-sales level-2 wins

Number of 26 lost-sales instances beaten by gpt-5.6-sol at level 2.

0.5%
Lost-sales maximum gap

gpt-5.6-sol remains within this gap on all 26 lost-sales instances.

628
MMNL benchmark

Number of MMNL assortment instances solved at the best known revenue.

1794
Constrained-MMNL benchmark

Number of constrained-MMNL instances matched at the best known revenue.

87.4%
Nested-logit near-match rate

Share of 971 nested-logit instances within 0.1% of the best comparator.

What the paper found

This paper tests whether frontier LLMs can design operations-research algorithms from a single untuned prompt, precise mathematical specifications, and a Python sandbox. It compares OpenAI’s gpt-5.1, gpt-5.4, and gpt-5.6-sol with Anthropic’s claude-fable-5 across inventory control, queueing-network control, and assortment optimization, using both level 1, instance-specific solutions, and level 2, reusable algorithms written before evaluation instances are revealed. On 26 lost-sales inventory instances, gpt-5.6-sol beats the best tuned benchmark on 23 at level 1 and 21 at level 2, while staying within 0.5% on all instances. Its level-2 inventory policy extends capped base-stock control with a projected inventory statistic that combines expected usable inventory and variance. In queueing, it recovers dynamic-programming optima on small criss-cross networks and beats per-instance PPO on five of six extended reentrant-line instances through pressure-based scheduling. For assortment, it reaches the best known revenue on all 628 MMNL instances and all 1794 constrained-MMNL instances, although nested logit remains difficult: it is within 0.1% of the comparator on 87.4% of 971 instances and loses on a hard tail. Overall, gpt-5.6-sol is no worse than the best existing method in eight of ten problem classes at both levels and matches it on every instance in six classes. The generated artifacts are inspectable combinations of dynamic programming, simulation-tuned policies, relaxations, greedy search, and local improvement rather than black-box neural controllers, but the evidence remains empirical and limited to well-specified benchmark families.

Original abstract

We ask whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems. We study inventory control, queueing network control, and assortment optimization. We evaluate two levels of LLM use: at level 1, the model receives one problem instance and returns a solution for that instance; at level 2, it receives only the problem class description and broad parameter ranges, and returns an algorithm that maps instance parameters to solutions. Human input is minimal: we give one untuned prompt that describes the problem, and the model has access to a Python sandbox tool with a fixed compute budget. The strongest model we test, gpt-5.6-sol, matches or outperforms the best existing method on almost all evaluated instances. This holds even at level 2, where the returned algorithm is fixed before seeing the evaluation instances. Performance also improves sharply across models released less than eight months apart, suggesting that this capability is moving quickly. Thus, for the well-specified operations problems we study, a single untuned LLM query can already produce algorithms competitive with specialized methods. These results suggest that frontier LLMs can be a serious empirical baseline for algorithm design in well-specified OR problems.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis