LLMs Can Design Near-Optimal OR Algorithms
AuthorsJackie Baek
Resources
The study finds that frontier LLMs can independently produce near-competitive algorithms for important operations-research problems.
Key results
Number of 26 lost-sales instances beaten by gpt-5.6-sol at level 1.
Number of 26 lost-sales instances beaten by gpt-5.6-sol at level 2.
gpt-5.6-sol remains within this gap on all 26 lost-sales instances.
Number of MMNL assortment instances solved at the best known revenue.
Number of constrained-MMNL instances matched at the best known revenue.
Share of 971 nested-logit instances within 0.1% of the best comparator.
What the paper found
This paper tests whether frontier LLMs can design operations-research algorithms from a single untuned prompt, precise mathematical specifications, and a Python sandbox. It compares OpenAI’s gpt-5.1, gpt-5.4, and gpt-5.6-sol with Anthropic’s claude-fable-5 across inventory control, queueing-network control, and assortment optimization, using both level 1, instance-specific solutions, and level 2, reusable algorithms written before evaluation instances are revealed. On 26 lost-sales inventory instances, gpt-5.6-sol beats the best tuned benchmark on 23 at level 1 and 21 at level 2, while staying within 0.5% on all instances. Its level-2 inventory policy extends capped base-stock control with a projected inventory statistic that combines expected usable inventory and variance. In queueing, it recovers dynamic-programming optima on small criss-cross networks and beats per-instance PPO on five of six extended reentrant-line instances through pressure-based scheduling. For assortment, it reaches the best known revenue on all 628 MMNL instances and all 1794 constrained-MMNL instances, although nested logit remains difficult: it is within 0.1% of the comparator on 87.4% of 971 instances and loses on a hard tail. Overall, gpt-5.6-sol is no worse than the best existing method in eight of ten problem classes at both levels and matches it on every instance in six classes. The generated artifacts are inspectable combinations of dynamic programming, simulation-tuned policies, relaxations, greedy search, and local improvement rather than black-box neural controllers, but the evidence remains empirical and limited to well-specified benchmark families.
Original abstract
We ask whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems. We study inventory control, queueing network control, and assortment optimization. We evaluate two levels of LLM use: at level 1, the model receives one problem instance and returns a solution for that instance; at level 2, it receives only the problem class description and broad parameter ranges, and returns an algorithm that maps instance parameters to solutions. Human input is minimal: we give one untuned prompt that describes the problem, and the model has access to a Python sandbox tool with a fixed compute budget. The strongest model we test, gpt-5.6-sol, matches or outperforms the best existing method on almost all evaluated instances. This holds even at level 2, where the returned algorithm is fixed before seeing the evaluation instances. Performance also improves sharply across models released less than eight months apart, suggesting that this capability is moving quickly. Thus, for the well-specified operations problems we study, a single untuned LLM query can already produce algorithms competitive with specialized methods. These results suggest that frontier LLMs can be a serious empirical baseline for algorithm design in well-specified OR problems.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.