NTH

LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28

AuthorsWes Sander

September 14, 2026 2 min read
Watch on YouTube
The one-line take

An LLM repeatedly rewrites a circle-packing solver, beats 10 established records, and does it for less than $28.

Key results

10
Packomania records improved

Number of csqv target records improved by Discovery Loop

5.4%
Maximum improvement

Largest gain over a prior Packomania record

27.72
LLM cost

Total cost in US dollars for the full run

15
Iterations completed

Iterations run before reaching the budget cap

13.95
Plateau-stop cost

Retrospective spending with adaptive plateau detection

0.01%
Plateau-stop quality loss

Estimated loss in final value under early stopping

What the paper found

The paper presents Discovery Loop, a roughly 400-line system that uses Anthropic’s Claude Fable 5.1 to evolve complete optimization solvers rather than isolated code patches. Each iteration supplies the current champion, a scoreboard, and the last 12 ideas; the model proposes a replacement Python solver, which is evaluated in parallel and checked by an independent zero-tolerance verifier. Starting from a multi-start penalty L-BFGS-B solver with LP-optimized radii, the system applied basin hopping, hexagonal-lattice initialization, island-model parallelism, KKT-Newton polishing, and defect migration to the Packomania csqv benchmark, which maximizes the sum of variable circle radii in a unit square. It improved 10 of 12 targets for N values from 101 to 114, achieving gains of 2.4% to 5.4% over accepted records, and completed 15 iterations for a total LLM cost of $27.72 on a consumer PC. The approach is a lightweight counterpart to Google DeepMind’s Gemini-powered AlphaEvolve and FunSearch: it uses one model call per iteration, no distributed cluster, and per-target best tracking. Cost-efficiency declined sharply after early gains, but retrospective plateau detection would have stopped at iteration 9, reducing spending to $13.95 while sacrificing only 0.01% of final value. The results suggest that independently verifiable, LLM-guided algorithm discovery can be performed at laptop scale, though performance remains sensitive to problem structure and code-generation reliability.

Original abstract

We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve optimization algorithms. Starting from a simple seed solver, the LLM proposes algorithmic improvements guided by a scoreboard of results and a history of prior ideas. Each candidate is evaluated against an independent verifier; improvements are kept and failures discarded. Applied to the Packomania circle-packing benchmark (csqv: maximize the sum of radii of N variable-radius circles in the unit square), the system improved the best known solutions for 10 values of N in the range 101-114, with gains of 2.4%-5.4% over prior records, all within 15 iterations and at a total LLM cost of $27.72. These results have been independently accepted by Packomania. We describe the method, analyze cost-efficiency dynamics including an adaptive plateau-detection mechanism, and discuss implications for democratizing automated scientific discovery.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis