LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
AuthorsWes Sander
Resources
An LLM repeatedly rewrites a circle-packing solver, beats 10 established records, and does it for less than $28.
Key results
Number of csqv target records improved by Discovery Loop
Largest gain over a prior Packomania record
Total cost in US dollars for the full run
Iterations run before reaching the budget cap
Retrospective spending with adaptive plateau detection
Estimated loss in final value under early stopping
What the paper found
The paper presents Discovery Loop, a roughly 400-line system that uses Anthropic’s Claude Fable 5.1 to evolve complete optimization solvers rather than isolated code patches. Each iteration supplies the current champion, a scoreboard, and the last 12 ideas; the model proposes a replacement Python solver, which is evaluated in parallel and checked by an independent zero-tolerance verifier. Starting from a multi-start penalty L-BFGS-B solver with LP-optimized radii, the system applied basin hopping, hexagonal-lattice initialization, island-model parallelism, KKT-Newton polishing, and defect migration to the Packomania csqv benchmark, which maximizes the sum of variable circle radii in a unit square. It improved 10 of 12 targets for N values from 101 to 114, achieving gains of 2.4% to 5.4% over accepted records, and completed 15 iterations for a total LLM cost of $27.72 on a consumer PC. The approach is a lightweight counterpart to Google DeepMind’s Gemini-powered AlphaEvolve and FunSearch: it uses one model call per iteration, no distributed cluster, and per-target best tracking. Cost-efficiency declined sharply after early gains, but retrospective plateau detection would have stopped at iteration 9, reducing spending to $13.95 while sacrificing only 0.01% of final value. The results suggest that independently verifiable, LLM-guided algorithm discovery can be performed at laptop scale, though performance remains sensitive to problem structure and code-generation reliability.
Original abstract
We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively evolve optimization algorithms. Starting from a simple seed solver, the LLM proposes algorithmic improvements guided by a scoreboard of results and a history of prior ideas. Each candidate is evaluated against an independent verifier; improvements are kept and failures discarded. Applied to the Packomania circle-packing benchmark (csqv: maximize the sum of radii of N variable-radius circles in the unit square), the system improved the best known solutions for 10 values of N in the range 101-114, with gains of 2.4%-5.4% over prior records, all within 15 iterations and at a total LLM cost of $27.72. These results have been independently accepted by Packomania. We describe the method, analyze cost-efficiency dynamics including an adaptive plateau-detection mechanism, and discuss implications for democratizing automated scientific discovery.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.