Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
AuthorsYoung-Jun Lee, Seungone Kim, Minki Kang, Alistair Cheong Liang Chuen, Zerui Chen, Seungho Han, Taehee Jung, Dongyeop Kang
Resources
This paper teaches language models how to 'evolve' solutions across hundreds of optimization problems, turning search experience into reusable skill.
Key results
Finch Collection after filtering
Tasks covered by Finch Collection
Distinct domains in Finch Collection
Average improvement over base models across 22 held-out tasks
Average held-out improvement when training tasks increase from 15 to 355
Finch-9B average score on competitive programming tasks
What the paper found
Evolution Fine-Tuning (EFT), developed by researchers at the University of Minnesota, Carnegie Mellon University, KAIST, the University of Cambridge, Hanyang University, and Amazon, reframes evolutionary search as training data so an LLM can internalize the act of discovery instead of relying on a separate scaffold at test time. The paper builds Finch Collection, a 156,731-trajectory corpus spanning 371 optimization tasks across 10 domains, using OpenEvolve and Qwen3.5-397B-A17B as the teacher mutation operator, then fine-tunes open-source Qwen3.5/Qwen3 models from 2B to 9B parameters. Across 22 held-out tasks, EFT improves the base models by 10.22% on average, with especially large gains on ahc058 at 290.59% and Transaction at 74.30%, and it scales from 15 to 355 training tasks with a 14.1% average held-out gain. On competitive programming, Finch-9B reaches 46.01 average score, versus 32.46 for the Qwen3.5-9B baseline. The method also pairs well with preference learning: KTO on Imp and Reg trajectories lifts Finch-8B to 0.381596 on Erdős and 0.9146 on AC2, and test-time RL with Finch matches state-of-the-art on two circle-packing tasks. Overall, the key novelty is turning long-horizon optimization traces into mid-training supervision so small open models can transfer mutation strategies across mathematics, algorithm engineering, GPU kernels, and scientific discovery tasks.
Original abstract
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimization tasks, including open mathematical conjectures, GPU kernel design, scientific law discovery, and combinatorial puzzles. To achieve this, prior work applied search scaffolds to one target task at a time, so every new problem is approached from scratch and the experience accumulated during search is discarded once the model finishes its attempt. This leaves the capability of iteratively evolving a solution (e.g., knowing which part to mutate and how, deciding when to backtrack) entirely in the scaffold rather than in the model itself. Whether the model itself could acquire this capability and reuse it across different tasks has been largely unexamined. To address this, we introduce Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. We construct Finch Collection, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, and fine-tune open-source LLMs from 2B to 9B parameters. Empirically, EFT confers cross-task generalization: across 22 held-out tasks, our models surpass their base counterparts by 10.22% on average. Furthermore, when paired with test-time RL, our model matches state-of-the-art performance on two circle-packing tasks and outperforms its base-model counterpart on the Erdős minimum-overlap problem. EFT thus serves as a "practice phase" for general-purpose discovery agents that do not solve new problems from scratch.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.