Self-Improving Language Models with Bidirectional Evolutionary Search
AuthorsGuowei Xu, Zhenting Qi, Huangyuan Su, Weirui Ye, Himabindu Lakkaraju, Sham M. Kakade, Yilun Du
This paper introduces a new way for language models to improve themselves by both evolving candidate answers forward and breaking problems into checkable subgoals backward, boosting performance on hard reasoning tasks.
Key results
Llama-3.2-3B-Instruct base model accuracy on MuSiQue
BES accuracy on MuSiQue with Llama-3.2-3B-Instruct
Llama-3.1-8B-Instruct base model accuracy on MuSiQue
BES accuracy on MuSiQue with Llama-3.1-8B-Instruct
BES average objective on Circle Packing (Square) with GPT-5
BES average objective on Heilbronn (Convex) with GPT-5
What the paper found
Self Improving Language Models with Bidirectional Evolutionary Search, from Harvard University and MIT, introduces BES, a search framework for self-improvement and test-time reasoning that targets two failure modes of best-of-N sampling and tree search: sparse verification and candidate generation trapped inside the model’s own probability mass. BES adds four evolution operators—combination, deletion, translocation, and crossover—to recombine partial trajectories, while a backward goal tree recursively decomposes the task into verifiable sub-goals and supplies dense intermediate scores. The paper proves that expansion-only rollouts remain confined to a narrow entropy shell, whereas evolution can escape it, and shows that backward sub-goal signals reduce the candidate requirement from a multiplicative terminal-hitting problem to a much cheaper local evidence-collection problem. In experiments, BES improves Knights-and-Knaves post-training on Gemma-3-1B-it, where GRPO and MaxRL show little to no progress, and on MuSiQue it raises accuracy from 4.0% to 7.0% on Llama-3.2-3B-Instruct and from 6.6% to 10.4% on Llama-3.1-8B-Instruct, while also increasing finish ratio to 0.97 and 0.94. At inference, using GPT-5 on open optimization problems, BES outperforms open-source frameworks like OpenEvolve, GEPA, and ShinkaEvolve on Circle Packing and Heilbronn, reaching 2.623 on Circle Packing (Square), 2.349 on Circle Packing (Rect.), and 0.026 on Heilbronn (Convex), with lower variance and modest extra API cost.
Original abstract
Search has been proposed as an effective method for self-improving language models and agentic systems, both for post-training sample generation and for inference. However, widely used methods such as best-of-N sampling and tree search face two fundamental limitations: they are guided by sparse verification signals, and they construct candidates primarily through autoregressive expansion, restricting exploration to regions with substantial model probability mass. To address these, we propose Bidirectional Evolutionary Search (BES), a search framework that couples forward candidate evolution with backward goal decomposition. In the forward search, BES augments standard expansion with evolution operators that recombine partial trajectories to generate candidates that are difficult to obtain from a single model rollout. In the backward search, BES recursively decomposes the original task into checkable subgoals, producing dense intermediate feedback that guides forward search. We provide theoretical motivation showing that candidates generated by expansion-only search are confined to a narrow entropy shell while evolutionary operators can escape it, and that backward search can exponentially reduce the number of required samples to find a correct answer. Experiments show that on challenging post-training tasks where mainstream post-training algorithms fail to improve, BES enables consistent gains, and on three open problem solving benchmarks at inference time, BES outperforms existing open-source frameworks in both average and best-case performance. Code and trained models are available at https://github.com/Embodied-Minds-Lab/BES.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.