The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
AuthorsAlex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccolò Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane
Resources
This paper introduces a self-improving agent framework where the evaluator changes too, letting writers, coders, and reviewers co-evolve in an ongoing Red Queen race.
Key results
RQGM coder held-out pass rate versus 69.9% prior SOTA
RQGM uses fewer blended search tokens than the HGM-H baseline
RQGM specialist writer acceptance under the four-reviewer panel
HGM-H writer acceptance under the same panel
RQGM grader improves ground-truth accuracy over static baselines
Search-time task-agent routing with NVIDIA Nemotron 3 Ultra lowers blended search cost
What the paper found
The Red Queen Gödel Machine (RQGM) extends recursive self-improvement to non-stationary utilities by co-evolving task agents and learned evaluators inside epoch-bounded search, so each epoch remains a fixed-criterion problem while evaluators can be replaced at checkpoints using ϵ-best-belief on a held-out ground-truth anchor. Built on Huxley-Gödel Machine and HyperAgents, the method uses selective erasure to delete utility records tied to displaced evaluators and exponential checkpoints to keep transition overhead linear in the evaluation budget. In GPT-5.5 (low) experiments, RQGM improves Polyglot coding from 69.9% to 71.7% held-out pass rate using 1.35×–1.72× fewer tokens, raises paper-review acceptance for co-evolved writers from 21.8% to 40.5% under a four-reviewer panel, and improves proof grading accuracy by 9% over static baselines while using 3× lower search cost at the best grader point. The framework also corrects evaluator bias: a paper reviewer trained with an adversarial objective reaches roughly 80% anchor accuracy while making acceptance rates for AI and human papers similar, instead of over-accepting AI-generated work at up to 1.91× the human rate. A cost ablation swaps search-time task-agent calls to NVIDIA Nemotron 3 Ultra and cuts search cost by about 13.0× while approaching the GPT-5.5-only endpoint performance, showing that most expense comes from repeated evaluation rather than expansion.
Original abstract
Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them. We aim to bring the same principle to recursive self-improvement, making evaluation part of the improvement loop and opening search to evolving evaluators, adversarial objectives, and dynamic utilities that may surpass static benchmarks. We introduce the Red Queen Godel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities. The RQGM makes this possible through controlled utility evolution: search is organized into epochs with a fixed within-epoch evaluation criterion, while the utility can be updated at epoch boundaries, so self-improvement guarantees hold per epoch as the objective evolves across them. We begin by showing that even on verifiable coding tasks, the RQGM improves test pass rate over the prior SOTA by adding a complementary agent-as-a-judge code-review signal. This signal is cheaper and the RQGM uses 1.35x-1.72x fewer tokens. We then turn to scientific paper writing and reviewing, and Olympiad-level proof writing and grading, where the RQGM improves performance over prior self-improving agents: co-evolved writers reach 1.78x-1.86x higher acceptance rates under a diverse agent-as-a-judge panel, while co-evolved graders reach 9% higher ground-truth accuracy. In paper reviewing, the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. The RQGM corrects this by introducing an adversarial objective that discovers reviewers equally stringent on AI and human work.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.