NTH

The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators

AuthorsAlex Iacob, Andrej Jovanović, William F. Shen, Daniel Burkhardt, Meghdad Kurmanji, Nurbek Tastan, Lorenzo Sani, Niccolò Alberto Elia Venanzi, Ambroise Odonnat, Zeyu Cao, Bill Marino, Xinchi Qiu, Nicholas D. Lane

July 8, 2026 3 min read
Watch on YouTube
The one-line take

This paper introduces a self-improving agent framework where the evaluator changes too, letting writers, coders, and reviewers co-evolve in an ongoing Red Queen race.

Key results

71.7%
Polyglot pass rate

RQGM coder held-out pass rate versus 69.9% prior SOTA

1.72×
Coding token savings

RQGM uses fewer blended search tokens than the HGM-H baseline

40.5%
Paper acceptance rate

RQGM specialist writer acceptance under the four-reviewer panel

21.8%
Paper acceptance baseline

HGM-H writer acceptance under the same panel

9%
Proof grading accuracy gain

RQGM grader improves ground-truth accuracy over static baselines

13.0×
Nemotron cost reduction

Search-time task-agent routing with NVIDIA Nemotron 3 Ultra lowers blended search cost

What the paper found

The Red Queen Gödel Machine (RQGM) extends recursive self-improvement to non-stationary utilities by co-evolving task agents and learned evaluators inside epoch-bounded search, so each epoch remains a fixed-criterion problem while evaluators can be replaced at checkpoints using ϵ-best-belief on a held-out ground-truth anchor. Built on Huxley-Gödel Machine and HyperAgents, the method uses selective erasure to delete utility records tied to displaced evaluators and exponential checkpoints to keep transition overhead linear in the evaluation budget. In GPT-5.5 (low) experiments, RQGM improves Polyglot coding from 69.9% to 71.7% held-out pass rate using 1.35×–1.72× fewer tokens, raises paper-review acceptance for co-evolved writers from 21.8% to 40.5% under a four-reviewer panel, and improves proof grading accuracy by 9% over static baselines while using 3× lower search cost at the best grader point. The framework also corrects evaluator bias: a paper reviewer trained with an adversarial objective reaches roughly 80% anchor accuracy while making acceptance rates for AI and human papers similar, instead of over-accepting AI-generated work at up to 1.91× the human rate. A cost ablation swaps search-time task-agent calls to NVIDIA Nemotron 3 Ultra and cuts search cost by about 13.0× while approaching the GPT-5.5-only endpoint performance, showing that most expense comes from repeated evaluation rather than expansion.

Original abstract

Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However, their search methods generally assume a stationary evaluation criterion: a fixed verifier, benchmark, or labeled dataset that remains valid as the agent improves. This ignores a central feature of evolution: species adapt as their environments change with them. We aim to bring the same principle to recursive self-improvement, making evaluation part of the improvement loop and opening search to evolving evaluators, adversarial objectives, and dynamic utilities that may surpass static benchmarks. We introduce the Red Queen Godel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities. The RQGM makes this possible through controlled utility evolution: search is organized into epochs with a fixed within-epoch evaluation criterion, while the utility can be updated at epoch boundaries, so self-improvement guarantees hold per epoch as the objective evolves across them. We begin by showing that even on verifiable coding tasks, the RQGM improves test pass rate over the prior SOTA by adding a complementary agent-as-a-judge code-review signal. This signal is cheaper and the RQGM uses 1.35x-1.72x fewer tokens. We then turn to scientific paper writing and reviewing, and Olympiad-level proof writing and grading, where the RQGM improves performance over prior self-improving agents: co-evolved writers reach 1.78x-1.86x higher acceptance rates under a diverse agent-as-a-judge panel, while co-evolved graders reach 9% higher ground-truth accuracy. In paper reviewing, the strongest baseline reviewer over-accepts AI-generated papers at up to 1.91x the human rate. The RQGM corrects this by introducing an adversarial objective that discovers reviewers equally stringent on AI and human work.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis