Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
AuthorsVladislav Beliaev
Resources
Agon trains two reasoning models to compete against each other, turning rivalry into an implicit judge that substantially improves problem-solving performance.
Key results
Qwen3-0.6B on 300 held-out hard DeepMath-103K problems
Baseline on the same Qwen3-0.6B DeepMath-103K evaluation
Noncompetitive cross-model exchange on the same benchmark
Average final-stage completion tokens, compared with 8.1k for GRPO
Zero-training two-agent inference control on DeepMath-103K
What the paper found
Agon proposes competitive cross-model reinforcement learning to address a weakness in GRPO and DeepSeek-R1-style outcome-based training: only the final answer is verified, so models can increase reward by producing longer, low-density reasoning. Two distinct policies instead solve the same problem in rotating drafter and challenger roles; the challenger reads the rival’s solution summary, with the final answer withheld, and earns a conversion bonus for being correct when the rival fails. This implicitly grades reasoning without process labels, a learned reward model, or changes to the verifier. The implementation uses two divergent LoRA adapters over one frozen base, adding about 2% memory overhead, and deploys as a two-stage cascade. On 300 held-out hard problems from DeepMath-103K, Qwen3-0.6B improved from 30% pass@1 with vanilla GRPO to 46% with cooperative exchange and 61% with Agon, while final-stage traces fell to 3.5k tokens from 8.1k for GRPO. The untrained Mixture-of-Agents baseline reached only 34%. Transfer checks on GSM8K and MATH-500 produced 75% and 64%, respectively, and CodeContests on Qwen3-1.7B preserved the ordering, with Agon reaching 34% pass@1. Scaling experiments also report gains across Qwen3.5 and Google DeepMind’s Gemma 4. The paper’s main limitation is that inference requires two sequential generations, and it does not yet quantify run-to-run variance or test latent-space communication, which is proposed as the next step.
Original abstract
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a two-stage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO's pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4). For now the models talk in text; the next step is to let them reason together in latent space.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.