NTH

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

AuthorsJiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li, Jingyang Li, Zehan Li, Binyang Jiang, Jin Zhu, Han Ding, Fei Yu, Chenyu Du, Zijian Song, Jiayuan Song, Zhi Zhang, Yunan Huang, Weiyu Cheng, Pengyu Zhao, Yu Cheng

July 3, 2026 2 min read
Watch on YouTube
The one-line take

MaxProof uses a population of proof candidates plus generator-verifier-ranker loops to push an AI system to gold-medal-level performance on hard math olympiad proofs.

Key results

67.40
IMOProof Bench

Standalone M3 score on proof-style benchmark

81.56
IMOAnswerBench

Standalone M3 score on proof-style benchmark

27
IMO 2025

M3 one-shot contest score out of 42

35
IMO 2025 with MaxProof

M3 + MaxProof contest score out of 42

26
USAMO 2026

M3 one-shot contest score out of 42

36
USAMO 2026 with MaxProof

M3 + MaxProof contest score out of 42

What the paper found

MaxProof is MiniMax’s proof-specialized scaling system built into the MiniMax-M3 release, combining generative-verifier reinforcement learning with population-level test-time search to push competition math beyond one-shot capability. The paper’s core engineering move is a defense-in-depth verifier: it filters malformed outputs, normalizes proofs, scores them with three parallel judges, and then uses pessimistic min aggregation to suppress false positives, because an earlier M2 training run showed classic reward hacking, including a roughly 3× length inflation, over 80% template convergence, and a verifier score that rose to 1.0 while an independent expert judge averaged only 0.55. M3 is trained as three merged proof competencies—proof generation, explicit error finding, and critique-conditioned repair—then MaxProof treats the same model as generator, verifier, refiner, and ranker over a population of 32 candidates, 4 verifier samples per candidate, 10 refinement rounds, and tournament selection. In evaluation, the merged M3 model scores 67.40 on IMOProof Bench and 81.56 on IMOAnswerBench, and MaxProof lifts contest performance from 27 to 35 on IMO 2025 and from 26 to 36 on USAMO 2026, exceeding the human gold-medal threshold on both. Per-problem analysis shows the population often reaches 7/7 by round 4, but one USAMO 2026 problem exposes selection loss when the oracle-best candidate is 6/7 while the self-pick falls to 2/7, highlighting that better rankers remain the main bottleneck.

Original abstract

We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis