MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
AuthorsJiacheng Chen, Xinyu Zhang, Shunkai Zhang, Yanmohan Wang, Lin Li, Tiancheng Qin, Qin Wang, Zhengmao Zhu, Tianle Li, Jingyang Li, Zehan Li, Binyang Jiang, Jin Zhu, Han Ding, Fei Yu, Chenyu Du, Zijian Song, Jiayuan Song, Zhi Zhang, Yunan Huang, Weiyu Cheng, Pengyu Zhao, Yu Cheng
Resources
MaxProof uses a population of proof candidates plus generator-verifier-ranker loops to push an AI system to gold-medal-level performance on hard math olympiad proofs.
Key results
Standalone M3 score on proof-style benchmark
Standalone M3 score on proof-style benchmark
M3 one-shot contest score out of 42
M3 + MaxProof contest score out of 42
M3 one-shot contest score out of 42
M3 + MaxProof contest score out of 42
What the paper found
MaxProof is MiniMax’s proof-specialized scaling system built into the MiniMax-M3 release, combining generative-verifier reinforcement learning with population-level test-time search to push competition math beyond one-shot capability. The paper’s core engineering move is a defense-in-depth verifier: it filters malformed outputs, normalizes proofs, scores them with three parallel judges, and then uses pessimistic min aggregation to suppress false positives, because an earlier M2 training run showed classic reward hacking, including a roughly 3× length inflation, over 80% template convergence, and a verifier score that rose to 1.0 while an independent expert judge averaged only 0.55. M3 is trained as three merged proof competencies—proof generation, explicit error finding, and critique-conditioned repair—then MaxProof treats the same model as generator, verifier, refiner, and ranker over a population of 32 candidates, 4 verifier samples per candidate, 10 refinement rounds, and tournament selection. In evaluation, the merged M3 model scores 67.40 on IMOProof Bench and 81.56 on IMOAnswerBench, and MaxProof lifts contest performance from 27 to 35 on IMO 2025 and from 26 to 36 on USAMO 2026, exceeding the human gold-medal threshold on both. Per-problem analysis shows the population often reaches 7/7 by round 4, but one USAMO 2026 problem exposes selection loss when the oracle-best candidate is 6/7 while the self-pick falls to 2/7, highlighting that better rankers remain the main bottleneck.
Original abstract
We present MaxProof, a population-level test-time scaling framework for competition-level mathematical proof in the MiniMax-M3 series. M3 first trains three proof-oriented capabilities -- proof generation, proof verification, and critique-conditioned proof repair -- using a defense-in-depth generative verifier engineered for low false-positive rate. These capabilities are merged into a single released M3 model. At test time, MaxProof treats the model as a generator, verifier, refiner, and ranker, searches over a population of candidate proofs, and returns one final proof through tournament selection. With MaxProof test-time scaling, the M3 model reaches 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the human gold-medal threshold on both.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.