An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
AuthorsIvan Moshkov, Stephen Ge, George Armstrong, Wei Du, Sadegh Mahdavi, Igor Gitman
Resources
An open Nemotron-based system uses training and test-time search to generate, check, and refine mathematical proofs, reaching the IMO gold-medal threshold.
Key results
Official score out of 42, exceeding the 29-point gold threshold
Nemotron 3 Ultra checkpoint size used by the three pipeline workers
Quality-filtered training examples for proof generation, refinement, and verification
End-to-end independent-jury score on the 30-problem development set
Rate for the unanimous 16-judgment RL-plus-SFT verification panel
Novel olympiad-level problems in the released benchmark
What the paper found
NVIDIA presents an open, natural-language recipe for olympiad-level proof generation built on three Nemotron 3 Ultra 550B-A55B checkpoints: the general-availability model, a supervised-fine-tuned specialist, and a reinforcement-learned specialist. The system performs iterative generate-verify-refine search, using complementary prompts and independent verifier panels, then applies a separate high-compute selection stage; it uses no formal prover, external tools, or internet access. Supervised fine-tuning combines proof construction, refinement, verification, and meta-verification traces into a corpus of 414890 quality-filtered examples, while reinforcement learning trains on difficult problems with asynchronous sampling. On a 30-problem development set, the full ensemble reaches a 188-point end-to-end jury score, outperforming any single checkpoint; the key benefit comes from checkpoint diversity rather than simply doubling samples from one model. For verification, unanimous agreement between the RL and SFT panels reduces false accepts to 1.1%, although it increases false rejects, favoring safe continued search. At IMO 2026, the pipeline scored 30 out of 42 points, above the 29-point gold threshold, earning full credit on Problems 1, 2, 4, and 5. The release includes the post-trained checkpoints, code, training data, submitted proofs, and Nemotron-IMO-Bench, a benchmark of 200 novel olympiad-level problems. The work contrasts with formal systems such as AlphaGeometry and DeepSeekProver, while its model-based judging is evaluated alongside GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.8.
Original abstract
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold. We release the two post-trained checkpoints as well as the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.