NTH

An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

AuthorsIvan Moshkov, Stephen Ge, George Armstrong, Wei Du, Sadegh Mahdavi, Igor Gitman

September 15, 2026 2 min read
Watch on YouTube
The one-line take

An open Nemotron-based system uses training and test-time search to generate, check, and refine mathematical proofs, reaching the IMO gold-medal threshold.

Key results

30
IMO score

Official score out of 42, exceeding the 29-point gold threshold

550B
Nemotron model size

Nemotron 3 Ultra checkpoint size used by the three pipeline workers

414890
SFT corpus

Quality-filtered training examples for proof generation, refinement, and verification

188
Development ensemble score

End-to-end independent-jury score on the 30-problem development set

1.1%
False-accept rate

Rate for the unanimous 16-judgment RL-plus-SFT verification panel

200
Nemotron-IMO-Bench

Novel olympiad-level problems in the released benchmark

What the paper found

NVIDIA presents an open, natural-language recipe for olympiad-level proof generation built on three Nemotron 3 Ultra 550B-A55B checkpoints: the general-availability model, a supervised-fine-tuned specialist, and a reinforcement-learned specialist. The system performs iterative generate-verify-refine search, using complementary prompts and independent verifier panels, then applies a separate high-compute selection stage; it uses no formal prover, external tools, or internet access. Supervised fine-tuning combines proof construction, refinement, verification, and meta-verification traces into a corpus of 414890 quality-filtered examples, while reinforcement learning trains on difficult problems with asynchronous sampling. On a 30-problem development set, the full ensemble reaches a 188-point end-to-end jury score, outperforming any single checkpoint; the key benefit comes from checkpoint diversity rather than simply doubling samples from one model. For verification, unanimous agreement between the RL and SFT panels reduces false accepts to 1.1%, although it increases false rejects, favoring safe continued search. At IMO 2026, the pipeline scored 30 out of 42 points, above the 29-point gold threshold, earning full credit on Problems 1, 2, 4, and 5. The release includes the post-trained checkpoints, code, training data, submitted proofs, and Nemotron-IMO-Bench, a benchmark of 200 novel olympiad-level problems. The work contrasts with formal systems such as AlphaGeometry and DeepSeekProver, while its model-based judging is evaluated alongside GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.8.

Original abstract

We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold. We release the two post-trained checkpoints as well as the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis