Post-Training Language Models for Gold-Medal Performance in Coding Competitions
AuthorsAleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, Boris Ginsburg
Resources
A specialized language model system combines curated coding data, reinforcement learning, and iterative solution refinement to surpass human gold-medal performance in programming competitions.
Key results
Executable competitive-programming problems used to build the specialization corpus.
Synthetic reasoning traces generated for Nemotron-3-Nano-CC.
Score after five GenCorrect rounds, above the 438.3 gold threshold.
Ultra-CC score out of 600 under matched contestant constraints.
Highest official human score surpassed by the live Ultra-CC run.
What the paper found
This paper presents an end-to-end specialization pipeline for competitive programming, combining curation of 22,000 executable problems, synthetic reasoning traces from DeepSeek-V4-Flash, supervised fine-tuning, executable-reward reinforcement learning with GRPO, and GenCorrect, a feedback-driven test-time strategy that generates diverse solutions, evaluates selected submissions, and iteratively refines them. NVIDIA’s Nemotron-3-Nano-CC uses 1.2M training traces plus reinforcement learning, while the larger Nemotron-3-Ultra-CC relies on supervised fine-tuning alone. On IOI 2025, Nano-CC improved from 130 points at Score@1 to 291 after post-training and 468 after five GenCorrect rounds, surpassing the 438.3 gold threshold despite having only 3B active parameters. Ultra-CC reached 502 after the same iterative inference. For a prospective IOI 2026 run, the system combined GLM-5.2-generated training data, expanded final-round sampling, execution-based candidate selection, and NVFP4 quantization; under the same time, internet, and submission constraints as contestants, it scored 535.4 out of 600, exceeding the 361.12 gold threshold and the 498.27 top-human score. The result is a system-level demonstration that post-training and test-time compute can transform general-purpose language models into gold-medal coding systems, although it required substantial computational resources and was not an official contestant.
Original abstract
Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pipeline combining large-scale problem curation, synthetic reasoning traces, supervised fine-tuning (SFT), and reinforcement learning (RL). Using 22,000 curated problems, we train Nemotron-3-Nano-CC (30B-A3B) with SFT and RL and Nemotron-3-Ultra-CC (550B-A55B) with SFT alone. We further introduce GenCorrect, a feedback-driven test-time compute strategy that iteratively generates, evaluates, and refines diverse solutions. On IOI 2025, Nano-CC improves from 130 points to 291 after post-training and to 468 with GenCorrect, exceeding the gold threshold of 438.3 while Ultra-CC reaches 502. Guided by these results, we develop a competition-specific Ultra-CC system and evaluate it prospectively during IOI 2026. Under the same time, internet-access, and submission constraints as human contestants, it scores 535.4 out of 600, exceeding both the gold threshold of 361.12 and the top human score of 498.27. To our knowledge, this is the first AI system to outscore the highest-scoring human contestant on an IOI problem set.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.