Nvidia Says Specialized AI Reached Gold-Level Scores in Math and Coding Olympiads
The results came from specialist models paired with repeated checking and revision—not fine-tuning alone. Official graders assessed the math proofs; the coding run was outside the official ranking.
Nvidia reports that specialized Nemotron systems cleared gold thresholds in 2026 math and programming olympiads through task-specific training and multi-stage search, rather than one-shot performance from a general model. IMO graders officially assessed the math proofs; the IOI result came from a live but unsupervised run and was excluded from competition rankings. Nvidia has released training and evaluation materials, offering researchers a basis to examine how specialist models, verification, and repeated refinement contribute to difficult reasoning results.
01
The IMO system scored 30/42, above the official gold threshold of 29; official graders assessed its submitted proofs.
02
The IOI system scored 535.4/600, above the 361.12 gold threshold and human high score of 498.27, but its unsupervised run was excluded from rankings.
03
The math dataset included 414,890 examples across 15,818 proof problems; a separate reinforcement-learning specialist trained on 9,597 problems.
Nvidia says its Nemotron AI family reached gold-level scores in two demanding student competitions—but the achievement belongs to specialized systems, not a model answering once. In an October 7 research report, the company explains how domain training and repeated checking and revision produced results above the gold thresholds in mathematics and competitive programming.
The International Mathematical Olympiad, or IMO, requires rigorous written proofs. The International Olympiad in Informatics, or IOI, demands algorithms and code that pass hidden tests under time and submission limits. Nvidia started both projects with Nemotron 3, then adapted the models and the process used to produce final answers.
Gold-level scores, different evaluation conditions
Nvidia’s mathematics system earned full credit on four of six problems. Its programming system exceeded both the gold cutoff and the top human score of 498.27. Those results have different standing: official IMO graders assessed the submitted proofs, while the IOI run was an unofficial, unsupervised benchmark excluded from the competition ranking.
The IOI test was live and prospective, meaning the system tackled the competition as it happened. Nvidia says it followed the same time, internet-access and submission constraints as human contestants.
Training the answer—and the critic
For programming, Nvidia curated 22,000 problems and generated synthetic reasoning examples. It trained a smaller Nano specialist with supervised fine-tuning—learning from examples—and reinforcement learning. The larger Ultra specialist received supervised fine-tuning only. GenCorrect, its generate-evaluate-refine process, then improved candidate solutions through repeated feedback rounds.
The mathematics training covered more than finished proofs. Its supervised dataset contained 414,890 filtered examples across 15,818 proof problems, including proof generation, refinement, verification and checking the verification itself. A separate reinforcement-learning specialist trained on 9,597 problems selected near the model’s capability frontier.
As the September 9 IMO paper explains, the final system combined those two specialists with the general Nemotron 3 Ultra model. They generated, verified and refined candidate proofs; a separate stage using substantial computation selected each submission. The system worked entirely in natural language, without a formal proof-checking program, external tools or internet access.
Nvidia says the mathematics checkpoints had complementary strengths: the supervised specialist led the first search round, while the reinforcement-learning specialist delivered the best overall single-checkpoint result. Combining them was more valuable than simply drawing more samples from one checkpoint.
The materials behind the scores
Mathematics: The researchers say they released both specialist checkpoints, training data, training and inference code, submitted solutions, and a 200-problem benchmark.
Programming: Nvidia says Nemotron-3-Ultra-CC is available on Hugging Face. The IOI paper describes its training recipe and GenCorrect, while NeMo-Skills provides the evaluation and inference pipeline.
Sources
arxiv.orgAn Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
huggingface.coOne Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
Reader comments
Newest comments first. Replies stay oldest first.