AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can One Nemotron Model Family Reach Gold-Level Results In IOI And IMO? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face says two specialized systems built from its Nemotron 3 family scored 30 of 42 at the 2026 IMO and 535.4 of 600 in an IOI benchmark run. IMO graders officially awarded the proofs 30 points; the IOI result was unofficial and excluded from the competition ranking.

Hugging Face says two specialized systems built from its Nemotron 3 model family reached the stated gold-medal thresholds in evaluations tied to the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO), as detailed in the original analysis. The IMO proofs received 30 of 42 points from official graders; the IOI system scored 535.4 of 600 in an unofficial run that did not count toward the contest’s official ranking.

For the IOI evaluation, Hugging Face says it used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning and combined with GenCorrect. That process generates candidate code, evaluates it and iteratively refines promising solutions. The company reports that the run took place prospectively under the same time, internet-access and submission constraints as human contestants. Its score was above the reported 361.12-point gold threshold and the top human score of 498.27. Those comparisons describe the company’s unofficial run, not an official placement or medal.

For the IMO, Hugging Face combined the general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints in a system for generating and revising written proofs. According to the company, official graders awarded the submitted proofs 30 points, exceeding the stated 29-point gold threshold and earning full credit on four of six problems. Hugging Face says the system used no formal prover, external tools or internet access.

The two projects used separate specialist training approaches and datasets. The IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO supervised fine-tuning data included 414,890 quality-filtered examples from 15,818 proof problems; its reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. These figures and results are reported by Hugging Face, which developed the systems.

At a glance
reportWhen: Reported in 2026; the IOI score was uno…
The developmentHugging Face reported that specialist systems based on Nemotron 3 reached the stated gold thresholds in mathematics and programming competition evaluations in 2026.
At a glance
reportWhen: Reported after the 2026 competitions
The developmentHugging Face reported that systems fine-tuned from Nemotron 3 scored above the gold thresholds at IOI 2026 and IMO 2026.

Two Contests, Different Evidence

The results show how a shared model family can be adapted for distinct tasks: producing code that performs against hidden tests and writing proofs that human graders can assess. Hugging Face’s approach pairs domain-specific fine-tuning with inference-time generation, checking and revision rather than relying on a model’s first answer. That is relevant to researchers testing whether targeted adaptation and additional reasoning steps can improve performance without building a separate foundation model for every subject.

The evidence is not equally strong across the two contests. The IMO score was awarded by official graders, while the IOI score came from a company-reported run outside the official ranking. Neither result by itself establishes broad ability across mathematics, programming or practical workloads. Competition scores measure performance on a particular set of tasks under particular conditions; they do not settle how the systems would handle unfamiliar problems or applications beyond those settings.

Amazon

AI programming problem solver

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From IOI Experiments to IMO Proofs

Hugging Face presents the 2026 work as an extension of its earlier IOI experiments. For IOI 2025, the company reported that a Nemotron-3-Nano-CC model increased from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. Adding GenCorrect brought the reported score to 468, above that year’s stated gold threshold of 438.3. An Ultra-CC version scored 502 using the same test-time strategy, according to the company.

The 2026 evaluations applied related ideas to two different disciplines. Programming performance depended on producing executable code that passed hidden tests, while the IMO required rigorous written proofs. Hugging Face says the IMO project found complementary strengths in its supervised fine-tuning and reinforcement-learning checkpoints, so the final system drew on both alongside the general model. The reported scores therefore reflect specialized systems and procedures, not an unmodified Nemotron model applied directly to both competitions.

“Success at both points to something broader.”

— Hugging Face

Amazon

AI proof generation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Far the Scores Can Be Verified

The IMO result has official grading, but the supplied account does not provide the graders’ detailed feedback or a full independent replication. The IOI result has a more direct limitation: it was an unofficial benchmark run and was not entered into the official ranking. Hugging Face describes it as unsupervised and run under contest-like constraints, but the available material does not explain how the run was audited or whether outside evaluators have reproduced it.

It is also unclear how reliably the systems would perform on other competitions, different proof styles, unseen programming problems or real-world tasks. Scores can depend on training data selection, compute budgets and evaluation details. The report does not provide enough information to compare those factors across the two systems or establish that their results generalize beyond the specific evaluations.

Amazon

AI code evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Checkpoints and Independent Tests

Hugging Face says it is releasing an IMO 2026 collection containing supervised fine-tuning and reinforcement-learning checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to a paper describing the IMO training and generate-verify-refine process, along with a NeMo-Skills repository. These resources may let other researchers inspect the methods and test the systems on additional problems.

The supplied account gives no timetable for further releases and does not describe an independent evaluation of the IOI system. Useful next evidence would include outside replication of the IMO benchmark, public scrutiny of the IOI run’s setup and evaluation, and tests on problems beyond the competitions. Until then, the IMO result is an officially graded competition score, while the IOI figure should be described as a company-reported unofficial benchmark result.

Amazon

AI competition preparation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Did a Nemotron system officially win an IOI medal?

No. Hugging Face reported a score of 535.4 out of 600 in an unofficial IOI run, but the result was not included in the competition’s official ranking.

Was the IMO result officially graded?

Yes. Hugging Face says official IMO graders awarded the submitted proofs 30 of 42 points, above the reported 29-point gold threshold. The system earned full credit on four of six problems, according to the company.

Were the same Nemotron system and training data used for both contests?

No. Hugging Face describes separate specialist systems and datasets: the IOI work used a competition-specific Nemotron-3-Ultra-CC version and GenCorrect, while the IMO system combined Nemotron 3 Ultra with supervised fine-tuning and reinforcement-learning checkpoints.

Can the reported scores establish broad reasoning ability?

Not on their own. They measure performance in specific competition evaluations, and the IOI result was unofficial. The company’s report does not establish how well either system generalizes to other competitions or practical tasks.

What materials does Hugging Face say it is releasing?

The company says its IMO collection includes model checkpoints, training datasets and a 200-problem olympiad benchmark, with a paper and NeMo-Skills repository also referenced. The supplied material does not give a release timetable or fully specify every repository detail.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Analyzing The Impact Of Anthropic’s 17% Cut On Claude Code’s Weekly Limits

Anthropic has reduced Claude Code’s weekly usage limits by 17%, affecting capacity but not model performance. Details on affected plans and timing are pending.

Discover How Anthropic’s Claude Contributes To The Next AI Breakthroughs

Anthropic claims its AI model, Claude, is actively contributing to building its successor through code, research, and analysis — a potential leap in AI self-improvement.

Kolibri: A Sovereign Open-Weight Model

Aleph Alpha says Kolibri has 78B total and 3B active parameters, a 1M-token context window and Apache 2.0 weights for download.

Run Kimi K3 Using 29 GB Of RAM At 0.50 Tok/s

Kimi K3 runs with 29 GB RAM at 0.50 tok/s, highlighting its high resource demands. Details are confirmed, but implications remain under discussion.