跳到正文
原文
Hugging Face Blog·· 3 小时前精选AI 评分65

Nemotron 经微调在 IOI 2026 与 IMO 2026 达到金牌水平

One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO

AI 导读

从 Nemotron 3 出发,研究团队用监督微调、强化学习与反馈驱动推理做出专项系统,在 IOI 2026 和 IMO 2026 都达到金牌水平。

推荐理由

从 Nemotron 基座出发,用监督微调、强化学习以及生成评估修正循环做专项化,使信息学与数学奥赛系统都越过金牌分数线。

正文 · 原文

The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) test different skills. IOI requires algorithms and code that pass hidden tests under strict time and submission limits. IMO demands rigorous natural-language proofs. Success at either competition is difficult. Success at both points to something broader.

Our recent results show that Nemotron is a strong, adaptable foundation for building world-class specialist models. Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.

fig_ioi_imo_gold_results_headline

Competition Nemotron specialization Result
IOI 2026 Nemotron-3-Ultra-CC with SFT and GenCorrect 535.4/600, above the 361.12 gold threshold and the top human score of 498.27
IMO 2026 Nemotron 3 Ultra general, SFT, and RL checkpoints in a generate-verify-refine system 30/42, above the official gold threshold of 29

The IOI result came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants. It was an unofficial, unsupervised benchmark and was not included in the official IOI ranking. The IMO system’s submitted proofs were graded by official IMO graders.

A reusable specialization recipe

"Easy to fine-tune" should mean more than making a checkpoint trainable. It should mean that a capable foundation model can be adapted to a demanding domain with a clear, reusable recipe.

Across the two projects, that recipe had four parts:

  1. Start with a strong Nemotron base model.
  2. Curate domain-specific problems and high-quality reasoning traces.
  3. Apply standard post-training methods such as SFT and, where useful, RL.
  4. Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers.

The training and inference runs were substantial, but the underlying approach is familiar and reproducible. We did not need to build a new foundation model for every challenge. We specialized Nemotron for the task.

From general coding ability to IOI gold

For competitive programming, we curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC, with 30 billion total parameters and 3 billion active parameters, received both SFT and RL. Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, received SFT.

The progression on IOI 2025 makes the value of specialization easy to see. Nano improved from 130 points before post-training to 280 after SFT and 291 after RL. With GenCorrect, our iterative generate-evaluate-refine strategy, it reached 468 points and crossed the gold threshold of 438.3. Ultra-CC reached 502 points with the same test-time strategy.

fig_main_capability_progression_previous_style

These experiments also showed that adaptation does not have to look the same at every scale. SFT produced most of Nano's gain, with RL adding a smaller but consistent improvement. For the stronger Ultra model, one SFT epoch was enough to outperform the fully post-trained Nano model across IOI, ICPC, and LiveCodeBench Pro. That finding guided the competition-specific Ultra-CC system used for IOI 2026, which scored 535.4 out of 600.

Teaching Nemotron to prove, check, and revise

The IMO project applied the same idea to olympiad mathematics. Starting from Nemotron 3 Ultra, we trained one specialist with SFT and another with RL.

The SFT corpus contained 414,890 quality-filtered examples across 15,818 unique proof problems. It did more than teach final answers. The data covered proof generation, refinement, verification, and meta-verification, so the model learned to construct arguments, identify gaps, respond to critiques, and judge whether a proof was complete. The RL model was trained on 9,597 proof problems selected near the model's capability frontier.

Both post-trained checkpoints outperformed the general-availability model in the development experiments. The SFT checkpoint was strongest in the first search round, while the RL checkpoint achieved the best overall single-checkpoint result. Their strengths were complementary, so the final system used both specialists alongside the general model.

For each IMO problem, the models generated candidate proofs, scored them, produced critiques, and refined the most promising attempts. A separate high-compute stage selected the final submission. The entire system worked in natural language, with no formal prover, external tools, or internet access. It scored 30 out of 42 points, including full credit on four of the six problems, and exceeded the official gold-medal threshold.

image

Fine-tuning and test-time compute work together

Our earlier IOI 2025 Hugging Face post showed how test-time compute can push open-weight models to gold-level performance. The new results add an important piece: better specialization gives the inference system better candidates, better critics, and better refinements.

At IOI, GenCorrect turned the gains from fine-tuning into larger improvements over multiple feedback rounds. At IMO, using complementary SFT and RL checkpoints was more valuable than simply drawing more samples from one checkpoint. In both cases, the best outcome came from combining a capable specialist with a system that could search, verify, and improve.

This distinction matters. The medals were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone. They came from co-designing the model, the data, and the inference loop.

Open models, data, and recipes on Hugging Face

We want these results to be useful beyond the competitions. The Nemotron Labs IMO 2026 collection brings together the SFT and RL checkpoints, both training datasets, and Nemotron-IMO-Bench, a new benchmark of 200 olympiad-level problems. The IMO paper describes the training approach and generate-verify-refine system, while the NeMo-Skills repository includes the IMO inference pipeline, prompts, submitted proofs, and a reproducible quickstart. For competitive programming, the Nemotron-3-Ultra-CC model is available on Hugging Face, and the IOI paper provides the training recipe and the GenCorrect methodology. The IOI evaluation and inference pipeline are also available in NeMo-Skills.

Together, IMO and IOI provide unusually demanding evidence for a simple idea: Nemotron can be fine-tuned into world-class domain specialists, then composed with transparent inference workflows to solve problems at the frontier of human competition.

We are excited to see what the Hugging Face community builds next.

来源:Hugging Face Blog · huggingface.co