英伟达开源 IMO 金牌配方:不仅是「人海战术」,1.5TB 显存做实 AI「 推恩令」?
Ranking
Overall
82
Content
95
Popularity
53
Observed public metrics from 1 member.
Merged summary
TL;DR - NVIDIA open-sourced much of the Nemotron 3 Ultra system that scored 30/42—above the gold-medal cutoff—at the 2026 IMO using natural-language proofs. The release exposes a powerful multi-checkpoint proof-search recipe, but also highlights correlated verifier errors and a roughly 1.5 TB VRAM barrier to reproduction.
- The system combines general, SFT, and RL checkpoints to diversify proof strategies; SFT emphasizes proof repair, while RL increases the probability of successful reasoning paths.
- It generates 384 initial proofs per problem, then iteratively verifies, retains, and revises promising candidates rather than repeatedly sampling from scratch.
- Acceptance requires unanimous repeated verification, yet checkpoints sharing the same base model can still endorse identical flawed assumptions or reject partially valid proofs.
- NVIDIA released expert checkpoints, training data, inference code, recipes, submitted proofs, and a 200-problem benchmark, but some intermediate checkpoints remain closed and full-scale reproduction requires substantial GB200-class compute.
Sources (1)
英伟达开源 IMO 金牌配方:不仅是「人海战术」,1.5TB 显存做实 AI「 推恩令」?
Public signals
Hugging Face upvotes 43
TL;DR - NVIDIA open-sourced much of the Nemotron 3 Ultra system that scored 30/42—above the gold-medal cutoff—at the 2026 IMO using natural-language proofs. The release exposes a powerful multi-checkpoint proof-search recipe, but also highlights correlated verifier errors and a roughly 1.5 TB VRAM barrier to reproduction.
- The system combines general, SFT, and RL checkpoints to diversify proof strategies; SFT emphasizes proof repair, while RL increases the probability of successful reasoning paths.
- It generates 384 initial proofs per problem, then iteratively verifies, retains, and revises promising candidates rather than repeatedly sampling from scratch.
- Acceptance requires unanimous repeated verification, yet checkpoints sharing the same base model can still endorse identical flawed assumptions or reject partially valid proofs.
- NVIDIA released expert checkpoints, training data, inference code, recipes, submitted proofs, and a 200-problem benchmark, but some intermediate checkpoints remain closed and full-scale reproduction requires substantial GB200-class compute.