🛰️ Daily AI Frontier
‹ back to 2026-09-03

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Research LLMs & Foundation Models

Ranking

Overall 85
Content 95
Popularity 62

Observed public metrics from 1 member.

Merged summary

TL;DR - Researchers developed a post-training and test-time refinement pipeline that enabled Nemotron models to achieve gold-medal-level competitive programming performance. Their specialized Ultra-CC system scored 535.4/600 on IOI 2026, surpassing both the gold threshold and the top human score under equivalent competition constraints.

  • The pipeline combines 22,000 curated problems, synthetic reasoning traces, supervised fine-tuning, and reinforcement learning.
  • GenCorrect iteratively generates, evaluates, and refines diverse candidate solutions using additional test-time compute.
  • On IOI 2025, Nano-CC rose from 130 to 291 points after post-training and reached 468 with GenCorrect; Ultra-CC scored 502.
  • On IOI 2026, the competition-specific Ultra-CC system exceeded the top human score of 498.27 while observing the same time, internet-access, and submission constraints.

Sources (1)

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

arXiv cs.LG Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar, Boris Ginsburg 2026-09-02 arXiv:2609.02849
Public signals Hugging Face upvotes 11 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 11 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:24:32.970623 UTC

TL;DR - Researchers developed a post-training and test-time refinement pipeline that enabled Nemotron models to achieve gold-medal-level competitive programming performance. Their specialized Ultra-CC system scored 535.4/600 on IOI 2026, surpassing both the gold threshold and the top human score under equivalent competition constraints.

  • The pipeline combines 22,000 curated problems, synthetic reasoning traces, supervised fine-tuning, and reinforcement learning.
  • GenCorrect iteratively generates, evaluates, and refines diverse candidate solutions using additional test-time compute.
  • On IOI 2025, Nano-CC rose from 130 to 291 points after post-training and reached 468 with GenCorrect; Ultra-CC scored 502.
  • On IOI 2026, the competition-specific Ultra-CC system exceeded the top human score of 498.27 while observing the same time, internet-access, and submission constraints.
item →