🛰️ Daily AI Frontier
‹ back to 2026-09-12

AI数学的最后一道高墙,塌了!GPT-6 Astra刷穿FrontierMath Tier 4

量子位 LLMs & Foundation Models henry 2026-09-12
Representative image for AI数学的最后一道高墙,塌了!GPT-6 Astra刷穿FrontierMath Tier 4

TL;DR - GPT-6 Astra reportedly scored 97.6% on FrontierMath Tier 4 and solved its last question never previously answered by AI, meaning every problem in the benchmark has now been solved at least once across model runs. The milestone highlights rapid gains in mathematical reasoning, but tougher open-problem and formally verified evaluations remain largely unsolved.

  • FrontierMath Tier 4 rose from roughly 5% top performance at its July 2025 launch to near-saturation in about 14 months.
  • After auditing and revising the benchmark, Epoch AI retained 43 Tier 4 problems; reported model scores include 83.0% for GPT-5.6 Sol, 90.2% for Claude Fable 5, and 97.6% for Astra.
  • “Saturated” refers to cumulative coverage across different models and attempts—not Astra achieving 100% in a single evaluation.
  • Astra solved only 2 of 68 FrontierMath Erdős problems, indicating substantial room for progress on open questions requiring formally verifiable proofs.

view merged work →