🛰️ Daily AI Frontier
‹ back to 2026-08-21

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

arXiv cs.CL LLMs & Foundation Models Dingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu, Wanxiang Che 2026-08-20

TL;DR - FormalTCS is an expert-validated benchmark of 175 frontier theoretical computer science problems with Lean formalizations, designed to test LLMs across the full research pipeline. Results show that autoformalization and selecting worthwhile research claims remain major barriers to autonomous TCS research.

  • Instances come from papers accepted to STOC, FOCS, SODA, and COLT in 2025–2026 and preserve paper-specific definitions, assumptions, and proof dependencies.
  • Autoformalization was the sharpest bottleneck: the best model scored 11.5 when translating natural-language claims into formal statements.
  • Models performed better when given human-written formal statements, reaching 28.6 Pass@8 on theorem proving.
  • An automated claim-generation and proving framework produced 64 claims, but only 6 passed both expert evaluation and proof verification.

view merged work →