🛰️ Daily AI Frontier
‹ back to 2026-08-21

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

Research LLMs & Foundation Models

Ranking

Overall 82
Content 100
Popularity 39

Observed public metrics from 1 member.

Merged summary

TL;DR - FormalTCS is an expert-validated benchmark of 175 frontier theoretical computer science problems with Lean formalizations, designed to test LLMs across the full research pipeline. Results show that autoformalization and selecting worthwhile research claims remain major barriers to autonomous TCS research.

  • Instances come from papers accepted to STOC, FOCS, SODA, and COLT in 2025–2026 and preserve paper-specific definitions, assumptions, and proof dependencies.
  • Autoformalization was the sharpest bottleneck: the best model scored 11.5 when translating natural-language claims into formal statements.
  • Models performed better when given human-written formal statements, reaching 28.6 Pass@8 on theorem proving.
  • An automated claim-generation and proving framework produced 64 claims, but only 6 passed both expert evaluation and proof verification.

Sources (1)

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

arXiv cs.CL Dingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu, Wanxiang Che 2026-08-20 arXiv:2608.20153
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-19 14:25:41.775960 UTC

TL;DR - FormalTCS is an expert-validated benchmark of 175 frontier theoretical computer science problems with Lean formalizations, designed to test LLMs across the full research pipeline. Results show that autoformalization and selecting worthwhile research claims remain major barriers to autonomous TCS research.

  • Instances come from papers accepted to STOC, FOCS, SODA, and COLT in 2025–2026 and preserve paper-specific definitions, assumptions, and proof dependencies.
  • Autoformalization was the sharpest bottleneck: the best model scored 11.5 when translating natural-language claims into formal statements.
  • Models performed better when given human-written formal statements, reaching 28.6 Pass@8 on theorem proving.
  • An automated claim-generation and proving framework produced 64 claims, but only 6 passed both expert evaluation and proof verification.
item →