FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
Ranking
Overall
82
Content
100
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - FormalTCS is an expert-validated benchmark of 175 frontier theoretical computer science problems with Lean formalizations, designed to test LLMs across the full research pipeline. Results show that autoformalization and selecting worthwhile research claims remain major barriers to autonomous TCS research.
- Instances come from papers accepted to STOC, FOCS, SODA, and COLT in 2025–2026 and preserve paper-specific definitions, assumptions, and proof dependencies.
- Autoformalization was the sharpest bottleneck: the best model scored 11.5 when translating natural-language claims into formal statements.
- Models performed better when given human-written formal statements, reaching 28.6 Pass@8 on theorem proving.
- An automated claim-generation and proving framework produced 64 claims, but only 6 passed both expert evaluation and proof verification.
Sources (1)
FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - FormalTCS is an expert-validated benchmark of 175 frontier theoretical computer science problems with Lean formalizations, designed to test LLMs across the full research pipeline. Results show that autoformalization and selecting worthwhile research claims remain major barriers to autonomous TCS research.
- Instances come from papers accepted to STOC, FOCS, SODA, and COLT in 2025–2026 and preserve paper-specific definitions, assumptions, and proof dependencies.
- Autoformalization was the sharpest bottleneck: the best model scored 11.5 when translating natural-language claims into formal statements.
- Models performed better when given human-written formal statements, reaching 28.6 Pass@8 on theorem proving.
- An automated claim-generation and proving framework produced 64 claims, but only 6 passed both expert evaluation and proof verification.