🛰️ Daily AI Frontier
‹ back to 2026-08-15

Vero: Can AI Agents Build Formally Verified Software Repositories?

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 40

Observed public metrics from 1 member.

Merged summary

TL;DR - Vero is the first benchmark for evaluating whether AI agents can jointly implement and formally verify entire multi-module software repositories. The best tested coding agent solved only 27 of 43 instances, showing substantial room for improvement.

  • Includes 43 Lean 4 repository tasks derived from real Python, Dafny, Verus, and Coq projects.
  • Supports proof-only and joint code-and-proof evaluation using fixed APIs and curated specifications.
  • Covers complex domains including cryptographic protocols and distributed systems.
  • Adds an audit mechanism for proving flawed specifications unsatisfiable or reference implementations incorrect.

Sources (1)

Vero: Can AI Agents Build Formally Verified Software Repositories?

arXiv cs.LG Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, Dawn Song 2026-08-13 arXiv:2608.13522
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-14 14:23:32.556207 UTC

TL;DR - Vero is the first benchmark for evaluating whether AI agents can jointly implement and formally verify entire multi-module software repositories. The best tested coding agent solved only 27 of 43 instances, showing substantial room for improvement.

  • Includes 43 Lean 4 repository tasks derived from real Python, Dafny, Verus, and Coq projects.
  • Supports proof-only and joint code-and-proof evaluation using fixed APIs and curated specifications.
  • Covers complex domains including cryptographic protocols and distributed systems.
  • Adds an audit mechanism for proving flawed specifications unsatisfiable or reference implementations incorrect.
item →