🛰️ Daily AI Frontier
‹ back to 2026-08-18

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 39

Observed public metrics from 1 member.

Merged summary

TL;DR - Reconstruction is a leakage-resistant benchmark testing whether LLMs can recover a paper’s research idea using only its pre-publication bibliography. Single models perform poorly, while a multi-agent review and tournament pipeline substantially improves idea matching.

  • Covers 643 papers across six scientific domains with temporal cutoffs, anonymized references, and frozen bibliographies.
  • Seven frontier models achieve only about 3–15% Match rates individually.
  • A reference-only top-four multi-agent pipeline reaches approximately 23–42% without web search.
  • Cross-model review and tournament selection yield an observed 2.4Ă— improvement over the best single-model baseline.

Sources (1)

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

arXiv cs.AI Shaolong Chen, Yanlin Fei, Nazhou Liu, Xinmiao Yu, Lei Li, Rahul Thapa, Madalina Ciobanu, Qingqing Mao, Ritankar Das 2026-08-17 arXiv:2608.16645
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-14 14:21:44.677808 UTC

TL;DR - Reconstruction is a leakage-resistant benchmark testing whether LLMs can recover a paper’s research idea using only its pre-publication bibliography. Single models perform poorly, while a multi-agent review and tournament pipeline substantially improves idea matching.

  • Covers 643 papers across six scientific domains with temporal cutoffs, anonymized references, and frozen bibliographies.
  • Seven frontier models achieve only about 3–15% Match rates individually.
  • A reference-only top-four multi-agent pipeline reaches approximately 23–42% without web search.
  • Cross-model review and tournament selection yield an observed 2.4Ă— improvement over the best single-model baseline.
item →