Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
TL;DR - Reconstruction is a leakage-resistant benchmark testing whether LLMs can recover a paper’s research idea using only its pre-publication bibliography. Single models perform poorly, while a multi-agent review and tournament pipeline substantially improves idea matching.
- Covers 643 papers across six scientific domains with temporal cutoffs, anonymized references, and frozen bibliographies.
- Seven frontier models achieve only about 3–15% Match rates individually.
- A reference-only top-four multi-agent pipeline reaches approximately 23–42% without web search.
- Cross-model review and tournament selection yield an observed 2.4Ă— improvement over the best single-model baseline.