RT by @_akhaliq: We just ran the largest reproducibility audit of an AI conference ever attempted…
TL;DR - A retweeted announcement claiming the largest AI-conference reproducibility audit to date, in which 1,200+ participants aimed coding agents at ICML 2026 papers. It matters because it tests whether LLM coding agents can scale peer-review-adjacent verification across an entire conference.
- Scale claimed: 1,200+ participants, 2,000+ ICML 2026 papers "reproduced or falsified" — roughly a third of the conference.
- Method: crowd-directed autonomous coding agents doing the reproduction work rather than manual human replication.
- Results are not yet public; a livestream (9am PT the following day) is slated to announce winners and findings.
- Content is thin — a teaser post only. No per-paper reproduction rates, agent stack, validation protocol, or falsification criteria are given, so the headline numbers are unverified.