TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
TL;DR - TruthInsightBench evaluates whether autonomous scientific agents can derive evidence-grounded discoveries rather than reproduce known results. Current coding agents analyze and document data competently but fall short on the scientific judgment needed to establish trustworthy claims.
- The benchmark contains 40 blind tasks spanning 10 scientific domains, exposing only neutral objectives and frozen data while withholding source conclusions and expected analysis paths.
- A fixed LLM judge scores claims across six dimensions using 29 artifact-grounded criteria, enabling repeatable automated evaluation without per-task human grading.
- Four coding agents using the same frozen base model clustered narrowly at 58.4–60.3/100, with no statistically reliable pairwise differences.
- Key weaknesses were controls, robustness checks, falsifiability, and cross-dataset generalization, indicating that scientific reasoning—not coding—is the primary bottleneck.