AutoSynthesis: An agentic system for automated meta-analysis
Merged summary
TL;DR - AutoSynthesis is an end-to-end multi-agent system that automates quantitative meta-analysis from a natural-language research question, and its pooled effect estimates closely match expert-conducted meta-analyses, suggesting evidence synthesis could become far more scalable.
- Covers the full evidence-synthesis pipeline: search strategy formulation, literature retrieval, study screening, full-text eligibility, statistic extraction, standardized effect-size computation, and random-effects meta-analysis.
- Adds heterogeneity analysis across moderators and risk-of-bias assessment, and outputs a transparent PRISMA-aligned report.
- In the reported application, it screened 28+ studies and extracted 20+ quantitative claims, with pooled Hedges' g estimates closely agreeing with manual expert meta-analyses.
- Positioned as a tool to scale evidence-based decision-making across science, medicine, education, and policy; note the validation appears limited to a single application, so broader benchmarking is not detailed here.
Sources (1)
AutoSynthesis: An agentic system for automated meta-analysis
TL;DR - AutoSynthesis is an end-to-end multi-agent system that automates quantitative meta-analysis from a natural-language research question, and its pooled effect estimates closely match expert-conducted meta-analyses, suggesting evidence synthesis could become far more scalable.
- Covers the full evidence-synthesis pipeline: search strategy formulation, literature retrieval, study screening, full-text eligibility, statistic extraction, standardized effect-size computation, and random-effects meta-analysis.
- Adds heterogeneity analysis across moderators and risk-of-bias assessment, and outputs a transparent PRISMA-aligned report.
- In the reported application, it screened 28+ studies and extracted 20+ quantitative claims, with pooled Hedges' g estimates closely agreeing with manual expert meta-analyses.
- Positioned as a tool to scale evidence-based decision-making across science, medicine, education, and policy; note the validation appears limited to a single application, so broader benchmarking is not detailed here.