Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
TL;DR - An expert-built benchmark tests whether frontier LLMs can extract and critically appraise evidence from microbial oncogenesis papers, using MMTV-LV and breast cancer as a case study. GPT-5 and GPT-5 Nano produced agreement distributions indistinguishable from human domain experts, supporting LLM-driven automated systematic evidence synthesis.
- Benchmark: 24 research papers, 77 question items spanning MCQ, Likert-scale, multi-select, and free-text formats, with a structured extraction/appraisal template built by recruited domain experts.
- Evaluation method: novel per-question agreement metrics comparing inter-expert agreement against expert-LLM agreement, testing whether an LLM behaves as "another expert" by maintaining or increasing agreement; free-text answers additionally scored qualitatively.
- Models compared: Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano. GPT-5 and GPT-5 Nano matched experts; Gemini models were similar but significantly more lenient in applying microbial oncogenicity criteria. Hallucinations were rare.
- Remaining weaknesses: methodological quality appraisal and detecting contradictions within full texts — the tasks requiring deeper reasoning over whole papers rather than fact extraction.