Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - An expert-built benchmark tests whether frontier LLMs can extract and critically appraise evidence from microbial oncogenesis papers, using MMTV-LV and breast cancer as a case study. GPT-5 and GPT-5 Nano produced agreement distributions indistinguishable from human domain experts, supporting LLM-driven automated systematic evidence synthesis.
- Benchmark: 24 research papers, 77 question items spanning MCQ, Likert-scale, multi-select, and free-text formats, with a structured extraction/appraisal template built by recruited domain experts.
- Evaluation method: novel per-question agreement metrics comparing inter-expert agreement against expert-LLM agreement, testing whether an LLM behaves as "another expert" by maintaining or increasing agreement; free-text answers additionally scored qualitatively.
- Models compared: Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano. GPT-5 and GPT-5 Nano matched experts; Gemini models were similar but significantly more lenient in applying microbial oncogenicity criteria. Hallucinations were rare.
- Remaining weaknesses: methodological quality appraisal and detecting contradictions within full texts — the tasks requiring deeper reasoning over whole papers rather than fact extraction.
Sources (1)
Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
TL;DR - An expert-built benchmark tests whether frontier LLMs can extract and critically appraise evidence from microbial oncogenesis papers, using MMTV-LV and breast cancer as a case study. GPT-5 and GPT-5 Nano produced agreement distributions indistinguishable from human domain experts, supporting LLM-driven automated systematic evidence synthesis.
- Benchmark: 24 research papers, 77 question items spanning MCQ, Likert-scale, multi-select, and free-text formats, with a structured extraction/appraisal template built by recruited domain experts.
- Evaluation method: novel per-question agreement metrics comparing inter-expert agreement against expert-LLM agreement, testing whether an LLM behaves as "another expert" by maintaining or increasing agreement; free-text answers additionally scored qualitatively.
- Models compared: Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano. GPT-5 and GPT-5 Nano matched experts; Gemini models were similar but significantly more lenient in applying microbial oncogenicity criteria. Hallucinations were rare.
- Remaining weaknesses: methodological quality appraisal and detecting contradictions within full texts — the tasks requiring deeper reasoning over whole papers rather than fact extraction.