BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
Merged summary
TL;DR - BioSecBench-Surveillance is a 100-evaluation benchmark measuring whether AI agents can design pathogen genomic-surveillance pipelines from raw sequencing data and context. Leading systems solved only about half the evaluations, highlighting significant reliability gaps.
- Tasks cover seven areas, including taxonomic classification and genetic-engineering detection.
- The benchmark deterministically grades structured answers using analyst-equivalent inputs.
- Opus 4.8 with PI and GPT-5.5 with Codex tied for the highest score at 50.2%.
- Errors often involved reference selection, thresholds, filtering, and normalization—even when agents chose the correct overall workflow.
Sources (1)
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
TL;DR - BioSecBench-Surveillance is a 100-evaluation benchmark measuring whether AI agents can design pathogen genomic-surveillance pipelines from raw sequencing data and context. Leading systems solved only about half the evaluations, highlighting significant reliability gaps.
- Tasks cover seven areas, including taxonomic classification and genetic-engineering detection.
- The benchmark deterministically grades structured answers using analyst-equivalent inputs.
- Opus 4.8 with PI and GPT-5.5 with Codex tied for the highest score at 50.2%.
- Errors often involved reference selection, thresholds, filtering, and normalization—even when agents chose the correct overall workflow.