BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
Ranking
Overall
82
Content
100
Popularity
40
Observed public metrics from 1 member.
Merged summary
TL;DR - BioSecBench-Surveillance is a 100-evaluation benchmark measuring whether AI agents can design pathogen genomic-surveillance pipelines from raw sequencing data and context. Leading systems solved only about half the evaluations, highlighting significant reliability gaps.
- Tasks cover seven areas, including taxonomic classification and genetic-engineering detection.
- The benchmark deterministically grades structured answers using analyst-equivalent inputs.
- Opus 4.8 with PI and GPT-5.5 with Codex tied for the highest score at 50.2%.
- Errors often involved reference selection, thresholds, filtering, and normalization—even when agents chose the correct overall workflow.
Sources (1)
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - BioSecBench-Surveillance is a 100-evaluation benchmark measuring whether AI agents can design pathogen genomic-surveillance pipelines from raw sequencing data and context. Leading systems solved only about half the evaluations, highlighting significant reliability gaps.
- Tasks cover seven areas, including taxonomic classification and genetic-engineering detection.
- The benchmark deterministically grades structured answers using analyst-equivalent inputs.
- Opus 4.8 with PI and GPT-5.5 with Codex tied for the highest score at 50.2%.
- Errors often involved reference selection, thresholds, filtering, and normalization—even when agents chose the correct overall workflow.