🛰️ Daily AI Frontier
‹ back to 2026-07-22

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

arXiv cs.AI Bioinformatics AI Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Amanda Darling, Joshua Stallings, David Stern, Shawn Higdon, Claire Duvallet, Bryan Tegomoh, Kenny Workman 2026-07-21

TL;DR - BioSecBench-Surveillance is a 100-evaluation benchmark measuring whether AI agents can design pathogen genomic-surveillance pipelines from raw sequencing data and context. Leading systems solved only about half the evaluations, highlighting significant reliability gaps.

  • Tasks cover seven areas, including taxonomic classification and genetic-engineering detection.
  • The benchmark deterministically grades structured answers using analyst-equivalent inputs.
  • Opus 4.8 with PI and GPT-5.5 with Codex tied for the highest score at 50.2%.
  • Errors often involved reference selection, thresholds, filtering, and normalization—even when agents chose the correct overall workflow.

view merged work →