EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?
Ranking
Overall
68
Content
80
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - EpiBench is a closed-book, sequence-only benchmark of 1,609 curated samples testing whether LLMs can reason about antibody epitopes, and it shows current general-purpose models fall short of reliable epitope understanding for antibody drug discovery.
- Data is grounded in structural antibody–antigen contacts, curated functional B-cell assays, and deep mutational scanning escape measurements; scoring is automatic.
- Covers five linked tasks spanning the development workflow: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to limit shortcut exploitation.
- Nine general-purpose LLMs were evaluated with task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection.
- Findings: models capture partial epitope signal but are weak at antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning.
Sources (1)
EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - EpiBench is a closed-book, sequence-only benchmark of 1,609 curated samples testing whether LLMs can reason about antibody epitopes, and it shows current general-purpose models fall short of reliable epitope understanding for antibody drug discovery.
- Data is grounded in structural antibody–antigen contacts, curated functional B-cell assays, and deep mutational scanning escape measurements; scoring is automatic.
- Covers five linked tasks spanning the development workflow: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to limit shortcut exploitation.
- Nine general-purpose LLMs were evaluated with task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection.
- Findings: models capture partial epitope signal but are weak at antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning.