Alignment with experimental data improves protein generative modeling
TL;DR - ProteinDPO uses direct preference optimization to align a pretrained protein language model with experimental stability data, improving its ability to score and generate thermostable proteins. For H5N1 hemagglutinin, it substantially increased thermal stability while preserving antibody recognition.
- Applies DPO using experimentally measured protein stability preferences.
- Supports both thermostability scoring and protein sequence generation.
- Demonstrates improved H5N1 hemagglutinin stability without sacrificing antigen recognition.