Drug firms’ secret data supercharge AI protein models
TL;DR - An AI protein-modeling system trained on more than 20,000 proprietary pharmaceutical-company structures reportedly outperforms AlphaFold-like models trained only on public data. The result highlights how access to private structural datasets can materially improve computational biology models.
- Training incorporated previously secret protein-structure data from drug companies.
- The proprietary dataset contains more than 20,000 structures.
- Reported gains are relative to AlphaFold-like systems using public data alone.
- The brief does not provide model architecture, benchmarks, or quantitative performance results.