🛰️ Daily AI Frontier
‹ back to 2026-08-07

Nat. Commun. | 排序引导学习加速自动化酶工程

WeChat: DrugAI Bioinformatics AI 2026-08-06
Representative image for Nat. Commun. | 排序引导学习加速自动化酶工程

TL;DR - A Nature Communications paper introduces REAP, a closed-loop automated enzyme engineering platform that couples a protein language model with joint rank–regression learning, active learning, and robotic experimentation. It matters because it turns sparse, noisy wet-lab feedback into rapid, data-driven navigation of vast enzyme sequence space.

  • PLM-RankReg: a frozen ESM2 encoder plus a lightweight MLP head trained on both pairwise ranking and absolute activity. On ProteinGym's 212 DMS datasets it beat MSE/Huber/MAE pointwise losses on rank correlation while also lowering normalized prediction error; in few-shot tests it surpassed pointwise baselines on all 8 datasets at 100 training samples.
  • P450 BM3 case: five closed-loop rounds on non-natural substrate deoxypodophyllotoxin C4β-hydroxylation yielded best-variant gains of 2.37×, 6.06×, 9.76×, 25.87×, 44.93×, culminating in a quintuple mutant (S81M/T180L/E207L/A330E/R498V) at 57× — ~16× higher turnover and ~64× better catalytic efficiency. >1300 variants were measured; model hit rate rose from 42.6% to 63.4% by round 3.
  • Generalization to sortase A: a mechanistically and assay-wise distinct enzyme reached ~104× activity, driven mainly by improved apparent substrate affinity rather than turnover, with positive epistasis rising from ~35% (double) to ~87% (quadruple mutants).
  • Throughput and limits: >2000 variants built/screened per week with ~one-week design-build-test-learn cycles, using ADE-MS or fluorescence readouts; authors note dependence on automatable assays, single-objective optimization, and the still-tiny fraction of sequence space explored.

view merged work →