🛰️ Daily AI Frontier
‹ back to 2026-08-20

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

arXiv cs.LG Bioinformatics AI Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov 2026-08-19
Representative image for Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

TL;DR - C3LM is a chemical-plausibility-aware language model for single-step retrosynthesis, trained on roughly 45.6 million verified reactions with a Top-K prediction paradigm. It achieves state-of-the-art results on an out-of-distribution benchmark and supports more diverse, plausible synthesis planning.

  • Top-K prompting addresses retrosynthesis’s one-to-many nature by producing multiple plausible reaction predictions rather than optimizing for one answer.
  • Training combines fine-tuning with ChemCensor-based chemical-plausibility rewards and novelty-oriented rewards.
  • C3LM achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark.
  • LLMs and conventional models explore complementary reaction spaces, suggesting potential gains from ensemble systems.

view merged work →