Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
Ranking
Overall
84
Content
90
Popularity
69
Observed public metrics from 1 member.
Merged summary
TL;DR - C3LM is a chemical-plausibility-aware language model for single-step retrosynthesis, trained on roughly 45.6 million verified reactions with a Top-K prediction paradigm. It achieves state-of-the-art results on an out-of-distribution benchmark and supports more diverse, plausible synthesis planning.
- Top-K prompting addresses retrosynthesis’s one-to-many nature by producing multiple plausible reaction predictions rather than optimizing for one answer.
- Training combines fine-tuning with ChemCensor-based chemical-plausibility rewards and novelty-oriented rewards.
- C3LM achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark.
- LLMs and conventional models explore complementary reaction spaces, suggesting potential gains from ensemble systems.
Sources (1)
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
Public signals
Hugging Face upvotes 35
TL;DR - C3LM is a chemical-plausibility-aware language model for single-step retrosynthesis, trained on roughly 45.6 million verified reactions with a Top-K prediction paradigm. It achieves state-of-the-art results on an out-of-distribution benchmark and supports more diverse, plausible synthesis planning.
- Top-K prompting addresses retrosynthesis’s one-to-many nature by producing multiple plausible reaction predictions rather than optimizing for one answer.
- Training combines fine-tuning with ChemCensor-based chemical-plausibility rewards and novelty-oriented rewards.
- C3LM achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark.
- LLMs and conventional models explore complementary reaction spaces, suggesting potential gains from ensemble systems.