Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
TL;DR - An audit of 22 frontier LLMs across 12 molecular regression benchmarks finds widespread, dataset-specific retrieval of exact published values, complicating claims of genuine predictive accuracy. Higher reasoning levels increased retrieval flags by 89%, showing that evaluation settings can substantially affect measured contamination.
- More than half of the tested LLMs exhibited verbatim retrieval on five datasets; retrieval was isolated on the other seven.
- Some strong models recognized transformed SMILES paired with original labels, suggesting memorization can persist despite input transformations.
- Suppressing retrieval made models’ prediction errors more similar, while differing reliance on memorized values exaggerated performance gaps.
- Molecular benchmark accuracy alone cannot distinguish property prediction from recall of published results.