Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - An audit of 22 frontier LLMs across 12 molecular regression benchmarks finds widespread, dataset-specific retrieval of exact published values, complicating claims of genuine predictive accuracy. Higher reasoning levels increased retrieval flags by 89%, showing that evaluation settings can substantially affect measured contamination.
- More than half of the tested LLMs exhibited verbatim retrieval on five datasets; retrieval was isolated on the other seven.
- Some strong models recognized transformed SMILES paired with original labels, suggesting memorization can persist despite input transformations.
- Suppressing retrieval made models’ prediction errors more similar, while differing reliance on memorized values exaggerated performance gaps.
- Molecular benchmark accuracy alone cannot distinguish property prediction from recall of published results.
Sources (1)
Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - An audit of 22 frontier LLMs across 12 molecular regression benchmarks finds widespread, dataset-specific retrieval of exact published values, complicating claims of genuine predictive accuracy. Higher reasoning levels increased retrieval flags by 89%, showing that evaluation settings can substantially affect measured contamination.
- More than half of the tested LLMs exhibited verbatim retrieval on five datasets; retrieval was isolated on the other seven.
- Some strong models recognized transformed SMILES paired with original labels, suggesting memorization can persist despite input transformations.
- Suppressing retrieval made models’ prediction errors more similar, while differing reliance on memorized values exaggerated performance gaps.
- Molecular benchmark accuracy alone cannot distinguish property prediction from recall of published results.