🛰️ Daily AI Frontier
‹ back to 2026-09-08

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Research Bioinformatics AI

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - An audit of 22 frontier LLMs across 12 molecular regression benchmarks finds widespread, dataset-specific retrieval of exact published values, complicating claims of genuine predictive accuracy. Higher reasoning levels increased retrieval flags by 89%, showing that evaluation settings can substantially affect measured contamination.

  • More than half of the tested LLMs exhibited verbatim retrieval on five datasets; retrieval was isolated on the other seven.
  • Some strong models recognized transformed SMILES paired with original labels, suggesting memorization can persist despite input transformations.
  • Suppressing retrieval made models’ prediction errors more similar, while differing reliance on memorized values exaggerated performance gaps.
  • Molecular benchmark accuracy alone cannot distinguish property prediction from recall of published results.

Sources (1)

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

arXiv cs.AI Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler 2026-09-04 arXiv:2609.05381
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-22 14:25:06.214480 UTC

TL;DR - An audit of 22 frontier LLMs across 12 molecular regression benchmarks finds widespread, dataset-specific retrieval of exact published values, complicating claims of genuine predictive accuracy. Higher reasoning levels increased retrieval flags by 89%, showing that evaluation settings can substantially affect measured contamination.

  • More than half of the tested LLMs exhibited verbatim retrieval on five datasets; retrieval was isolated on the other seven.
  • Some strong models recognized transformed SMILES paired with original labels, suggesting memorization can persist despite input transformations.
  • Suppressing retrieval made models’ prediction errors more similar, while differing reliance on memorized values exaggerated performance gaps.
  • Molecular benchmark accuracy alone cannot distinguish property prediction from recall of published results.
item →