onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
TL;DR - onepot-Bench 0 is a proprietary benchmark suite for measuring whether language models can handle synthetic chemistry tasks that matter for real wet-lab execution, built partly on private in-house experimental data to avoid training-corpus contamination. It matters because existing chemistry evals rarely test the decision-making reliability needed in a physical laboratory.
- Three complementary evaluations: ChemAbacus (tool-free cheminformatics literacy and numerical reasoning), SynthRefusal (safety/refusal behavior across benign, controlled, and designer-drug targets), and SynthBench (reaction-outcome prediction and catalyst selection).
- SynthBench uses private experimental data generated in the authors' own lab, explicitly addressing the contamination risk of public-data benchmarks.
- The stated framing is that lab-relevant capability requires both general problem-solving and domain-specific intuition, so the suite targets basic competency, reliability, and deeper chemical knowledge separately.
- No model scores or empirical results are included in the provided abstract — this is a benchmark-description item only.