Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
TL;DR - SEE is a multimodal, expert-curated benchmark testing whether MLLMs can draw evidence-bounded inferences from real experimental data in chemistry, biology, and materials science. Even the best of 19 models scores only 48.7%, showing current models fall well short of supporting real laboratory discovery.
- Questions are grounded in peer-reviewed literature and actual experimental practice, targeting inference from experimental results rather than recall of established concepts.
- Across 19 MLLMs, top accuracy is 48.7%; general-purpose models outperform science-specialized models on average.
- A visual-agent setting with tool use raises best accuracy to 52.7%, but the authors note more information does not translate into reliable scientific reasoning.
- The core failure mode identified is managing tool-derived information within the boundaries of the original experimental evidence — i.e., justified, evidence-bounded inference.