🛰️ Daily AI Frontier
‹ back to 2026-08-10

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

arXiv cs.AI AI for Science Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao 2026-08-07
Representative image for Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

TL;DR - SEE is a multimodal, expert-curated benchmark testing whether MLLMs can draw evidence-bounded inferences from real experimental data in chemistry, biology, and materials science. Even the best of 19 models scores only 48.7%, showing current models fall well short of supporting real laboratory discovery.

  • Questions are grounded in peer-reviewed literature and actual experimental practice, targeting inference from experimental results rather than recall of established concepts.
  • Across 19 MLLMs, top accuracy is 48.7%; general-purpose models outperform science-specialized models on average.
  • A visual-agent setting with tool use raises best accuracy to 52.7%, but the authors note more information does not translate into reliable scientific reasoning.
  • The core failure mode identified is managing tool-derived information within the boundaries of the original experimental evidence — i.e., justified, evidence-bounded inference.

view merged work →