ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
TL;DR - ToolSciVer trains vision-language models to verify scientific claims using specialized tools for extracting evidence from tables, charts, and high-resolution figure regions. It outperforms prompting- and reinforcement-learning-based tool-use baselines across two datasets and five VLMs.
- Provides type-aware tools for table focusing, chart parsing, and region zooming.
- Uses GRPO with rewards covering correctness, formatting, response length, and efficient, valid tool use.
- Converts dense scientific visuals into explicit evidence relevant to each claim.
- Evaluated on SciVer and MuSciClaims with Qwen, InternVL, and Gemma model families.