EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models
TL;DR - EviSafe is an evidence-grounded framework and benchmark that tests whether vision-language models make safe decisions based on the correct textual and visual evidence. Results across 11 VLMs reveal large gaps between apparently safe responses and genuinely grounded safety reasoning.
- EviSafeBench contains 1,181 gold image-text scenarios and 2,452 targeted counterfactual variants spanning eight safety domains and eight risk-source types.
- Its three-probe protocol evaluates natural responses, evidence reporting, and reactions to counterfactual changes in safety-critical evidence.
- Natural severity accuracy ranged from 27.6% to 52.8%, while relaxed diagnostic consistency reached only 6.1% to 29.3%.
- Unsafe-to-safe counterfactual transition success ranged from 30.4% to 58.4%, motivating evaluation beyond refusal rates alone.