TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
TL;DR - TRAPSBench shows that vision-language models can internally recognize when visual evidence is insufficient, yet often answer instead of abstaining. This representation–output gap suggests that improving epistemic restraint may require output-stage interventions.
- TRAPSBench contains 1,404 matched video-physics pairs, with targeted changes making outcomes visually undeterminable.
- Across 16 VLMs from five families, spontaneous restraint was poor; the best Penalized Epistemic Calibration Score was 0.292.
- Linear probes decoded answerability from hidden states at up to 0.91 AUROC, while single-layer steering causally altered abstention.
- Models recognized textual impossibility roughly four times more readily than missing visual evidence.