TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
Ranking
Overall
80
Content
100
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - TRAPSBench shows that vision-language models can internally recognize when visual evidence is insufficient, yet often answer instead of abstaining. This representation–output gap suggests that improving epistemic restraint may require output-stage interventions.
- TRAPSBench contains 1,404 matched video-physics pairs, with targeted changes making outcomes visually undeterminable.
- Across 16 VLMs from five families, spontaneous restraint was poor; the best Penalized Epistemic Calibration Score was 0.292.
- Linear probes decoded answerability from hidden states at up to 0.91 AUROC, while single-layer steering causally altered abstention.
- Models recognized textual impossibility roughly four times more readily than missing visual evidence.
Sources (1)
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - TRAPSBench shows that vision-language models can internally recognize when visual evidence is insufficient, yet often answer instead of abstaining. This representation–output gap suggests that improving epistemic restraint may require output-stage interventions.
- TRAPSBench contains 1,404 matched video-physics pairs, with targeted changes making outcomes visually undeterminable.
- Across 16 VLMs from five families, spontaneous restraint was poor; the best Penalized Epistemic Calibration Score was 0.292.
- Linear probes decoded answerability from hidden states at up to 0.91 AUROC, while single-layer steering causally altered abstention.
- Models recognized textual impossibility roughly four times more readily than missing visual evidence.