IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals
TL;DR - IntroConformal is a training-free conformal risk control framework that uses a vision-language model’s internal signals to provide finite-sample, distribution-free factuality guarantees. It reduces reliance on external verifiers and better handles confidently incorrect outputs.
- Derives conformity scores from layer-wise semantic stability in hidden-state representations.
- Introduces verification probability, based on the model’s self-assessment of claim factuality, as a stronger introspective score.
- Satisfies conformal risk guarantees across multiple large vision-language model architectures.
- Reduces abstention while matching or exceeding external-verifier baselines in claim-level discrimination.