Test-Time Training for Modality Order Consistency in Vision-Language Models
TL;DR - Vision-language models perform worse when a question precedes its image, despite equivalent semantics. An asymmetric test-time training method reduces this modality-order gap and can also improve the stronger image-first baseline.
- The failure appeared consistently across three models and three benchmarks.
- Activation patching localized order-dependent representation divergence to a narrow mid-network region.
- Test-time adaptation repaired the misalignment across layers and improved consistency between both prompt orders.
- The findings frame modality-order sensitivity as a circuit-level VLM failure rather than a semantic one.