🛰️ Daily AI Frontier
‹ back to 2026-07-23

Test-Time Training for Modality Order Consistency in Vision-Language Models

Research Multimodal & Generative

Merged summary

TL;DR - Vision-language models perform worse when a question precedes its image, despite equivalent semantics. An asymmetric test-time training method reduces this modality-order gap and can also improve the stronger image-first baseline.

  • The failure appeared consistently across three models and three benchmarks.
  • Activation patching localized order-dependent representation divergence to a narrow mid-network region.
  • Test-time adaptation repaired the misalignment across layers and improved consistency between both prompt orders.
  • The findings frame modality-order sensitivity as a circuit-level VLM failure rather than a semantic one.

Sources (1)

Test-Time Training for Modality Order Consistency in Vision-Language Models

arXiv cs.CV Aditi Gupta, Yossi Gandelsman 2026-07-22 arXiv:2607.20351

TL;DR - Vision-language models perform worse when a question precedes its image, despite equivalent semantics. An asymmetric test-time training method reduces this modality-order gap and can also improve the stronger image-first baseline.

  • The failure appeared consistently across three models and three benchmarks.
  • Activation patching localized order-dependent representation divergence to a narrow mid-network region.
  • Test-time adaptation repaired the misalignment across layers and improved consistency between both prompt orders.
  • The findings frame modality-order sensitivity as a circuit-level VLM failure rather than a semantic one.
item →