🛰️ Daily AI Frontier
‹ back to 2026-07-23

Test-Time Training for Modality Order Consistency in Vision-Language Models

Research Multimodal & Generative

Ranking

Overall 76
Content 90
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - Vision-language models perform worse when a question precedes its image, despite equivalent semantics. An asymmetric test-time training method reduces this modality-order gap and can also improve the stronger image-first baseline.

  • The failure appeared consistently across three models and three benchmarks.
  • Activation patching localized order-dependent representation divergence to a narrow mid-network region.
  • Test-time adaptation repaired the misalignment across layers and improved consistency between both prompt orders.
  • The findings frame modality-order sensitivity as a circuit-level VLM failure rather than a semantic one.

Sources (1)

Test-Time Training for Modality Order Consistency in Vision-Language Models

arXiv cs.CV Aditi Gupta, Yossi Gandelsman 2026-07-22 arXiv:2607.20351
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-06 16:13:00.648089 UTC

TL;DR - Vision-language models perform worse when a question precedes its image, despite equivalent semantics. An asymmetric test-time training method reduces this modality-order gap and can also improve the stronger image-first baseline.

  • The failure appeared consistently across three models and three benchmarks.
  • Activation patching localized order-dependent representation divergence to a narrow mid-network region.
  • Test-time adaptation repaired the misalignment across layers and improved consistency between both prompt orders.
  • The findings frame modality-order sensitivity as a circuit-level VLM failure rather than a semantic one.
item →