🛰️ Daily AI Frontier
‹ back to 2026-08-27

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

Research Multimodal & Generative

Ranking

Overall 83
Content 100
Popularity 42

Observed public metrics from 1 member.

Merged summary

TL;DR - LongVU-TTT introduces a causal test-time-training resampler that adapts visual features to each long video before compressing them for a multimodal LLM. It processes up to 512 frames into 128 LLM frames while retaining selected evidence needed for long-range reasoning.

  • Grouped 2D convolutional fast weights aggregate temporal context between the vision encoder and LLM.
  • A hybrid uniform and change-aware selector explicitly preserves frames because fast-weight benefits weaken as evidence becomes more distant.
  • TTT-Conv outperforms TTT-MLP by up to 2.12 points and bidirectional Mamba2 by up to 3.04 points on MLVU under controlled conditions.
  • The approach is competitive across five video-understanding benchmarks and stronger than attention-based and fixed-state recurrent resamplers across three.

Sources (1)

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

arXiv cs.CV Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs, Mathis Bode, Mohamed Elhoseiny 2026-08-26 arXiv:2608.25729
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-12 14:20:13.431285 UTC

TL;DR - LongVU-TTT introduces a causal test-time-training resampler that adapts visual features to each long video before compressing them for a multimodal LLM. It processes up to 512 frames into 128 LLM frames while retaining selected evidence needed for long-range reasoning.

  • Grouped 2D convolutional fast weights aggregate temporal context between the vision encoder and LLM.
  • A hybrid uniform and change-aware selector explicitly preserves frames because fast-weight benefits weaken as evidence becomes more distant.
  • TTT-Conv outperforms TTT-MLP by up to 2.12 points and bidirectional Mamba2 by up to 3.04 points on MLVU under controlled conditions.
  • The approach is competitive across five video-understanding benchmarks and stronger than attention-based and fixed-state recurrent resamplers across three.
item →