LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding
Ranking
Overall
83
Content
100
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - LongVU-TTT introduces a causal test-time-training resampler that adapts visual features to each long video before compressing them for a multimodal LLM. It processes up to 512 frames into 128 LLM frames while retaining selected evidence needed for long-range reasoning.
- Grouped 2D convolutional fast weights aggregate temporal context between the vision encoder and LLM.
- A hybrid uniform and change-aware selector explicitly preserves frames because fast-weight benefits weaken as evidence becomes more distant.
- TTT-Conv outperforms TTT-MLP by up to 2.12 points and bidirectional Mamba2 by up to 3.04 points on MLVU under controlled conditions.
- The approach is competitive across five video-understanding benchmarks and stronger than attention-based and fixed-state recurrent resamplers across three.
Sources (1)
LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - LongVU-TTT introduces a causal test-time-training resampler that adapts visual features to each long video before compressing them for a multimodal LLM. It processes up to 512 frames into 128 LLM frames while retaining selected evidence needed for long-range reasoning.
- Grouped 2D convolutional fast weights aggregate temporal context between the vision encoder and LLM.
- A hybrid uniform and change-aware selector explicitly preserves frames because fast-weight benefits weaken as evidence becomes more distant.
- TTT-Conv outperforms TTT-MLP by up to 2.12 points and bidirectional Mamba2 by up to 3.04 points on MLVU under controlled conditions.
- The approach is competitive across five video-understanding benchmarks and stronger than attention-based and fixed-state recurrent resamplers across three.