🛰️ Daily AI Frontier
‹ back to 2026-08-20

ECCV 2026 | 长视频Token剪枝新范式:从关键帧到证据链

Research Efficiency & Systems

Ranking

Overall 74
Content 85
Popularity 47

Observed public metrics from 1 member.

Representative image for ECCV 2026 | 长视频Token剪枝新范式:从关键帧到证据链

Merged summary

TL;DR - SemVID is a training-free visual-token pruning method for long-video temporal grounding that preserves an “evidence chain” rather than isolated keyframes. It reduces inference cost while maintaining accurate event-boundary localization under aggressive token compression.

  • SemVID allocates per-frame token budgets using both query relevance and inter-frame changes, retaining evidence across the event timeline.
  • It preserves complementary object, motion, and context tokens to capture relevant entities, temporal transitions, and scene continuity.
  • Motion tokens serve as cross-frame relay nodes, while MMR selection prevents redundant object patches from consuming the token budget.
  • On Charades-STA and ActivityNet-Grounding with Qwen3-VL and Qwen2.5-VL, SemVID outperformed existing pruning methods in localization and evidence-retention/connectivity metrics at equal budgets, particularly at low retention rates.

Sources (1)

ECCV 2026 | 长视频Token剪枝新范式:从关键帧到证据链

WeChat: PaperWeekly 2026-08-19 arXiv:2603.05663
Public signals Semantic Scholar citations 4 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 4 · Influential citations 0 X · N/A Fetched 2026-09-19 14:26:05.947360 UTC

TL;DR - SemVID is a training-free visual-token pruning method for long-video temporal grounding that preserves an “evidence chain” rather than isolated keyframes. It reduces inference cost while maintaining accurate event-boundary localization under aggressive token compression.

  • SemVID allocates per-frame token budgets using both query relevance and inter-frame changes, retaining evidence across the event timeline.
  • It preserves complementary object, motion, and context tokens to capture relevant entities, temporal transitions, and scene continuity.
  • Motion tokens serve as cross-frame relay nodes, while MMR selection prevents redundant object patches from consuming the token budget.
  • On Charades-STA and ActivityNet-Grounding with Qwen3-VL and Qwen2.5-VL, SemVID outperformed existing pruning methods in localization and evidence-retention/connectivity metrics at equal budgets, particularly at low retention rates.
item →