🛰️ Daily AI Frontier
‹ back to 2026-07-28

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Research Multimodal & Generative

Ranking

Overall 70
Content 70
Popularity 70

Observed public metrics from 1 member.

Representative image for Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Merged summary

TL;DR - SCDT is a denoising transformer for RGB-thermal tracking that reconstructs missing modality information and strengthens weak features using spatial and temporal context. It matters because one model handles both incomplete and complete inputs without architecture or parameter changes.

  • Combines recent-frame cues with long-term modality evolution for temporally consistent representations.
  • Progressively denoises available-modality features to recover reliable multimodal information.
  • Uses noise-modulated adaptation to adjust dynamically to modality availability.
  • Reportedly outperforms prior methods across three public benchmarks.

Sources (1)

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

arXiv cs.CV Andong Lu, Ziyi Zha, Jiandong Jin, Shihao Li, Chenglong Li, Jin Tang, Bin Luo 2026-07-27 arXiv:2607.24701
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-08-24 14:32:57.141669 UTC

TL;DR - SCDT is a denoising transformer for RGB-thermal tracking that reconstructs missing modality information and strengthens weak features using spatial and temporal context. It matters because one model handles both incomplete and complete inputs without architecture or parameter changes.

  • Combines recent-frame cues with long-term modality evolution for temporally consistent representations.
  • Progressively denoises available-modality features to recover reliable multimodal information.
  • Uses noise-modulated adaptation to adjust dynamically to modality availability.
  • Reportedly outperforms prior methods across three public benchmarks.
item →