Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking
TL;DR - SCDT is a denoising transformer for RGB-thermal tracking that reconstructs missing modality information and strengthens weak features using spatial and temporal context. It matters because one model handles both incomplete and complete inputs without architecture or parameter changes.
- Combines recent-frame cues with long-term modality evolution for temporally consistent representations.
- Progressively denoises available-modality features to recover reliable multimodal information.
- Uses noise-modulated adaptation to adjust dynamically to modality availability.
- Reportedly outperforms prior methods across three public benchmarks.