综述 | 终身视觉表征:持续自监督学习CSSL系统综述
TL;DR - This survey reviews continual self-supervised learning (CSSL) for vision models, where representations are updated from unlabeled, non-stationary data while limiting catastrophic forgetting. It frames CSSL as a key foundation for lifelong vision and vision-language models.
- Self-supervised objectives often preserve more transferable features than supervised classification and may converge to flatter, more update-resistant solutions.
- The survey organizes CSSL methods into distillation, weight regularization, replay, architectural approaches, model merging, and objective-level adaptation.
- Major challenges include inconsistent evaluation, replay overfitting, stability–plasticity tradeoffs, computational cost, and cross-modal alignment drift.
- Future research should move beyond short, small-scale benchmarks toward budget-aware, long-term continual pretraining of large multimodal models.