5篇论文被ACM MM 2026录用
TL;DR - ACM MM 2026 accepted five highlighted papers spanning multimodal scene-text spotting, robust watermarking, social-media prediction, visual-token compression, child-oriented AIGC video safety, and explainable continual learning.
- SPaTS/SPaSO uses single-token grounding and reinforcement-optimized patch selection to improve scene-text localization and recognition.
- DualAlign jointly aligns watermark embedding and extraction, achieving robustness after training with only two distortion types.
- VisionSelector learns adaptive visual-token selection, reporting 90% compression, a 12-point gain over heuristic baselines, and 1.44× faster prefill.
- Other work introduces relation-enhanced RAG for popularity prediction, a child-focused AIGC video-risk benchmark and agentic evaluator, and explainable class-incremental learning.