🛰️ Daily AI Frontier
‹ back to 2026-08-14

5篇论文被ACM MM 2026录用

Research Multimodal & Generative

Ranking

Overall 62
Content 75
Popularity 33

Observed public metrics from 1 member.

Representative image for 5篇论文被ACM MM 2026录用

Merged summary

TL;DR - ACM MM 2026 accepted five highlighted papers spanning multimodal scene-text spotting, robust watermarking, social-media prediction, visual-token compression, child-oriented AIGC video safety, and explainable continual learning.

  • SPaTS/SPaSO uses single-token grounding and reinforcement-optimized patch selection to improve scene-text localization and recognition.
  • DualAlign jointly aligns watermark embedding and extraction, achieving robustness after training with only two distortion types.
  • VisionSelector learns adaptive visual-token selection, reporting 90% compression, a 12-point gain over heuristic baselines, and 1.44× faster prefill.
  • Other work introduces relation-enhanced RAG for popularity prediction, a child-focused AIGC video-risk benchmark and agentic evaluator, and explainable class-incremental learning.

Sources (1)

5篇论文被ACM MM 2026录用

WeChat: CVer 2026-08-12 arXiv:2607.19200
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-02 14:21:10.102228 UTC

TL;DR - ACM MM 2026 accepted five highlighted papers spanning multimodal scene-text spotting, robust watermarking, social-media prediction, visual-token compression, child-oriented AIGC video safety, and explainable continual learning.

  • SPaTS/SPaSO uses single-token grounding and reinforcement-optimized patch selection to improve scene-text localization and recognition.
  • DualAlign jointly aligns watermark embedding and extraction, achieving robustness after training with only two distortion types.
  • VisionSelector learns adaptive visual-token selection, reporting 90% compression, a 12-point gain over heuristic baselines, and 1.44× faster prefill.
  • Other work introduces relation-enhanced RAG for popularity prediction, a child-focused AIGC video-risk benchmark and agentic evaluator, and explainable class-incremental learning.
item →