🛰️ Daily AI Frontier
‹ back to 2026-08-05

实验室4篇论文被ACM MM 2026录用

Industry & News Multimodal & Generative 🔗 2 sources

Ranking

Overall 72
Content 80
Popularity 54

Observed public metrics from 1 member.

Representative image for 实验室4篇论文被ACM MM 2026录用

Merged summary

TL;DR — 该实验室有4篇论文被ACM MM 2026录用,分别推进长时序预测、道路异常检测、多模态推荐与室内设计推理。

  • F-LLM通过反馈校正和Lipschitz正则化抑制长时预测误差累积。
  • SHIELD增强小型道路危险物检测,AP达到84.10%。
  • D3ER动态融合共享特征与模态专属特征,提升多模态推荐。
  • DART-I无需微调,将空间与审美先验注入冻结的多模态大模型。
  • 更广泛的录用工作还涉及多模态感知、生成式成像、自主智能体、遥感、医学影像,以及新数据集和评测基准。

注:一则来源聚焦该实验室的4篇论文,另一则将其置于涵盖11篇录用论文的更大技术综述中。

Sources (2)

实验室4篇论文被ACM MM 2026录用

WeChat: CVer 2026-08-03
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-04 14:20:27.431192 UTC

TL;DR - A research lab announced four ACM MM 2026 acceptances spanning LLM forecasting, road-anomaly segmentation, multimodal recommendation, and interior-design reasoning.

  • F-LLM uses feedback correction and Lipschitz regularization to limit long-horizon forecasting errors.
  • SHIELD improves detection of small road hazards, reaching 84.10% AP.
  • D3ER dynamically combines shared and modality-specific features for recommendation.
  • DART-I injects spatial and aesthetic priors into frozen multimodal LLMs without fine-tuning.
item →

11篇论文被ACM MM 2026录用

WeChat: CVer 2026-08-05 arXiv:2503.06993
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-08-30 14:25:37.271151 UTC

TL;DR - A roundup of 11 ACM MM 2026 accepted papers spanning multimodal perception, generative imaging, autonomous agents, remote sensing, medical imaging, and visual benchmarks. The collection highlights new datasets, evaluation frameworks, and efficient or training-free methods.

  • New benchmarks target satellite multi-object tracking, RGB-D video saliency, mobile GUI-agent memory, and classical-painting reconstruction.
  • Agent-focused work includes multi-agent audio-visual segmentation and an evaluation of memory failures across 11 mobile GUI agents.
  • New methods address open-vocabulary detection, federated long-tailed learning, diffusion-based lighting adaptation, and hyperspectral change detection.
  • Other contributions cover unified CT reconstruction and scalable multi-view, multi-label feature selection.
item →