🛰️ Daily AI Frontier
‹ back to 2026-09-09

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

arXiv cs.RO Multimodal & Generative Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang 2026-09-08
Representative image for DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

TL;DR - DeCAL is a vision-language-action model for contact-rich dexterous manipulation that adaptively combines visual and tactile sensing while learning their joint physical dynamics. It reports state-of-the-art results across evaluated tasks, including a 71% average success rate and strong generalization to unseen scenarios.

  • Uses a Mixture-of-Transformers architecture with specialized experts for understanding, latent imagination, and action generation.
  • Introduces contact-aware gating to dynamically regulate tactile input during visuo-tactile fusion.
  • Jointly models visual and tactile dynamics through latent co-imagination, giving the policy implicit knowledge of physical interactions.
  • Achieves a 71% average success rate and an 83.4% progress success rate across the reported tasks.

view merged work →