🛰️ Daily AI Frontier
‹ back to 2026-09-09

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

Research Multimodal & Generative

Ranking

Overall 85
Content 95
Popularity 61

Observed public metrics from 1 member.

Representative image for DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

Merged summary

TL;DR - DeCAL is a vision-language-action model for contact-rich dexterous manipulation that adaptively combines visual and tactile sensing while learning their joint physical dynamics. It reports state-of-the-art results across evaluated tasks, including a 71% average success rate and strong generalization to unseen scenarios.

  • Uses a Mixture-of-Transformers architecture with specialized experts for understanding, latent imagination, and action generation.
  • Introduces contact-aware gating to dynamically regulate tactile input during visuo-tactile fusion.
  • Jointly models visual and tactile dynamics through latent co-imagination, giving the policy implicit knowledge of physical interactions.
  • Achieves a 71% average success rate and an 83.4% progress success rate across the reported tasks.

Sources (1)

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

arXiv cs.RO Yankai Fu, Ning Chen, Junkai Zhao, Heng Zhang, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang 2026-09-08 arXiv:2609.09119
Public signals Hugging Face upvotes 1 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 1 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:22:29.428714 UTC

TL;DR - DeCAL is a vision-language-action model for contact-rich dexterous manipulation that adaptively combines visual and tactile sensing while learning their joint physical dynamics. It reports state-of-the-art results across evaluated tasks, including a 71% average success rate and strong generalization to unseen scenarios.

  • Uses a Mixture-of-Transformers architecture with specialized experts for understanding, latent imagination, and action generation.
  • Introduces contact-aware gating to dynamically regulate tactile input during visuo-tactile fusion.
  • Jointly models visual and tactile dynamics through latent co-imagination, giving the policy implicit knowledge of physical interactions.
  • Achieves a 71% average success rate and an 83.4% progress success rate across the reported tasks.
item →