DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
TL;DR - DeCAL is a vision-language-action model for contact-rich dexterous manipulation that adaptively combines visual and tactile sensing while learning their joint physical dynamics. It reports state-of-the-art results across evaluated tasks, including a 71% average success rate and strong generalization to unseen scenarios.
- Uses a Mixture-of-Transformers architecture with specialized experts for understanding, latent imagination, and action generation.
- Introduces contact-aware gating to dynamically regulate tactile input during visuo-tactile fusion.
- Jointly models visual and tactile dynamics through latent co-imagination, giving the policy implicit knowledge of physical interactions.
- Achieves a 71% average success rate and an 83.4% progress success rate across the reported tasks.