🛰️ Daily AI Frontier
‹ back to 2026-08-20

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

arXiv cs.CV Multimodal & Generative Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella, Dulanga Weerakoon, Vigneshwaran Subbaraju, Ranga Rodrigo 2026-08-19
Representative image for CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

TL;DR - CL4D is a contrastively pretrained vision encoder that aligns dynamic 4D point clouds with language, enabling geometric and temporal reasoning without relying on images or video. Its companion model, 4DVLM, uses these representations for language generation and reportedly surpasses prior 4D methods and frontier video VLMs on evaluated tasks.

  • CL4D jointly models spatial geometry and motion evolution directly from dynamic point clouds.
  • Contrastive language-4D pretraining supports zero-shot motion-to-text and text-to-motion retrieval.
  • The authors introduce DynAction4D, a dataset covering human motions, object interactions, and varied environments.
  • Across multiple 4D action benchmarks, CL4D reports an approximately 16.75% improvement over prior methods; 4DVLM also outperforms Gemini and GPT-5 given corresponding RGB videos.

view merged work →