CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - CL4D is a contrastively pretrained vision encoder that aligns dynamic 4D point clouds with language, enabling geometric and temporal reasoning without relying on images or video. Its companion model, 4DVLM, uses these representations for language generation and reportedly surpasses prior 4D methods and frontier video VLMs on evaluated tasks.
- CL4D jointly models spatial geometry and motion evolution directly from dynamic point clouds.
- Contrastive language-4D pretraining supports zero-shot motion-to-text and text-to-motion retrieval.
- The authors introduce DynAction4D, a dataset covering human motions, object interactions, and varied environments.
- Across multiple 4D action benchmarks, CL4D reports an approximately 16.75% improvement over prior methods; 4DVLM also outperforms Gemini and GPT-5 given corresponding RGB videos.
Sources (1)
CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
TL;DR - CL4D is a contrastively pretrained vision encoder that aligns dynamic 4D point clouds with language, enabling geometric and temporal reasoning without relying on images or video. Its companion model, 4DVLM, uses these representations for language generation and reportedly surpasses prior 4D methods and frontier video VLMs on evaluated tasks.
- CL4D jointly models spatial geometry and motion evolution directly from dynamic point clouds.
- Contrastive language-4D pretraining supports zero-shot motion-to-text and text-to-motion retrieval.
- The authors introduce DynAction4D, a dataset covering human motions, object interactions, and varied environments.
- Across multiple 4D action benchmarks, CL4D reports an approximately 16.75% improvement over prior methods; 4DVLM also outperforms Gemini and GPT-5 given corresponding RGB videos.