Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
TL;DR - Exo2EgoPose forecasts egocentric 3D hand poses from visual observations, language instructions, and pose states, using exocentric demonstrations to overcome limited and unstable first-person views. This representation could improve human-to-robot action transfer for manipulation.
- Treats 3D hand pose as a shared representation bridging human and robot actions.
- Reconstructs video- and frame-chunk-level exocentric features to capture spatial context and motion dynamics.
- Progressively integrates these features through attention and adaptive modulation.
- Reports substantial gains on three pose benchmarks and improved transfer performance on CALVIN.