PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics
TL;DR - PointZero learns transferable 3D dynamics by completing point trajectories from RGB-D observations and sparse partial tracks, avoiding the need for robot action labels during pre-training. This enables broader data use and improves downstream 3D prediction and robot manipulation.
- Pre-training uses 2.9 million synthetic frames spanning deformable, articulated, and rigid objects.
- A transformer predicts future 3D trajectories for all observed points from a single RGB-D frame and sparse tracks.
- After action-conditioned fine-tuning, PointZero outperforms baselines on the PGND 3D dynamics benchmark.
- For imitation learning, it matches or exceeds baselines on 6 of 7 simulated and real-world manipulation tasks.