TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
TL;DR - TANGO is a whole-body vision-language-action framework that maps natural-language instructions and egocentric RGB observations directly to 29-DoF humanoid joint actions. It enables geometry-aware navigation through cluttered 3D environments and transfers zero-shot from simulation to a real Unitree G1 robot.
- Coordinates arm placement, torso adjustment, and gait modulation rather than treating navigation as 2D path planning.
- Uses simulation-generated, dynamically feasible supervision built from path planning, whole-body motion generation, obstacle-aware editing, and reinforcement-learning-based tracking.
- Achieves state-of-the-art simulation performance in vision-language navigation and outperforms modular baselines on obstacle negotiation.
- Demonstrates robust real-world traversal without training on real-world navigation data.