🛰️ Daily AI Frontier
‹ back to 2026-09-09

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

Research Humanoid Robotics

Ranking

Overall 85
Content 95
Popularity 63

Observed public metrics from 1 member.

Representative image for TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

Merged summary

TL;DR - TANGO is a whole-body vision-language-action framework that maps natural-language instructions and egocentric RGB observations directly to 29-DoF humanoid joint actions. It enables geometry-aware navigation through cluttered 3D environments and transfers zero-shot from simulation to a real Unitree G1 robot.

  • Coordinates arm placement, torso adjustment, and gait modulation rather than treating navigation as 2D path planning.
  • Uses simulation-generated, dynamically feasible supervision built from path planning, whole-body motion generation, obstacle-aware editing, and reinforcement-learning-based tracking.
  • Achieves state-of-the-art simulation performance in vision-language navigation and outperforms modular baselines on obstacle negotiation.
  • Demonstrates robust real-world traversal without training on real-world navigation data.

Sources (1)

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

arXiv cs.RO Anqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao, Junli Ren, Masayoshi Tomizuka, Dhruv Shah 2026-09-08 arXiv:2609.09158
Public signals Hugging Face upvotes 21
Providers: Hugging Face · Upvotes 21 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:22:30.158279 UTC

TL;DR - TANGO is a whole-body vision-language-action framework that maps natural-language instructions and egocentric RGB observations directly to 29-DoF humanoid joint actions. It enables geometry-aware navigation through cluttered 3D environments and transfers zero-shot from simulation to a real Unitree G1 robot.

  • Coordinates arm placement, torso adjustment, and gait modulation rather than treating navigation as 2D path planning.
  • Uses simulation-generated, dynamically feasible supervision built from path planning, whole-body motion generation, obstacle-aware editing, and reinforcement-learning-based tracking.
  • Achieves state-of-the-art simulation performance in vision-language navigation and outperforms modular baselines on obstacle negotiation.
  • Demonstrates robust real-world traversal without training on real-world navigation data.
item →