🛰️ Daily AI Frontier
‹ back to 2026-07-27

机器人需要一个「思考系统」:τ0-VLA让具身智能迈向长程任务时代

WeChat: 机器之心 Embodied AI 2026-07-27
Representative image for 机器人需要一个「思考系统」:τ0-VLA让具身智能迈向长程任务时代

TL;DR - τ0-VLA is a hierarchical vision-language-action model for long-horizon robotic tasks that separates high-level planning from low-level control. Its world-model-guided test-time computation helps robots plan, execute, and correct multi-step tasks in real environments.

  • A “slow thinking, fast execution” architecture combines subtask planning and memory with high-frequency closed-loop control.
  • High-level planning uses proposal, world, value, and reflection models to predict and rank future outcomes via subtask-level beam search.
  • The model was pretrained on 40,115 hours of real-world interaction data, including more than 20,000 hours of physical-robot data across multiple platforms.
  • On AGIBOT G1 long-horizon tasks, hierarchical planning raised average success from 27.5% to 45.0% and task progress from 80.10% to 87.85%.

view merged work →