🛰️ Daily AI Frontier
‹ back to 2026-08-13

G0.5: One Autoregressive Stream for Robot Reasoning and Action

arXiv cs.RO Robot Foundation Models Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao 2026-08-12
Representative image for G0.5: One Autoregressive Stream for Robot Reasoning and Action

TL;DR - G0.5 is an autoregressive vision-language-action model that generates reasoning and robot actions from one transformer decoder. Its unified design enables prompt-steerable behavior and strong results across seven robotic evaluation regimes.

  • Uses a learned tokenizer to represent actions from heterogeneous robots in a shared vocabulary.
  • Interleaves task decomposition, object grounding, action hints, and action tokens in one chain-of-thought stream.
  • Adds visual memory through the vision encoder to incorporate multi-second histories.
  • Outperforms cited baselines across real-world fine-tuning, long-horizon manipulation, zero-shot transfer, and simulation benchmarks.

view merged work →