🛰️ Daily AI Frontier
‹ back to 2026-09-03

机械臂做到第10步就容易出错?一个 2B 模型靠动态调整注意力解决了 | IJCAI 2026

Research Embodied AI

Ranking

Overall 73
Content 85
Popularity 46

Observed public metrics from 1 member.

Representative image for 机械臂做到第10步就容易出错?一个 2B 模型靠动态调整注意力解决了 | IJCAI 2026

Merged summary

TL;DR - S²-VLA is a 2B-parameter vision-language-action model that uses a learned belief state and dynamic attention gating to reduce cumulative errors in long-horizon robotic manipulation. It reports 96.4% success on LIBERO-Long while using about 7GB of inference memory, outperforming many 7B–8.5B models.

  • A lightweight recurrent network derives a belief state from action history and joint-sensor feedback, learning task progress and execution quality without explicit stage labels.
  • Its SSGAA module dynamically balances local visual attention, global intent attention, and action self-attention according to the current manipulation stage.
  • Visual weighting rises for precise alignment, intent weighting rises during grasping and subtask transitions, and action self-attention dominates steady movement.
  • The model reportedly runs at 80.8Hz, highlighting stage-adaptive computation as an efficient alternative to simply scaling VLA parameter counts.

Sources (1)

机械臂做到第10步就容易出错?一个 2B 模型靠动态调整注意力解决了 | IJCAI 2026

雷峰网 (AI科技评论) 2026-09-03 arXiv:2606.27872
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-14 14:12:41.222041 UTC

TL;DR - S²-VLA is a 2B-parameter vision-language-action model that uses a learned belief state and dynamic attention gating to reduce cumulative errors in long-horizon robotic manipulation. It reports 96.4% success on LIBERO-Long while using about 7GB of inference memory, outperforming many 7B–8.5B models.

  • A lightweight recurrent network derives a belief state from action history and joint-sensor feedback, learning task progress and execution quality without explicit stage labels.
  • Its SSGAA module dynamically balances local visual attention, global intent attention, and action self-attention according to the current manipulation stage.
  • Visual weighting rises for precise alignment, intent weighting rises during grasping and subtask transitions, and action self-attention dominates steady movement.
  • The model reportedly runs at 80.8Hz, highlighting stage-adaptive computation as an efficient alternative to simply scaling VLA parameter counts.
item →