🛰️ Daily AI Frontier
‹ back to 2026-09-12

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

Research LLM Agents

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Representative image for 2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

Merged summary

TL;DR - 2AM separates memory from robotic control: a multimodal agent retains interaction history and steers an episodically stateless, RGB-only vision-language-action model through language and optional 2D spatial hints. On LIBERO-Mem, it achieves 76.3% average completion, suggesting that richer agent-policy interfaces can enable long-horizon manipulation without embedding memory in the action model or relying on depth and geometric planning.

  • The multimodal agent is the sole holder of task memory, while one action model executes all task-relevant motion.
  • The agent translates history into subtask instructions plus optional 2D grasp, placement, and movement hints at multiple time scales.
  • Training uses structured hint labels, condition dropout, spatial noise, and temporal jitter to make execution robust to imperfect guidance.
  • 2AM reports 76.3% average completion, 63.0% relaxed success, and 11.8% strict success, versus 14.8% completion for the strongest reported baseline.

Sources (1)

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

arXiv cs.RO Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry 2026-09-10 arXiv:2609.11308
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-12 14:12:55.988285 UTC

TL;DR - 2AM separates memory from robotic control: a multimodal agent retains interaction history and steers an episodically stateless, RGB-only vision-language-action model through language and optional 2D spatial hints. On LIBERO-Mem, it achieves 76.3% average completion, suggesting that richer agent-policy interfaces can enable long-horizon manipulation without embedding memory in the action model or relying on depth and geometric planning.

  • The multimodal agent is the sole holder of task memory, while one action model executes all task-relevant motion.
  • The agent translates history into subtask instructions plus optional 2D grasp, placement, and movement hints at multiple time scales.
  • Training uses structured hint labels, condition dropout, spatial noise, and temporal jitter to make execution robust to imperfect guidance.
  • 2AM reports 76.3% average completion, 63.0% relaxed success, and 11.8% strict success, versus 14.8% completion for the strongest reported baseline.
item →