🛰️ Daily AI Frontier
‹ back to 2026-08-24

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Research LLM Agents

Ranking

Overall 82
Content 100
Popularity 39

Observed public metrics from 1 member.

Representative image for Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Merged summary

TL;DR - TOWN-VLA is a prompt-authority interface that lets retrieval modify a frozen vision-language-action policy only when a compatibility rule approves a canonical instruction. It prevents prompt-form collapse and improves simulated and physical robot manipulation without retraining.

  • Appending text indiscriminately reduced mean success from 92.47% to 3.00%; both meaningful and length-matched meaningless additions failed on all 500 audited states.
  • TOWN-VLA restores the exact base prompt when intervention is unauthorized; across 900 routes, 525 matched the base prompt by hash and 375 authorized prompts preserved the task signature.
  • On LIBERO-Plus, success increased from 69.5% to 73.1% across 10,030 episodes per method, improving on all four suites and six perturbation axes.
  • On a physical PiPER arm, success rose from 52.7% to 78.7% over 150 trials per method.

Sources (1)

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

arXiv cs.RO Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li, Chenwei Wang, Huiming Chen, Xiaojun Zhu 2026-08-24 arXiv:2608.23224
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-13 14:09:54.340888 UTC

TL;DR - TOWN-VLA is a prompt-authority interface that lets retrieval modify a frozen vision-language-action policy only when a compatibility rule approves a canonical instruction. It prevents prompt-form collapse and improves simulated and physical robot manipulation without retraining.

  • Appending text indiscriminately reduced mean success from 92.47% to 3.00%; both meaningful and length-matched meaningless additions failed on all 500 audited states.
  • TOWN-VLA restores the exact base prompt when intervention is unauthorized; across 900 routes, 525 matched the base prompt by hash and 375 authorized prompts preserved the task signature.
  • On LIBERO-Plus, success increased from 69.5% to 73.1% across 10,030 episodes per method, improving on all four suites and six perturbation axes.
  • On a physical PiPER arm, success rose from 52.7% to 78.7% over 150 trials per method.
item →