Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation
Ranking
Overall
82
Content
100
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - TOWN-VLA is a prompt-authority interface that lets retrieval modify a frozen vision-language-action policy only when a compatibility rule approves a canonical instruction. It prevents prompt-form collapse and improves simulated and physical robot manipulation without retraining.
- Appending text indiscriminately reduced mean success from 92.47% to 3.00%; both meaningful and length-matched meaningless additions failed on all 500 audited states.
- TOWN-VLA restores the exact base prompt when intervention is unauthorized; across 900 routes, 525 matched the base prompt by hash and 375 authorized prompts preserved the task signature.
- On LIBERO-Plus, success increased from 69.5% to 73.1% across 10,030 episodes per method, improving on all four suites and six perturbation axes.
- On a physical PiPER arm, success rose from 52.7% to 78.7% over 150 trials per method.
Sources (1)
Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - TOWN-VLA is a prompt-authority interface that lets retrieval modify a frozen vision-language-action policy only when a compatibility rule approves a canonical instruction. It prevents prompt-form collapse and improves simulated and physical robot manipulation without retraining.
- Appending text indiscriminately reduced mean success from 92.47% to 3.00%; both meaningful and length-matched meaningless additions failed on all 500 audited states.
- TOWN-VLA restores the exact base prompt when intervention is unauthorized; across 900 routes, 525 matched the base prompt by hash and 375 authorized prompts preserved the task signature.
- On LIBERO-Plus, success increased from 69.5% to 73.1% across 10,030 episodes per method, improving on all four suites and six perturbation axes.
- On a physical PiPER arm, success rose from 52.7% to 78.7% over 150 trials per method.