Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
TL;DR - ROBORMBENCH reveals that vision-language reward models can assign contradictory rewards to identical robot trajectories when goal instructions are merely paraphrased. This fragility threatens the reliability of VLM-guided robot learning.
- The benchmark contains 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases across lexical, syntactic, and action-goal rewrites.
- Paraphrases can substantially change predicted progress scores and even flip the same behavior between failure and success.
- Instability affects both proprietary and open-source VLMs, worsens with more divergent rewrites, and is not consistently mitigated by model scale or explicit reasoning.
- Dedicated reward models using trajectory-grounded supervision are substantially more stable.