🛰️ Daily AI Frontier
‹ back to 2026-09-07

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

arXiv cs.RO Robotics AI Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No 2026-09-04

TL;DR - ROBORMBENCH reveals that vision-language reward models can assign contradictory rewards to identical robot trajectories when goal instructions are merely paraphrased. This fragility threatens the reliability of VLM-guided robot learning.

  • The benchmark contains 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases across lexical, syntactic, and action-goal rewrites.
  • Paraphrases can substantially change predicted progress scores and even flip the same behavior between failure and success.
  • Instability affects both proprietary and open-source VLMs, worsens with more divergent rewrites, and is not consistently mitigated by model scale or explicit reasoning.
  • Dedicated reward models using trajectory-grounded supervision are substantially more stable.

view merged work →